audiotranskription Hintergrund

AI Research

Privacy Policy

“Can I have my interviews processed by an AI system?” is one of the most common questions asked in our webinars. There is no one-size-fits-all answer to that. The key factors are the type of data, the purpose of the research, the legal basis, institutional requirements, and the specific technical processing method.

Qualitative interview data, in particular, often remains personally identifiable even after names and other unique identifiers have been removed. That is why we consider not only data masking but the entire processing workflow: from the consent text to the service used, and on to storage, access, and deletion.

The following articles provide guidance on this topic: from preprints to data masking, covering the prerequisites that always need to be clarified, through a comparison of processing methods, to a workshop report on working locally. They do not replace the review of the respective research project by the relevant authorities.

Data masking through pseudo-anonymization and anonymization is not enough! [Preprint]

When dealing with qualitative interview data, simply replacing names is often not enough to reliably rule out the identification of individuals. Biographical characteristics, case context, rare events, and the combination of various pieces of information can still make re-identification possible.

Based on this, the preprint by Dresing, Pehl, and Krähnke develops methodological, ethical, and data protection criteria for AI-supported processing of qualitative data. The key conclusion: Data masking is an important protective measure, but it does not automatically make the transfer of data to an AI system permissible.

Clarify the following before any AI processing: consent, ethics, and responsibilities

Before deploying an AI system, it must be clarified on what basis personal research data will be processed. In many qualitative research projects, this requires informed consent that specifically describes the intended use of AI.

Participants should be able to understand what data is being processed, for what purpose, whether an external provider is involved, where the processing takes place, and what conditions apply to data retention and deletion.

Depending on the field, project, and institution, an ethics committee and the relevant data protection officers must also be involved. When using external services, it is important to determine what role the provider plays and whether a data processing agreement is required and feasible.

These questions should be addressed during the planning phase of the research project, not just during the analysis phase. We provide additional guidance on how to word consent forms in the guide“Citing & Documenting AI.”

A Comparison of Processing Methods

On-premises: Data remains within your own infrastructure

In a properly configured local environment, interview transcripts are not transferred to an external AI service. Models with open weights can be run directly on your own computer or within a controlled institutional infrastructure.

This provides greater control over the transfer, storage, and deletion of data. The other requirements of the research project (such as consent, access controls, secure backups, and documentation of the models used) still apply, however.

Our workshop report on local hybrid interpreting illustrates what such a setup might look like in practice.

Servers in Germany with a clearly defined contractual framework

Processing data on servers in Germany may be a suitable solution for some research projects. However, the server location alone is not sufficient to make this decision.

Additional factors to be reviewed include, among others, the specific provider, potential subcontractors, access rights, retention and deletion periods, the potential use of the data for model development, and the contractual framework for processing.

Examples of services organized in different ways include the automatic transcription service from Audiotranskription and the GWDG’s Academic Cloud. Whether a service is suitable for a specific research project should be clarified with the relevant data protection officers before using it.

International Cloud-Based LLMs: Not Suitable for Personal Research Data

For interviews involving specific individuals, publicly available cloud services such as ChatGPT, Claude, or Gemini are not a viable solution. It is nearly impossible to keep track of data flows, storage practices, subcontractors, and transfers to third countries, and even prior anonymization does not reliably prevent qualitative data from being linked to specific individuals.

Such services remain usable for material that does not refer to specific individuals, such as publicly available texts, simulated interviews, or specially created practice materials.
Before each use, it is necessary to verify whether the specific material and the terms of the selected service are consistent with the research project’s data protection policy.

Hybrid Interpretation with Locally Run LLMs [Workshop Report]

The lab report documents a specific local test setup using four LLMs on a MacBook Pro and LM Studio.

This document describes the technical requirements, the selection of models, relevant settings, and findings from fourteen interpretation runs. The focus is not only on technical feasibility, but also on the question of how local models contribute to the focused interpretation of short text passages and where their limitations lie.

The report thus provides a clear framework for researchers who wish to use generative AI without transferring their high-quality data to external cloud services.