The Law and Technology Collaboration on Privacy by Design: A Research Collaboration Report Submitted for Conference Review
The EU General Data Protection Regulation (GDPR) is not self-executing. Knowing what the GDPR requires is one thing; building a system that actually reflects those requirements is another. That rarely happens without lawyers and engineers working together from the start. Our collaboration at the University of Turku was an attempt to make that happen.
Introduction
This post shares the experience of a joint initiative inside the University of Turku between the Faculty of Law and the Faculty of Technology. In spring 2026, the tech team – Jouni Isoaho, Antti Hakkala, Aisvarya Adeseye, and Jiling Zhou – invited us as legal experts to work together to develop an expert-annotated benchmark for an early system design. A local language model was designed and trained on a set of European data protection materials, with legal researchers reviewing and validating the outputs. The experience taught us as much about the limits of the exercise as about its possibilities, and that, from a legal research perspective, is worth sharing.
How It Began: A Meeting in Calonia
The collaboration started with a joint meeting between the legal and technology research teams. The central question was whether a local LLM could engage meaningfully with legal materials, and whether its outputs could be validated against a catalogue built by legal researchers. The ambition was proactive: to get data protection as a fundamental right consideration into the pre-design phase, before architecture decisions are made and compliance becomes remedial. The immediate goal was a conference submission to the 33rd ACM Conference on Computer and Communications Security (ACM CCS 2026 B), with a deeper and more scenario-specific legal analysis to follow. By the time of writing, June 2026, our submission has been accepted for the second round of review.
In the initial session, three agreements came out that shaped everything that followed. First, risk would be understood from the perspective of the data subject, and from a technical perspective, this translates into a constraint on the system. Second, the scope would be strictly limited to privacy, data protection and related security, including EU General Data Protection Regulation 2016/679 (GDPR), EDPB Guidelines 07/2020 on the concepts of controller and processor, DPIA Guidelines WP248 rev.01 — Data Protection Impact Assessment, EDPB Recommendations 01/2020 on measures for transfers to third countries, and Directive (EU) 2022/2555 (NIS2) — Network and Information Security; other legal frameworks would be acknowledged but not analysed. Third, the project would not conduct a risk impact assessment, but the DPIA framework would serve as a structuring lens throughout.
Building the Constraint Catalogue
The collaboration’s methodology was the construction of a shared constraint catalogue, developed across structured components. We, as legal researchers, engaged directly with each component, contributing expertise that the model alone could not supply. Our experience of turning a legal norm into a system constraint was one of the most instructive aspects of the entire process.
The working dataset comprised 50 scenarios in total, 25 drawn from the health domain and 25 from social and messaging platforms, covering both AI-enabled and non-AI cases. Each scenario came with an initial description and a general data flow.
We started reviewing 50 case studies and mapping the applicable legal obligations to each, considering not only what data is processed but how and under what conditions.
Then, under the constraint catalogue, we verified the constraints identified by the tech team. Each constraint in the catalogue was structured based on the type (the compliance category), a description of the requirement, an evidence source traceable to a specific legal provision, an applicability rule defining when and to whom the constraint applies, a baseline risk level, and an applicability basis explaining whether the constraint is always required or context-dependent. This schema was designed to be both legally and technically actionable. We started analyzing and adding obligations that had been missed, removing those that were inapplicable, rephrasing descriptions for legal accuracy, and revising baseline risk levels where appropriate.
Key Findings & Reflections: A Legal Perspective on the Limits of the Analysis
For us, one of the most valuable outputs of the collaboration was a statement of structural limitations. We consider it worth highlighting because it reflects the kind of honest interdisciplinary dialogue this research requires.
Legal compliance is not a point-in-space assessment but an exercise in navigating overlapping, interacting, and sometimes competing normative layers. We identified four structural limitations that affect any approach, including LLM-based analysis, that attempts to assess legal obligations in relative isolation:
First, abstract lawfulness is not concrete lawfulness. A system that satisfies GDPR requirements in the abstract may be unlawful in its specific implementation, depending on architecture, storage location, access controls, pseudonymisation practices, and whether AI processing occurs on-premises or via an API call to a third country.
Second, GDPR compliance is not legal compliance. Satisfying every GDPR requirement does not resolve potential conflicts with medical device regulation, national patient data law, professional secrecy obligations, consumer protection rules, or the EU AI Act. A positive GDPR assessment may create a false sense of assurance.
Third, legal meaning exceeds legal text. The GDPR has applied since 2018, and its interpretation has been substantially shaped by CJEU case law and EDPB guidance. Analysing the Regulation’s articles without accounting for how they have been interpreted, for example, the expansive reading of ‘joint controllership’ under Article 26, or the evolving guidance on Article 22’s scope, risks conclusions that are textually defensible but legally incorrect.
Last but not least, the GDPR is not uniformly harmonised across Member States. Member State choices on health and genetic data, research exemptions, children’s age of consent, and employment-related processing mean that the same GDPR provision may produce different legal outcomes depending on jurisdiction.
We also note that a legal assessment of the full scenario set would, in practice, require a multidisciplinary team of legal professionals with domain-specific expertise in each relevant field. This is not a failure of the methodology; it is an honest account of what law requires.
What Worked Well and What Was Hard
The collaboration succeeded precisely because the two teams approached the same problem from different directions. That difference was productive: as legal researchers, we had to articulate our reasoning in terms that could be operationalised; technology researchers, in turn, had to sit with the fact that legal analysis does not always produce determined answers. Neither team could have produced the constraint catalogue alone.
The most significant methodological challenge was the translation layer. A legal norm lives in language that is deliberately open-textured, shaped over years by guidance documents, court decisions, and regulatory practice. A system constraint needs to be specific enough to implement. Moving from one to the other is never a clean process, and something is always lost in the crossing. That is precisely why the human-in-the-loop design was not a methodological preference but a necessity.
Next Steps
The immediate goal of the collaboration is conference acceptance. Reviewers of the previous submission, written solely from a tech perspective, have requested a strong legal perspective. Together with the tech partners, we responded by developing the constraint catalogue, the scenario mapping, and the reflections on structural limitations.
Future stages of the project will move from the general to the particular: deeper legal analysis of specific scenarios, engagement with sector-specific regulatory frameworks, and a more granular examination of how national law shapes GDPR obligations in the health data domain. We also intend to explore how the methodology can be extended to accommodate the AI Act’s layered requirements for high-risk AI systems.
Colleagues from law faculties with interests in data protection, digital health law, or AI regulation who wish to engage with this work are warmly invited to reach out.
Information about the authors in alphabetical order:

Karimi, Sahar
sahar.s.karimi@utu.fi
Sahar Karimi is a doctoral researcher at the Faculty of Law, University of Turku. She is funded by Finland’s Ministry of Education and Culture’s Doctoral Education Pilot under Decision No. VN/3137/2024-OKM-6 (The Finnish Doctoral Program Network in Artificial Intelligence, AI-DOC project). She works at the intersection of data protection law, AI governance, and EU consumer protection.

Lepinkäinen, Nea
nemaol@utu.fi
Nea Lepinkäinen is a postdoctoral researcher at the Faculty of Law, University of Turku. Her research focuses on the intersection of AI, society, and law.

Maunula, Gail L.
gaimau@utu.fi
Gail Maunula is a doctoral researcher at the Faculty of Law, University of Turku. She is also a practicing legal expert at Privaon OY. She specializes in legal issues involving Sharing Economy.

Sivetc, Liudmila
liusiv@utu.fi
Liudmila Sivetc is a postdoctoral researcher at the Faculty of Law, University of Turku. Her research is supported by the project Dynamics of Digital Rights in Europe, funded by the Research Council of Finland grant (no 362853). She focuses on issues of online free expression and online platform governance from a user-centred perspective.

Sun, Xiaotong
xiaots@utu.fi
Xiaotong Sun is a PhD researcher in AI governance and law at the University of Turku and a Marie Skłodowska-Curie Fellow under the EU Horizon Europe programme (grant agreement No 101177564—HAIF). Her research focuses on the legal and governance challenges of AI, particularly regarding human oversight.
Heading image: AI-generated image (created with ChatGPT image generator, 2026).
