AI-Assisted Privacy Document Analysis
Collecting, analyzing and answering questions about privacy policies with AI
Privacy policies are the main channel through which organizations tell people what happens to their data, yet they are long, intricate, and routinely skipped or misunderstood — and for smart devices they are scattered across manufacturers and e-commerce platforms, so even finding the right policy is hard.
This project applies AI to that document landscape from both ends. PrivacyLens collects, analyzes and publishes smart-device privacy policies at scale, giving consumers, policy authors and regulators a way to understand and compare them. GenAIPABench measures how well generative-AI privacy assistants answer people’s questions about policies and regulations — because genAI’s tendency to produce inaccurate information has to be quantified before such assistants can be trusted.
PrivacyLens — collecting and analyzing smart-device privacy policies
A framework that automatically collects, analyzes and publishes the privacy policies of smart IoT devices. It crawls e-commerce sites (and the Wayback Machine for historical versions), uses NLP and machine learning to assess qualities such as readability and ambiguity, and publishes the results monthly — over 1,200 policies covering 7,300 devices so far.
GenAIPABench — benchmarking generative-AI privacy assistants
A benchmark for generative-AI privacy assistants: annotated questions about privacy policies and data-protection regulations, metrics for the accuracy, relevance and consistency of answers, and a tool that generates prompts to test robustness. Applied to ChatGPT-4, Bard and Bing AI, it showed real promise alongside clear weaknesses on complex queries, consistency and source accuracy.
Highlights
- PrivacyLens: automated collection, analysis and monthly publication of smart-device privacy policies — over 1,200 policies for 7,300 devices, with historical versions via the Wayback Machine
- GenAIPABench: a benchmark for generative-AI privacy assistants — annotated questions on privacy policies and data-protection regulations, metrics for accuracy, relevance and consistency, and a prompt-generation tool
- Evaluation of ChatGPT-4, Bard and Bing AI as privacy assistants: real promise, but open challenges with complex queries, consistency and source accuracy
Team
-
Aamir HamidPh.D Student, University of Maryland, Baltimore County
-
Hemanth Reddy SamidiMS Student, University of Maryland, Baltimore County
-
Primal PappachanAssistant Professor, Portland State University
-
Tim FininProfessor, University of Maryland, Baltimore County
-
Roberto YusAssistant Professor, University of Maryland, Baltimore County

