AI-Assisted Privacy Document Analysis

Collecting, analyzing and answering questions about privacy policies with AI

Privacy policies are the main channel through which organizations tell people what happens to their data, yet they are long, intricate, and routinely skipped or misunderstood — and for smart devices they are scattered across manufacturers and e-commerce platforms, so even finding the right policy is hard.

This project applies AI to that document landscape from both ends. PrivacyLens collects, analyzes and publishes smart-device privacy policies at scale, giving consumers, policy authors and regulators a way to understand and compare them. GenAIPABench measures how well generative-AI privacy assistants answer people’s questions about policies and regulations — because genAI’s tendency to produce inaccurate information has to be quantified before such assistants can be trusted.

PrivacyLens — collecting and analyzing smart-device privacy policies

PrivacyLens — collecting and analyzing smart-device privacy policies

A framework that automatically collects, analyzes and publishes the privacy policies of smart IoT devices. It crawls e-commerce sites (and the Wayback Machine for historical versions), uses NLP and machine learning to assess qualities such as readability and ambiguity, and publishes the results monthly — over 1,200 policies covering 7,300 devices so far.

Code: github.com/Aamir7693/privacy-lens

GenAIPABench — benchmarking generative-AI privacy assistants

GenAIPABench — benchmarking generative-AI privacy assistants

A benchmark for generative-AI privacy assistants: annotated questions about privacy policies and data-protection regulations, metrics for the accuracy, relevance and consistency of answers, and a tool that generates prompts to test robustness. Applied to ChatGPT-4, Bard and Bing AI, it showed real promise alongside clear weaknesses on complex queries, consistency and source accuracy.

Code: github.com/Aamir7693/GenAIPABench

Highlights

  • PrivacyLens: automated collection, analysis and monthly publication of smart-device privacy policies — over 1,200 policies for 7,300 devices, with historical versions via the Wayback Machine
  • GenAIPABench: a benchmark for generative-AI privacy assistants — annotated questions on privacy policies and data-protection regulations, metrics for accuracy, relevance and consistency, and a prompt-generation tool
  • Evaluation of ChatGPT-4, Bard and Bing AI as privacy assistants: real promise, but open challenges with complex queries, consistency and source accuracy

Team

  • Aamir HamidPh.D Student, University of Maryland, Baltimore County
  • Hemanth Reddy SamidiMS Student, University of Maryland, Baltimore County
  • Primal PappachanAssistant Professor, Portland State University
  • Tim FininProfessor, University of Maryland, Baltimore County
  • Roberto YusAssistant Professor, University of Maryland, Baltimore County
Roberto Yus
Roberto Yus
Assistant Professor

My research interests include Data Management, Knowledge Representation, the Internet of Things, and Privacy.

Related