Can AI help us make sense of what we’re really agreeing to?
Executive Summary
In the first Mellonhead Labs experiment, participants explored how large language models (LLMs) handle complex, dense legal documents - specifically the Terms & Conditions from Ulta and Sephora. The aim was to discover what was similar, what was unique, and whether AI tools could help make these documents easier to digest for non-legal readers.
Participants used a range of models - ChatGPT, NotebookLM, and Microsoft Copilot - and tested prompt techniques including chaining, context-setting, role assignment, and comparative formatting. Results varied depending on the model and how prompts were constructed. NotebookLM stood out for its citation fidelity and structure; ChatGPT was conversational but skewed toward user assumptions; and Copilot surprised with emoji-enhanced, UI-focused output.
The group worked collaboratively - some preparing detailed flows, others experimenting live - and contributed shared learnings around prompt design, model behavior, and user bias. This community-led, real-time testing delivered on Mellonhead's mission to provide practical, accessible, people-focused AI education without the hype.
Compare the Terms & Conditions from Ulta and Sephora using AI.
Test different prompting techniques and tools (ChatGPT, NotebookLM, Copilot).
Surface what’s similar, what’s unique—and what matters to real people.
“The goal wasn’t to get it perfect—it was to try, learn, and reflect together.”
Our experiment highlighted several important data related themes present in these documents that take liberty with what consumers voluntarily provide while using these brands' services.
"Loyalty cards... you're getting coupons and in exchange your data is just going to everybody."
"Ulta, for instance, I think didn't hold on to the data that long. It was pretty restricted who could have it."
"Sephora can use your photos on their socials... I don't think anybody submitting a review would reasonably expect that."
“After reading the Ulta policy, I was actually more comfortable. They didn’t hold onto the data that long. It was pretty restricted who could have it.”
✅ Most accurate and structured
✅ Cites sources directly
✅ Summarizes without needing a prompt
“NotebookLM was very different—very clinical and just the answers and citing its sources.”
“It summarized the docs without asking, which helped clarify how I needed to prompt.”
“In it's output, it cited the location in the source document . So I was confident in that.”
“It brought back too much unless I gave it categories—but still more trustworthy.”
💬 Summary:
Best for legal document accuracy. Most structured. High citation fidelity. Requires some prompt refinement to avoid info overload.
✅ Easy to use and conversational
⚠️ Can skew results based on user phrasing
⚠️ Does not cite sources unless explicitly asked
⚠️ May carry over prior context
“ChatGPT became more casual and started to skew answers based on my prompts.”
“It was giving me what it thought I wanted to hear.”
“ChatGPT doesn’t check its work unless you tell it.”
“In a new chat, it found an error it missed before.”
“Without categories, NotebookLM returned everything; ChatGPT reworded it with less control.”
💬 Summary:
Useful for summarization, but less reliable for legal nuance. Requires precise prompts and separate sessions for accuracy.
✅ Simple summaries with visual formatting (emojis, icons)
⚠️ No mention of citations or document sourcing
⚠️ Light on legal nuance, more UX-focused
“Copilot... especially loves putting in little emojis, check marks.”
“It was trying to make this visually stimulating for me.”
💬 Summary:
Good for readability and visual learners. Not built for legal precision. Best used for casual summaries or UX-first experiences.
Clarity, context, and chaining made all the difference.
“Less conversation and more direct instructions gave clearer answers.”
“I told it to reread the doc five times. That changed the output.”
If you're concerned about your face, voice, images, or identity being stored, reused, or monetized, Sephora offers a slightly stronger legal posture for user protection.
Ulta's use of biometric tools and community content grants them more control and introduces higher risk, especially for underage users or those unaware of how their likeness can be used.
Exp.1: Understanding Legalese with LLMs. Debrief