Mood-based discovery
A discovery tool that surfaces archive content by emotional resonance, not keyword match.
San Antonio Review has thousands of pieces in its archive, but WordPress search isn't built for subjective content, tagging doesn't scale, and older work goes unnoticed.
Let readers describe a feeling in their own words and find published work that fits, without the AI overstepping into judgment, therapy or surveillance.
I built a discovery tool that uses emotional resonance as the search primitive, with an LLM suggesting classifications and editors approving every one.
About 800 pieces are now discoverable by feeling, and no classification reaches the index without human review.
The design problem
When content resists categories, keyword search and manual curation both fail.
A poem can be melancholic and hopeful and political at once. Readers searching the archive aren't typing keywords; they're typing feelings, often complicated ones: “My dog died last week. I'm heartbroken but relieved he's no longer suffering.” That isn't a “grief” category. It's an emotional state that resists categorization, and better indexing doesn't solve a vocabulary problem.
The conventional answers (tag everything, add more category pages, curate manually) fail at the levels they need to succeed. Tags are inconsistent across years of contributors and content managers, categories flatten what's interesting about the work, and manual curation works but doesn't scale. An archive becomes invisible the moment its discoverability depends on someone remembering it's there.
“Tag consistently. Rely on the search bar. Curate themed packages. The archive's discoverability is the content manager's work, so we just need to do more of it.”
“Readers don't search keywords; they search feelings. Tags can't capture overlapping moods, and manual curation needs resources the team doesn't have. The problem isn't volume, it's vocabulary.”
What I designed
Embeddings of classification, not text, with content managers as the classification authority.
Two architectural decisions did most of the work. First, every piece is classified offline by an LLM and then reviewed by an editor. The LLM suggests a primary, secondary and optional tertiary mood plus a central theme; the content manager confirms or changes each one.
Second, the embedding stored in the vector index isn't generated from the poem's raw text. It's generated from a structured representation (summary, moods, themes), so the human-reviewed classifications are baked directly into the semantic fingerprint. A search for “bad but relieved about losing a pet” finds pieces whose classification matches that emotional shape, not pieces that happen to share surface vocabulary.
What I deliberately did not build
The places where this tool could have crossed into human judgment, therapy or surveillance, but didn't.
Each decision below is a place where automation would have done more work but made the system worse. Each one defends a relationship: content manager authority, user safety, or user privacy.
- Rejected
Fully automated classificationLLMs are good at the grunt work of suggesting moods and themes, but they aren't reliable judges of what a piece is actually about. Every mood and theme assignment passes through a human before it reaches the index. The tool's discovery quality depends on that decision being made by people.
- Rejected
Therapy mode or emotional adviceThe character is a guide to the archive, not a chatbot or a therapist. When the safety layer detects expressed self-harm or violent intent, the system stops recommending poems and points to crisis resources (988 in the U.S.). It doesn't offer comfort, coping strategies or interpretation of what someone is going through. The boundary is explicit in the system prompt and enforced at the backend, so the assistant can't override it.
- Rejected
Embedding the raw text of piecesRaw-text embeddings would reflect surface word patterns. Embedding a structured representation means similarity reflects the human-reviewed classification: a query about ambivalent grief surfaces pieces classified as ambivalent grief, not pieces that happen to mention the word “grief.”
- Rejected
Long-term conversation storageDiscovery happens in moments of real emotional vulnerability, when people describe grief, anxiety or longing. Storing those conversations would turn that vulnerability into data. Conversations are ephemeral; nothing about a user's emotional state persists in a database.
How it works
A human-curated classification system, a structured embedding pipeline and a conversational interface with a defined character.
Classification and human review
Each piece gets a primary mood, a secondary mood, an optional tertiary mood and a central theme. The vocabulary is deliberately broad, because narrow categories like “sad” and “happy” can't carry what readers actually search for, which is closer to “conflicted relief” or “ambivalent grief.” It was bootstrapped from an LLM analysis of a sample of the archive, then reviewed and refined by people, and it keeps evolving as the archive grows. Claude suggests; an editor approves or sends it back for revision.
Structured embeddings and backend
Each approved record (title, author, summary, moods, themes) is embedded with OpenAI's text-embedding-3-large and stored in Pinecone. At runtime, a FastAPI service embeds the reader's message, finds the nearest records and passes them to the guide for its reply. Classification happens offline; search happens live.
A conversational guide with a defined character
The guide speaks for the archive, not as a corporate assistant or a therapist. It's warm but grounded in its job as a literary guide: it connects readers to poems and explains why a poem might resonate with what they described. Designing the character, not the writing, was the hard part.
Boundaries and two-layer safety
A keyword check catches common phrases around self-harm and violence, and a moderation endpoint classifies messages as a second layer. When either detects real risk, recommendations stop and crisis resources take over. Off-topic requests are handled in character: asked for directions to a taco shop, the guide says it can't help with that and offers poems about food and place instead.
Result
An archive that's discoverable by feeling, and a content pipeline that scales.
What it preserved, and what it made possible
- PreservedHuman judgment determines what every piece is about
- PreservedThe publication's voice: the guide speaks on its behalf
- PreservedUser safety and privacy: distress gets crisis resources, and conversations don't persist
- PossibleContent managers can build collections by searching for emotional shape
- PossibleOlder work stays in rotation and keeps finding new readers
- PossibleThe pattern applies to any archive of subjective content where category vocabulary doesn't fit the search vocabulary
Deliverables
- Reader persona and guide persona2 boards
- Conceptual model1 pp
- Interaction model1 pp
- System architecture1 pp
Full decks and specifications available on request.
Redacted for client confidentiality.
mistycripps@protonmail.com