ScreenJSON can carry its own embeddings under analysis.embeddings, keyed by
the UUID of the scene, element or character they describe
(how to generate them). The free
screenjson-db-importer can put those
vectors into a vector database’s native vector field, so they’re searchable
straight away, with the scene’s text and metadata alongside.
1. Declare the vector
Use a layout map, and give the level you embedded a vector with the model
name and dimensions you used:
data_dir: ./.screenjson-importer
storage:
driver: postgres
url: postgres://screenjson:[email protected]:5432/screenplays?sslmode=disable
layout:
root: {collection: screenplays}
scenes: {collection: scenes, vector: {model: text-embedding-3-small, dimensions: 1536}}
scenes, elements and characters can each declare a vector.
2. Import
screenjson-db-importer --config importer.yaml ./embedded/
For each scene, the importer takes the first embedding made with
text-embedding-3-small whose length is 1536, and stores it in the scene
record’s vector column. The rest of the embedding (its ID, source, date)
stays in the record, so the original document can be rebuilt exactly.
3. Search
On PostgreSQL, the scenes table now has a pgvector column. Embed a query
with the same model and look for the nearest scenes:
SELECT id, node->'heading' AS heading
FROM scenes
ORDER BY vec <=> $1
LIMIT 10;
Each hit is a scene ID you can cite, with the scene’s ScreenJSON in node
and its document in doc.
Other vector databases
| Database | Change |
|---|---|
| Chroma | driver: chroma, url: http://127.0.0.1:8000 |
| Weaviate | driver: weaviate, url: http://127.0.0.1:8080 |
| Pinecone | driver: pinecone, url: your index host, pinecone.api_key. The index’s dimension must match the declared one. |
Scenes without a matching embedding are still imported. Vector databases store a zero vector for them, marked as not native, so nothing is lost.
Next
- AI playbook: retrieval, RAG and citations.
- screenjson-server: vector search over the same records through its API and MCP tools.