Skip to main content

Command Palette

Search for a command to run...

Chroma DB

Updated
•6 min read•View as Markdown
Y

Developer | Adept in software development | Building expertise in machine learning and deep learning

Chroma DB Filtering

Filtering in Chroma DB fundamentally differs from traditional SQL based filtering due to its emphasis on vector similarity and flexible metadata querying.

Chroma DB supports two primary types of filtering:

Filter Type Description Comparison to SQL
Metadata Filtering Filters based on document metadata, such as "topic": "history" or "date": "2023-01-15". Similar to SQL WHERE clauses, but more flexible and can be combined with vector search.
Document Filtering Filters based on document content using keyword presence (e.g., \(contains, \)not_contains). Comparable to SQL's CONTAINS or LIKE operators, but more powerful when integrated with vector search.

Metadata filtering in Chroma DB can be performed by using the where parameter inside the .query(), .get(), or .delete() methods.

Document filtering in Chroma DB can be performed by supplying \(contains or \)not_contains to the where_document parameter inside the .query(), .get(), or .delete() methods using the following syntax:

    where_document={"$contains":"value"}

Moreover, note that you can combine multiple document filters using the \(and and \)or document operators in an analogous way to the metadata filters.

A full example of metadata and document filtering in Chroma DB

Setup

import chromadb
from chromadb.utils import embedding_functions

ef = embedding_functions.SentenceTransformerEmbeddingFunction(
    model_name="all-MiniLM-L6-v2"
)

# Create a Collection
client = chromadb.Client()
collection = client.create_collection(
    name="filter_demo",
    metadata={"description": "Used to demo filtering in ChromaDB"},
    configuration={
        "embedding_function": ef
    }
)
print(f"Collection created: {collection.name}")

# adding documents to Collections
collection.add(
    documents=[
        "This is a document about LangChain",
        "This is a reading about LlamaIndex",
        "This is a book about Python",
        "This is a document about pandas",
        "This is another document about LangChain"
    ],
    metadatas=[
        {"source": "langchain.com", "version": 0.1},
        {"source": "llamaindex.ai", "version": 0.2},
        {"source": "python.org", "version": 0.3},
        {"source": "pandas.pydata.org", "version": 0.4},
        {"source": "langchain.com", "version": 0.5},
    ],
    ids=["id1", "id2", "id3", "id4", "id5"]
)

# filter using metaData
collection.get(
    where={"source": {"$eq": "langchain.com"}}
)

# filter using document content
collection.get(
    where_document={"$contains":"pandas"}
)

Similarity Search and HNSW

In Chorma DB

What is a vector Index?

it is very slow to compare with each embedding of all documents, to address this, applying vector indexes.

A vector index is specifically designed to store and orgnize high-dimensional embeddings for fast similarity and nearest neighbor searches. Instead of treating the dataset as a flat list, the index structures the data in a way that reflects geometry of vector space, clustering similar vectors together or linking them through proximity-based graphs. so searching and focuing on the most promising regions.

What is (H)ierarchical (N)avigable (S)mall (W)orld

Chroma DB uses HNSW as the sole indexing method, it is a graph-based vector index designed for approximate nearest neighbor(ANN) search in high dimensional space.

How It Works

HNSW builds a multi-layered graph where:

  • The upper layers contain a sparse overview of the data for fast navigation.

  • The bottom layer holds all vectors for detailed search.

Each vector connects to a few nearby neighbors, forming a "small world" network—meaning most vectors can be reached in just a few steps.

Why Use HNSW?

  • Fast: Avoids scanning the entire dataset.

  • Accurate: Delivers near-exact results.

  • Scalable: Handles millions to billions of vectors.

  • Versatile: Works with various similarity metrics.

# Setup
import chromadb
from chromadb.utils import embedding_functions
ef = embedding_functions.SentenceTransformerEmbeddingFunction(
    model_name="all-MiniLM-L6-v2"
)

# Collection creation
client = chromadb.Client()
collection = client.create_collection(
    name="my_collection_name",
    metadata={"topic": "query testing"},
    configuration={
        "hnsw": {
            "space": "cosine",
            "ef_search": 100,
            "ef_construction": 100,
            "max_neighbors": 16
        },
        "embedding_function": ef
    }
)

The key configuration parameters are:

  • space: selects the distance metric. Possible options include:

    • l2: squared L2 (Euclidean) distance (default)

    • ip: inner (dot) product distance

    • cosine: cosine distance

  • ef_search: the size of the candidate list used to search for nearest neighbors when a nearest neighbor search is performed. The default value is 100. Higher values improve both accuracy and recall, but at the cost of slower performance and increased computatuonal cost.

  • ef_construction: the size of the candidate list used to select neighbors when a node is inserted during index construction. The default value is 100. Higher values improve the quality of the index and accuracy, but at the cost of slower performance and increased memory usage.

  • max_neighbors: the maximum number of connections each node can have during construction. The defualt value is 16. Higher values lead to denser graphs that perform better during searches at the cost of higher memory usage and construction time.

Peforming Similarity Searches in Chroma DB

# before performing we must add data to cour collection
# documents id1 and id3 refer to pandas as animals, while id2 and id4 refer to the Python library
collection.add(
    documents=[
        "Giant pandas are a bear species that lives in mountainous areas.",
        "A pandas DataFrame stores two-dimensional, tabular data",
        "I think everyone agrees that pandas are some of the cutest animals on the planet",
        "A direct comparison between pandas and polars indicates that polars is a more efficient library than pandas.",
    ],
    metadatas=[
        {"topic": "animals"},
        {"topic": "data analysis"},
        {"topic": "animals"},
        {"topic": "data analysis"},
    ],
    ids=["id1", "id2", "id3", "id4"]
)

# once collection created
# n_results control returned results
# query_text is a list
collection.query(
    query_texts=["cats"],
    n_results=10,
)

# results
{'ids': [['id3', 'id1', 'id2', 'id4']],
 'embeddings': None,
 'documents': [['I think everyone agrees that pandas are some of the cutest animals on the planet',
   'Giant pandas are a bear species that lives in mountainous areas.',
   'A pandas DataFrame stores two-dimensional, tabular data',
   'A direct comparison between pandas and polars indicates that polars is a more efficient library than pandas.']],
 'uris': None,
 'included': ['metadatas', 'documents', 'distances'],
 'data': None,
 'metadatas': [[{'topic': 'animals'},
   {'topic': 'animals'},
   {'topic': 'data analysis'},
   {'topic': 'data analysis'}]],
 'distances': [[0.7380143404006958,
   0.8351750373840332,
   0.8634340167045593,
   0.9299634695053101]]}

Query with Filters

how to improve query results, 1) updating the query giving more context 2) change to antoher embedding model that may better capture the meaning 3) adding filter, for instance filter topic about 'animals'

# add a specific topic filter
collection.query(
    query_texts=["polar bear"],
    n_results=1,
    where={'topic': 'animals'}
)
# result
'ids': [['id1']],
 'embeddings': None,
 'documents': [['Giant pandas are a bear species that lives in mountainous areas.']],
 'uris': None,
 'included': ['metadatas', 'documents', 'distances'],
 'data': None,
 'metadatas': [[{'topic': 'animals'}]],
 'distances': [[0.7096824645996094]]}

# optionally, instead of meat data filtering, performing a full content search to include or exclude documents based on specific words or phrases
collection.query(
    query_texts=["polar bear"],
    n_results=1,
    where_document={'$not_contains': 'library'}
)

# output
'ids': [['id1']],
 'embeddings': None,
 'documents': [['Giant pandas are a bear species that lives in mountainous areas.']],
 'uris': None,
 'included': ['metadatas', 'documents', 'distances'],
 'data': None,
 'metadatas': [[{'topic': 'animals'}]],
 'distances': [[0.7096824645996094]]}

# can combine both meta data and full text earch in Chroma DB
collection.query(
    query_texts=["polar bear"],
    n_results=1,
    where={'topic': 'animals'},
    where_document={'$not_contains': 'library'}
)
#output
'ids': [['id1']],
 'embeddings': None,
 'documents': [['Giant pandas are a bear species that lives in mountainous areas.']],
 'uris': None,
 'included': ['metadatas', 'documents', 'distances'],
 'data': None,
 'metadatas': [[{'topic': 'animals'}]],
 'distances': [[0.7096824645996094]]}