Python
Install candding from PyPI, run its five model types from Python with NumPy output, and find the Rust name behind each Python call.
Install candding from PyPI, then embed and rerank from Python: the package wraps the Rust crate, keeps its method names, and returns NumPy arrays.
Install
pip install candding| Platform | Wheel tag | Python | Devices |
|---|---|---|---|
| Linux x86-64 | manylinux_2_28_x86_64 | CPython 3.9 and newer | CPU |
| Linux aarch64 | manylinux_2_28_aarch64 | CPython 3.9 and newer | CPU |
| macOS on Apple Silicon | macosx_11_0_arm64 | CPython 3.9 and newer | CPU and Metal |
Each wheel targets CPython's stable ABI, abi3, so one wheel per platform serves every Python version in its row. On any other platform pip builds the source distribution, which needs the Rust toolchain and C compiler that getting started lists; no wheel carries CUDA, so a CUDA build comes from the source distribution:
MATURIN_PEP517_ARGS="--features cuda" pip install candding --no-binary canddingDense vectors
from candding import TextEmbedding
model = TextEmbedding("BAAI/bge-small-en-v1.5")
passages = model.passage_embed(["Embeddings map text to vectors.", "The cat sat on the mat."])
queries = model.query_embed(["what do embeddings do"])
print(passages.shape, passages.dtype)
print(queries @ passages.T)passage_embed and query_embed apply the model's templates and embed applies none, as in Rust. Each row has unit length unless the model was built with normalize=False, so a dot product is a cosine similarity.
Sparse rows
from candding import SparseTextEmbedding
model = SparseTextEmbedding("prithivida/Splade_PP_en_v1")
passages = model.passage_embed(["Embeddings map text to vectors.", "The cat sat on the mat."])
query = model.query_embed(["what do embeddings do"])[0]
print(len(passages[0]), passages[0].indices.dtype, passages[0].values.dtype)
print([query.dot(passage) for passage in passages])BM25
from candding import Bm25TextEmbedding
bm25 = Bm25TextEmbedding()
passages = bm25.passage_embed(["Embeddings map text to vectors.", "The cat sat on the mat."])
query = bm25.query_embed(["what do embeddings do"])[0]
print([query.dot(passage) for passage in passages])Bm25TextEmbedding loads no weights and takes no device, dtype or batch size.
Multi-vector matrices
from candding import MultiVectorTextEmbedding
model = MultiVectorTextEmbedding("answerdotai/answerai-colbert-small-v1")
passages = model.passage_embed(["Embeddings map text to vectors.", "The cat sat on the mat."])
query = model.query_embed(["what do embeddings do"])[0]
print(query.shape, passages[0].shape)
print([model.score(query, passage) for passage in passages])score is the late-interaction score under the checkpoint's own reduction, the score its reference ranks by; multi-vector explains the reductions.
Reranking
from candding import TextCrossEncoder
model = TextCrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
documents = ["The cat sat on the mat.", "Embeddings map text to vectors."]
print(model.rerank("what do embeddings do", documents))rerank returns (document index, score) pairs best first; tied scores keep their input order and a NaN ranks last.
API
Every keyword after the model id is keyword-only.
| Python | Returns |
|---|---|
TextEmbedding(model_id, *, device="auto", dtype=None, max_length=None, batch_size=None, revision=None, normalize=True) | a dense model |
.embed(texts, *, dim=None, batch_size=None), .query_embed(texts, *, instruction=None, dim=None, batch_size=None), .passage_embed(texts, *, dim=None, batch_size=None) | a float32 array of shape (len(texts), dim); query_embed takes instruction or dim, not both, as no Rust call takes both |
.token_count(texts), .query_token_count(texts, *, instruction=None), .passage_token_count(texts) | list[int] |
SparseTextEmbedding(model_id, *, device="auto", dtype=None, max_length=None, batch_size=None, revision=None) | a sparse model; .embed, .query_embed and .passage_embed take texts and batch_size and return list[SparseEmbedding], and .token_count(texts) returns list[int] |
SparseEmbedding | one row: .indices, a uint32 array ascending, .values, a float32 array, .dot(other) and len(row) |
Bm25TextEmbedding(model_id="Qdrant/bm25") | .embed, .query_embed and .passage_embed take texts and return list[SparseEmbedding] |
MultiVectorTextEmbedding(model_id, *, device="auto", dtype=None, max_length=None, batch_size=None, revision=None, normalize=True) | .embed, .query_embed and .passage_embed return one float32 array of shape (tokens, dim) per text; .score(query, document) returns a float; .token_count(texts) |
TextCrossEncoder(model_id, *, device="auto", dtype=None, max_length=None, batch_size=None, revision=None) | .rerank(query, documents, *, instruction=None, batch_size=None) returns list[tuple[int, float]] best first; .pair_token_count(query, documents, *, instruction=None) and .pair_window_count(query, documents) return list[int] |
list_models(kind=None) | list[dict], every catalog entry with its statuses per device and dtype and a kind key: "dense", "sparse", "multi_vector" or "rerank"; kind keeps one catalog |
describe(model_id) | a dict from each catalog that carries the id to its descriptor |
Every model but Bm25TextEmbedding has the properties max_length, device, dtype, weights_dtype and descriptor, and TextEmbedding also has dim; Bm25TextEmbedding has descriptor. device and the dtypes are strings, and descriptor is a dict holding the fields candding describe --json prints. The model catalog lists every id.
Devices and dtypes
device | Backend |
|---|---|
"auto" | Metal, then CUDA, then the CPU, among the backends compiled in |
"cpu" | the CPU, in every wheel |
"metal" | the Apple GPU, in the macOS wheel |
"cuda" | the NVIDIA GPU, in a source build with the cuda feature |
dtype | Weights |
|---|---|
None | the model's default dtype |
"f32", "f16" or "bf16" | that dtype, where the catalog allows it on the device; "bf16" on the CPU raises CanddingError |
A device the build lacks raises CanddingError with the kind BackendNotCompiled rather than falling back to the CPU. Devices and dtypes has each dtype's limits.
Errors
| Exception | When |
|---|---|
candding.CanddingError | any failure the Rust crate reports: kind is the Rust variant's name, such as UnsupportedOnDevice, NonDenseModel or GatedModel, and str(error) is the Rust message |
ValueError | an unknown device, dtype or kind; an id that describe or Bm25TextEmbedding does not carry; query_embed with both instruction and dim; score on arrays of different widths |
TypeError | a single string where a list of texts belongs |
The API page lists the variants.
Threads
Loading a model and every embed, count, score and rerank call release the GIL while the Rust code runs, so Python threads can call one model at once. On the CPU their forward passes run in parallel; on Metal the crate serializes every forward pass in the process behind one lock.
The Rust names
| Rust | Python |
|---|---|
TextEmbedding::builder(id).device(device::cpu()).dtype(DType::F16).build()? | TextEmbedding(id, device="cpu", dtype="f16") |
embed_dim(texts, dim, None), query_embed_dim, passage_embed_dim | embed(texts, dim=dim), query_embed(texts, dim=dim), passage_embed(texts, dim=dim) |
query_embed_with_instruction(texts, instruction, None), query_token_count_with_instruction | query_embed(texts, instruction=instruction), query_token_count(texts, instruction=instruction) |
rerank_with_instruction, pair_token_count_with_instruction | rerank(query, documents, instruction=instruction), pair_token_count(query, documents, instruction=instruction) |
rerank, one RerankScore per document in input order | rerank, (index, score) pairs best first |
Vec<Vec<f32>> from a dense call | one float32 array |
MultiVectorEmbedding | a float32 array of shape (rows, dim) |
multi_vector::late_interaction_score_for(reduction, query, document, fill) | model.score(query, document), with the reduction and fill of the model's descriptor |
Bm25TextEmbedding::builder().build() | Bm25TextEmbedding() |
registry::dense::all() and the other catalogs' all | list_models("dense") and the other kinds |
registry::dense::lookup(id) and the other catalogs' lookup | describe(id) |
CanddingError::UnsupportedOnDevice { .. } | CanddingError whose kind is "UnsupportedOnDevice" |
embed is the passage side on SparseTextEmbedding, Bm25TextEmbedding and MultiVectorTextEmbedding, and applies no template on TextEmbedding, as in Rust. The API page has the Rust types.
Verify
The dense example prints the array's shape and dtype, then one row of scores in which the first passage scores higher; the last digits can differ between machines:
(2, 384) float32
[[0.7862278 0.35534123]]The pytest suite in candding-py/tests checks that every Python output equals the Rust crate's bit for bit on the CPU, and that the dense vectors match the reference implementation's within the F32 tolerance.