Skip to content

Elasticsearch

The Elasticsearch provider uses the DSL module that ships with the elasticsearch client (elasticsearch.dsl) for document store operations, making it suitable for search and analytics workloads.

Overview

Elasticsearch is a document-oriented provider designed for:

  • Full-text search across domain data
  • Analytics and aggregation workloads
  • Read-optimized views when combined with projections

Unlike relational providers, Elasticsearch does not support real transactions or raw queries. It is best used for read-heavy workloads where eventual consistency is acceptable, or alongside a relational provider for write operations.

Installation

pip install "protean[elasticsearch]"

# Or install the client separately
pip install "elasticsearch>=8.18.0,<9.0.0"

Configuration

[databases.default]
provider = "memory"

[databases.search]
provider = "elasticsearch"
database_uri = { hosts = ["${ELASTICSEARCH_HOST|http://localhost:9200}"] }
NAMESPACE_PREFIX = "${PROTEAN_ENV}"
NAMESPACE_SEPARATOR = "-"
SETTINGS = { number_of_shards = 3 }

Aggregates declared with provider="search" are stored in Elasticsearch. ${ELASTICSEARCH_HOST|...} reads the host from the ELASTICSEARCH_HOST environment variable, and uses the value after | when it is not set.

The provider reads NAMESPACE_PREFIX, NAMESPACE_SEPARATOR and SETTINGS in upper case only. In lower case they are ignored.

Configuration Options

Option Default Description
provider Required Must be "elasticsearch" for Elasticsearch
database_uri Required A table with a hosts list. Each host is a scheme://host:port URL, or host:port (the scheme is then http)
NAMESPACE_PREFIX None Prefix for index names (e.g. prod → prod_person)
NAMESPACE_SEPARATOR "_" Character joining prefix and index name
SETTINGS None Index settings passed as-is to Elasticsearch

Namespace Prefixing

Index names are derived from aggregate class names. When NAMESPACE_PREFIX is set, it is prepended to every index name:

Prefix Separator Aggregate Index Name
prod _ (default) Person prod_person
prod - Person prod-person
(none) — Person person

Using NAMESPACE_PREFIX = "${PROTEAN_ENV}" lets you share a single Elasticsearch cluster across environments by giving each environment a distinct prefix.

Capabilities

The Elasticsearch provider supports the following capabilities:

  • ✅ CRUD: Create, Read, Update, Delete single records
  • ✅ FILTER: Query/filter records with lookup criteria
  • ✅ BULK_OPERATIONS: update_all(), delete_all()
  • ✅ ORDERING: Server-side ordering of results
  • ✅ SCHEMA_MANAGEMENT: Create/drop indices
  • ✅ OPTIMISTIC_LOCKING: Version-based concurrency control
  • ❌ TRANSACTIONS: No transaction support (session has no-op commit/rollback)
  • ❌ SIMULATED_TRANSACTIONS: Not applicable
  • ❌ RAW_QUERIES: Not supported
  • ❌ CONNECTION_POOLING: Managed by the elasticsearch client internally
  • ❌ NATIVE_JSON: Elasticsearch stores JSON natively, but this flag refers to SQL-style JSON columns
  • ❌ NATIVE_ARRAY: No SQL-style array columns

Indexes

Elasticsearch does not use relational indexes, so portable Index declarations are not translated into DDL here. Configure search behavior through the Elasticsearch field mapping (below) instead. Index declarations on an aggregate remain valid (they are honored by SQL providers); they are not applied by this adapter.

Field Mapping

Protean auto-generates an explicit Elasticsearch mapping for every aggregate. Each Protean field type is mapped to an appropriate elasticsearch.dsl field type:

Protean Field ES Mapping Type Notes
String keyword Exact match, sortable, aggregatable
Identifier / Auto keyword Identity fields
Integer integer
Float float
Boolean boolean
DateTime date
Date date
Dict (dynamic) Uses ES dynamic mapping
List (dynamic) Uses ES dynamic mapping
ValueObjectList nested Nested objects

String fields default to keyword. They support exact matching, sorting, and aggregations. If you need full-text search with analyzers, define a custom Elasticsearch Model (see below).

Custom Elasticsearch Model

Supply a custom @domain.database_model when you need ES-specific field tuning (analyzers, multi-fields, normalizers, etc.). User-defined fields take precedence; unmapped attributes are filled in automatically from the aggregate. This example needs a running Elasticsearch server:

# fragment
import os

from elasticsearch import dsl

from protean import Domain
from protean.fields import String

domain = Domain(name="Articles")
domain.config["databases"]["default"] = {
    "provider": "elasticsearch",
    "database_uri": {
        "hosts": [os.environ.get("ELASTICSEARCH_HOST", "http://localhost:9200")]
    },
}


@domain.aggregate(schema_name="articles")
class Article:
    title: String()
    body: String()
    category: String()


@domain.database_model(part_of=Article)
class ArticleModel:
    # Full-text search with .keyword subfield for exact match
    title = dsl.Text(analyzer="standard", fields={"keyword": dsl.Keyword()})
    # Full-text search only
    body = dsl.Text(analyzer="english")
    # category is not listed: auto-mapped as Keyword from the aggregate


domain.init(traverse=False)

The index is named articles, from the aggregate's schema_name option. title and body are mapped as text, and category as keyword.

Note

Index settings come from the SETTINGS configuration option. For a custom model written as a plain class, as above, Protean builds the model's Index inner class itself, so settings declared on that class are not used. A custom model that already subclasses Protean's ElasticsearchModel keeps its own Index class.

Note

When a custom model defines Text-type fields, lookups like exact, in, contains, startswith, and endswith on those fields automatically use the .keyword subfield for exact matching.

Limitations

  • No Real Transactions: Elasticsearch does not support ACID transactions. The session object provides no-op commit() and rollback() methods. Data is indexed immediately on write.
  • Eventual Consistency: Newly indexed documents may not be immediately searchable. Elasticsearch refreshes indices periodically (default: 1 second).
  • No Raw Queries: The raw() method is not supported. Use the Elasticsearch DSL directly if you need advanced query features.
  • Requires Running Service: Elasticsearch must be installed and running. Use make up to start Protean's Docker-based development services.