Rails Hybrid Search: Combining pgvector Semantic Search with Postgres Full-Text for Better Recall
Rails hybrid search combines pgvector semantic vectors with Postgres full-text search and Reciprocal Rank Fusion for dramatically better search recall.
Six months after we shipped semantic search on a client’s knowledge base, their support team came back with a strange complaint. Users searching for “404 error” were getting results about HTTP status codes, REST APIs, and web infrastructure — technically related but completely useless when the user needed the article titled “What to do when your integration returns 404.” Pure keyword search would have surfaced that article on the first result. Our vectors were doing something else entirely.
That is the failure mode nobody warns you about when you swap your search for embeddings. Semantic search is brilliant at finding conceptually related content when users express intent in natural language. It falls apart on exact-match queries: product codes, error codes, proper nouns, technical acronyms, anything where the actual characters matter more than meaning. Keyword search has the inverse problem. “Machine learning model deployment” finds documents with those exact words but misses the article that discusses “deploying your ML pipeline to production” without using the phrase “model deployment.”
Rails hybrid search solves both. You run semantic search and full-text search in parallel, then combine the rankings with a small algorithm called Reciprocal Rank Fusion. The result consistently outperforms either approach alone. After implementing this in production several times, I have stopped shipping pure semantic search entirely. Hybrid is the default.
Why Semantic-Only Search Fails in Production
Semantic search works by converting queries and documents into embedding vectors — arrays of floating-point numbers that encode meaning. Documents with similar meaning have vectors that are close together in vector space. When a user searches, you embed the query and find the nearest document vectors.
The weakness is precision on exact terms. Your embedding model has never seen “ERR_INVALID_RESPONSE” in meaningful training context, so the embedding for that string is arbitrary. A user pasting an error code gets results that are semantically adjacent — articles about browser errors, network problems, debugging — but not the specific article with that exact error code in the title. Same problem with product SKUs, version numbers, names of people, and company-specific jargon your model has never encountered.
The other failure mode is document length. Embedding models typically have a context window of 512 to 8192 tokens. Long documents get truncated or chunked. A search phrase appearing only at the end of a long document might not be retrieved because the chunk containing that phrase ranked poorly or got cut entirely.
Why Full-Text-Only Search Also Fails
Full-text search tokenizes documents into words, builds an inverted index, and scores by term frequency and inverse document frequency (TF-IDF). Postgres does this natively with tsvector and tsquery. The pg_search gem wraps it cleanly for Rails.
The weakness is vocabulary mismatch. “How do I make my app faster?” does not match an article titled “Optimizing Rails performance for high-traffic applications” unless those exact words appear in it. Full-text search has no concept of synonyms, related terms, or paraphrases.
It also struggles with natural language questions. Users increasingly type queries the way they talk — especially in AI-adjacent apps where a chat metaphor sets expectations. Keyword search returns nothing meaningful for “why is my background job not processing” unless a document contains exactly those words.
What Hybrid Search Is: The Architecture
Hybrid search runs both queries and merges the ranked lists. The algorithm for merging is called Reciprocal Rank Fusion (RRF). It is deceptively simple:
score(document) = Σ 1 / (k + rank_i(document))
Where rank_i(document) is the document’s rank in the i-th result list (1-based), and k is a smoothing constant — typically 60. Documents ranked highly in both lists get the highest combined score. Documents absent from one list still contribute their score from the other. No per-corpus weight tuning required, no score normalization needed.
The full implementation in Rails runs two queries — one vector similarity query via pgvector, one full-text query using tsvector — and merges them in a single SQL CTE.
Setting Up pgvector for the Semantic Arm
If you do not have pgvector running, the pgvector and embeddings post covers the setup end to end. Short version: add the pgvector gem, enable the extension, add a vector column, generate embeddings with OpenAI or another model, and store them.
# Migration
class AddEmbeddingToDocuments < ActiveRecord::Migration[8.0]
def change
add_column :documents, :embedding, :vector, limit: 1536
add_index :documents, :embedding,
using: :hnsw,
opclass: :vector_cosine_ops,
name: "index_documents_embedding_hnsw"
end
end
limit: 1536 matches OpenAI’s text-embedding-3-small dimensions. The HNSW index makes approximate nearest neighbor search fast enough for production — on a corpus of a million documents, under 20ms on a warm Postgres cache.
For generating embeddings on save:
class Document < ApplicationRecord
after_save :schedule_embedding, if: :body_previously_changed?
private
def schedule_embedding
GenerateDocumentEmbeddingJob.perform_later(id)
end
end
class GenerateDocumentEmbeddingJob < ApplicationJob
def perform(document_id)
document = Document.find(document_id)
client = OpenAI::Client.new
response = client.embeddings(
parameters: {
model: "text-embedding-3-small",
input: document.body.truncate(8000)
}
)
embedding = response.dig("data", 0, "embedding")
document.update_column(:embedding, embedding)
end
end
Setting Up tsvector for the Keyword Arm
For the keyword arm, raw tsvector columns integrate more cleanly into the hybrid CTE than pg_search does. The pg_search post covers the gem approach if you prefer it.
class AddSearchVectorToDocuments < ActiveRecord::Migration[8.0]
def up
add_column :documents, :search_vector, :tsvector
add_index :documents, :search_vector, using: :gin
execute <<~SQL
UPDATE documents
SET search_vector =
setweight(to_tsvector('english', coalesce(title, '')), 'A') ||
setweight(to_tsvector('english', coalesce(body, '')), 'B')
SQL
execute <<~SQL
CREATE FUNCTION documents_search_vector_update() RETURNS trigger AS $$
BEGIN
NEW.search_vector :=
setweight(to_tsvector('english', coalesce(NEW.title, '')), 'A') ||
setweight(to_tsvector('english', coalesce(NEW.body, '')), 'B');
RETURN NEW;
END
$$ LANGUAGE plpgsql;
CREATE TRIGGER documents_search_vector_update
BEFORE INSERT OR UPDATE ON documents
FOR EACH ROW EXECUTE FUNCTION documents_search_vector_update();
SQL
end
def down
execute "DROP TRIGGER IF EXISTS documents_search_vector_update ON documents"
execute "DROP FUNCTION IF EXISTS documents_search_vector_update"
remove_column :documents, :search_vector
end
end
setweight with 'A' for title and 'B' for body means title matches rank higher than body-only matches. The GIN index makes the @@ operator fast. Both are non-negotiable for production.
The Hybrid Query: Reciprocal Rank Fusion in SQL
Here is the full service object. It runs both queries in a single SQL CTE and applies RRF in the database, which avoids loading two full result sets into Ruby just to merge them:
class Document::HybridSearch
K = 60
DEFAULT_LIMIT = 20
def initialize(query:, limit: DEFAULT_LIMIT)
@query = query
@limit = limit
@embedding = embed(query)
end
def call
Document.find_by_sql([<<~SQL, embedding: @embedding, tsquery: tsquery, limit: @limit, k: K])
WITH semantic AS (
SELECT
id,
ROW_NUMBER() OVER (ORDER BY embedding <=> :embedding::vector) AS rank
FROM documents
WHERE embedding IS NOT NULL
ORDER BY embedding <=> :embedding::vector
LIMIT 60
),
keyword AS (
SELECT
id,
ROW_NUMBER() OVER (ORDER BY ts_rank_cd(search_vector, query) DESC) AS rank
FROM documents,
to_tsquery('english', :tsquery) query
WHERE search_vector @@ query
ORDER BY ts_rank_cd(search_vector, query) DESC
LIMIT 60
),
rrf AS (
SELECT
COALESCE(s.id, k.id) AS id,
COALESCE(1.0 / (:k + s.rank), 0.0)
+ COALESCE(1.0 / (:k + k.rank), 0.0) AS score
FROM semantic s
FULL OUTER JOIN keyword k ON s.id = k.id
)
SELECT documents.*
FROM documents
JOIN rrf ON documents.id = rrf.id
ORDER BY rrf.score DESC
LIMIT :limit
SQL
end
private
def embed(text)
client = OpenAI::Client.new
response = client.embeddings(
parameters: { model: "text-embedding-3-small", input: text.truncate(8000) }
)
response.dig("data", 0, "embedding")
end
def tsquery
@query
.gsub(/[^a-z0-9\s]/i, " ")
.split
.reject(&:empty?)
.map { |term| "#{term}:*" }
.join(" & ")
end
end
Usage from a controller:
class SearchController < ApplicationController
def show
@results = Document::HybridSearch.new(query: params[:q], limit: 10).call
end
end
A few details worth unpacking:
The LIMIT 60 on each subquery is the candidate pool per arm. Increasing it improves recall at the cost of query time. Sixty candidates is enough for most apps. On a million-document corpus with proper HNSW and GIN indexes, this query runs in 50–100ms cold. Warm, under 20ms.
The tsquery method converts free-form input into a Postgres tsquery. The :* suffix enables prefix matching — “deploy” matches “deployment” and “deploying”. The & between terms requires all terms to match. Change to | for OR semantics if your queries tend to be multi-word with optional terms, but AND gives better precision for most knowledge-base use cases.
The FULL OUTER JOIN in the RRF CTE ensures documents appearing in only one arm still get a score from that arm. A document ranked first in semantic search that does not appear in the keyword results at all still scores 1 / (60 + 1) ≈ 0.016. One that ranks top in both scores 1/61 + 1/61 ≈ 0.033. The combined ordering naturally promotes documents that are both semantically close and keyword-matched.
Hybrid Search Rails: Tuning the Weights
Standard RRF treats both arms equally. In practice you may want to weight them differently — give semantic search more influence on a general knowledge base, give full-text more influence on a technical documentation site where error codes and exact-match queries dominate.
The cleanest approach is a per-arm multiplier, not changing the RRF formula:
SEMANTIC_WEIGHT = 1.0
KEYWORD_WEIGHT = 1.5
# In the RRF CTE, replace the score calculation with:
# COALESCE(:semantic_weight / (:k + s.rank), 0.0)
# + COALESCE(:keyword_weight / (:k + k.rank), 0.0) AS score
Expose these as configuration on the service object and log them alongside query results. After a week of production traffic, look at which arm contributed the relevant result on your worst-performing queries and adjust from there. I rarely tune in practice — equal weights work for most corpora — but having the lever available has saved me more than once.
Performance and Indexing Considerations
Both arms need their own indexes or this degrades to full table scans at scale.
For the vector arm, the HNSW index handles approximate nearest neighbor search. The defaults work for up to a few hundred thousand documents. At millions of rows, tune ef_construction upward (longer index build, better recall) and benchmark with SET hnsw.ef_search = 100 in a session to see the recall/latency tradeoff.
For the keyword arm, the GIN index on search_vector is required. Ensure your Postgres work_mem is high enough for GIN builds on large corpora — 2–4GB per connection is reasonable. Without enough working memory, Postgres degrades to a lossy index and query performance degrades.
One thing that will silently kill query performance: running the vector arm without WHERE embedding IS NOT NULL. Documents waiting for background embedding generation have a NULL embedding. Postgres has to scan those rows, attempt the comparison, return NULL, and filter them out. Always add the filter.
Profile with EXPLAIN (ANALYZE, BUFFERS) before and after any index change. On Postgres 16+ the optimizer is good, but hybrid queries have enough moving parts that the plan can surprise you. The CTE scans especially — FULL OUTER JOIN on two subqueries sometimes gets a bad hash join estimate. If you see that, materialized CTEs (WITH ... AS MATERIALIZED) force the subqueries to run first:
WITH semantic AS MATERIALIZED (...),
keyword AS MATERIALIZED (...),
rrf AS (...)
SELECT ...
Testing Hybrid Search in Rails
The hard part about testing search quality is that “better” is subjective. The practical approach I use is a golden set: fixed queries paired with expected top results, checked in CI as a regression suite.
RSpec.describe Document::HybridSearch do
fixtures :documents
GOLDEN_QUERIES = [
{ query: "404 error not found", expected_id: :http_404_article },
{ query: "how to deploy rails app", expected_id: :kamal_deployment_guide },
{ query: "ERR_INVALID_RESPONSE", expected_id: :chrome_errors_reference },
].freeze
it "returns expected top result for golden queries" do
GOLDEN_QUERIES.each do |pair|
results = described_class.new(query: pair[:query], limit: 5).call
expect(results.first).to eq(documents(pair[:expected_id])),
"Query '#{pair[:query]}' expected #{documents(pair[:expected_id]).title}, got #{results.first&.title}"
end
end
end
Store real embeddings in the fixtures. Generate them once and commit the YAML to the repo. Tests run without hitting any API, and regressions surface when someone changes the schema or the search logic. The RAG with pgvector post covers the embedding fixture strategy in detail.
For benchmarking recall improvement, measure Mean Reciprocal Rank (MRR) on the golden set:
def mean_reciprocal_rank(query_pairs, k: 10)
scores = query_pairs.map do |query, relevant_id|
results = Document::HybridSearch.new(query: query, limit: k).call
rank = results.index { |r| r.id == relevant_id }
rank ? 1.0 / (rank + 1) : 0.0
end
scores.sum / scores.size
end
On every corpus I have measured, hybrid beats semantic-only by 8–25% on MRR and beats keyword-only by 15–40%. The exact numbers depend entirely on your query distribution. Apps with a lot of technical exact-match queries benefit more from the keyword arm. General knowledge bases benefit more from the semantic arm. Hybrid beats both in every case I have seen.
When to Use Hybrid, Semantic Only, or Keyword Only
Use semantic only when queries are almost always natural language, the corpus is homogeneous (no proper nouns or codes), and recall precision matters less than covering vague queries well. Marketing copy discovery. Blog post recommendations. Conceptual similarity search.
Use keyword only when the corpus has many exact-match terms — legal references, part numbers, code identifiers — users are technical and type precise queries, and latency is a hard constraint. Developer portal code search. Legal document lookup by statute number.
Use hybrid for everything else in production. Support knowledge bases. Documentation sites. E-commerce product search combined with attribute filtering. Internal wikis. Anything where your users have mixed query types — some typing precise technical terms, others asking conversational questions. The latency overhead over pure semantic or pure keyword is one additional SQL subquery. The recall improvement is consistent.
After nineteen years of building Rails apps and the last few years adding AI-powered search to many of them, hybrid is the architecture I default to. I stopped asking “do we need hybrid?” and started asking “is there a specific reason to skip one of the arms?”
FAQ
How does Rails hybrid search compare to using Elasticsearch?
Elasticsearch’s multi_match does keyword search well. Its knn query does dense vector search. Their hybrid search support as of 8.x is less cleanly integrated than Postgres and requires a separate infrastructure dependency. For Rails apps already on Postgres with pgvector, the approach in this post avoids that dependency with comparable quality. Elasticsearch makes sense when you need it for other reasons — massive scale beyond what Postgres handles, complex aggregations, cross-index federation. Do not add it just for hybrid search if Postgres already serves you.
What embedding model should I use for Rails hybrid search?
For English-language content, OpenAI’s text-embedding-3-small at 1536 dimensions gives excellent results at low cost. For multilingual content, text-embedding-3-large or a multilingual sentence-transformers model. For on-premise or latency-constrained setups, a locally-hosted model via Ollama works with the same code — swap the embedding call. The specific model matters less than consistency: always embed with the same model at index time and at query time, or your similarity scores become meaningless.
Does Reciprocal Rank Fusion work better than linear score combination?
RRF is more robust than linear combination. Linear combination requires normalizing scores from both arms onto the same scale — cosine similarity from pgvector sits in [-1, 1], ts_rank_cd from Postgres full-text sits in [0, 1] with different distributions at that. Normalization is fragile when the corpus distribution changes. RRF uses rank positions, which are scale-independent and require no calibration. The practical quality difference is small but RRF requires zero tuning. I use it by default and switch to weighted linear combination only with a specific reason, such as wanting to heavily boost recency or document freshness.
How do I handle documents without embeddings in the hybrid query?
The CTE handles this automatically. Documents without embeddings are excluded from the semantic arm by WHERE embedding IS NOT NULL but still appear in keyword results if they match. Freshly saved documents that haven’t yet been processed by the embedding job participate in keyword search immediately. Once the background job generates the embedding, they join the semantic arm on the next query. No special handling required — the FULL OUTER JOIN in the RRF CTE covers the asymmetry.
Building search into your Rails app and want it right the first time — hybrid retrieval, proper pgvector indexing, and benchmarked recall? TTB Software builds production-grade search and AI integrations on Rails. We have been doing this for nineteen years.
Related Articles
Rails LLM Conversation History: Persisting Context, Summarization, and Managing the Context Window
LLM conversation history in Rails: persist message turns, manage context windows, implement summarization, and build ...
Rails API Serialization: Blueprinter, Alba, and JSONAPI-Serializer Compared for Production APIs
Rails API serialization done right: compare Blueprinter, Alba, and jsonapi-serializer with real code, N+1 traps, cach...
Rails Data Migrations: Safe Backfills with data-migrate, Maintenance Tasks, and Batched Updates
Rails data migrations done right: use data-migrate or maintenance_tasks for safe, resumable backfills that don't lock...