RUBY ON RAILS · 19 MIN READ ·

Rails Hybrid Search: Combining pgvector Semantic Search with Postgres Full-Text for Better Recall

Rails hybrid search combines pgvector semantic vectors with Postgres full-text search and Reciprocal Rank Fusion for dramatically better search recall.

Six months after we shipped semantic search on a client’s knowledge base, their support team came back with a strange complaint. Users searching for “404 error” were getting results about HTTP status codes, REST APIs, and web infrastructure — technically related but completely useless when the user needed the article titled “What to do when your integration returns 404.” Pure keyword search would have surfaced that article on the first result. Our vectors were doing something else entirely.

That is the failure mode nobody warns you about when you swap your search for embeddings. Semantic search is brilliant at finding conceptually related content when users express intent in natural language. It falls apart on exact-match queries: product codes, error codes, proper nouns, technical acronyms, anything where the actual characters matter more than meaning. Keyword search has the inverse problem. “Machine learning model deployment” finds documents with those exact words but misses the article that discusses “deploying your ML pipeline to production” without using the phrase “model deployment.”

Rails hybrid search solves both. You run semantic search and full-text search in parallel, then combine the rankings with a small algorithm called Reciprocal Rank Fusion. The result consistently outperforms either approach alone. After implementing this in production several times, I have stopped shipping pure semantic search entirely. Hybrid is the default.

Why Semantic-Only Search Fails in Production

Semantic search works by converting queries and documents into embedding vectors — arrays of floating-point numbers that encode meaning. Documents with similar meaning have vectors that are close together in vector space. When a user searches, you embed the query and find the nearest document vectors.

The weakness is precision on exact terms. Your embedding model has never seen “ERR_INVALID_RESPONSE” in meaningful training context, so the embedding for that string is arbitrary. A user pasting an error code gets results that are semantically adjacent — articles about browser errors, network problems, debugging — but not the specific article with that exact error code in the title. Same problem with product SKUs, version numbers, names of people, and company-specific jargon your model has never encountered.

The other failure mode is document length. Embedding models typically have a context window of 512 to 8192 tokens. Long documents get truncated or chunked. A search phrase appearing only at the end of a long document might not be retrieved because the chunk containing that phrase ranked poorly or got cut entirely.

Why Full-Text-Only Search Also Fails

Full-text search tokenizes documents into words, builds an inverted index, and scores by term frequency and inverse document frequency (TF-IDF). Postgres does this natively with tsvector and tsquery. The pg_search gem wraps it cleanly for Rails.

The weakness is vocabulary mismatch. “How do I make my app faster?” does not match an article titled “Optimizing Rails performance for high-traffic applications” unless those exact words appear in it. Full-text search has no concept of synonyms, related terms, or paraphrases.

It also struggles with natural language questions. Users increasingly type queries the way they talk — especially in AI-adjacent apps where a chat metaphor sets expectations. Keyword search returns nothing meaningful for “why is my background job not processing” unless a document contains exactly those words.

What Hybrid Search Is: The Architecture

Hybrid search runs both queries and merges the ranked lists. The algorithm for merging is called Reciprocal Rank Fusion (RRF). It is deceptively simple:

score(document) = Σ 1 / (k + rank_i(document))

Where rank_i(document) is the document’s rank in the i-th result list (1-based), and k is a smoothing constant — typically 60. Documents ranked highly in both lists get the highest combined score. Documents absent from one list still contribute their score from the other. No per-corpus weight tuning required, no score normalization needed.

The full implementation in Rails runs two queries — one vector similarity query via pgvector, one full-text query using tsvector — and merges them in a single SQL CTE.

Setting Up pgvector for the Semantic Arm

If you do not have pgvector running, the pgvector and embeddings post covers the setup end to end. Short version: add the pgvector gem, enable the extension, add a vector column, generate embeddings with OpenAI or another model, and store them.

# Migration
class AddEmbeddingToDocuments < ActiveRecord::Migration[8.0]
  def change
    add_column :documents, :embedding, :vector, limit: 1536
    add_index :documents, :embedding,
              using: :hnsw,
              opclass: :vector_cosine_ops,
              name: "index_documents_embedding_hnsw"
  end
end

limit: 1536 matches OpenAI’s text-embedding-3-small dimensions. The HNSW index makes approximate nearest neighbor search fast enough for production — on a corpus of a million documents, under 20ms on a warm Postgres cache.

For generating embeddings on save:

class Document < ApplicationRecord
  after_save :schedule_embedding, if: :body_previously_changed?

  private

  def schedule_embedding
    GenerateDocumentEmbeddingJob.perform_later(id)
  end
end

class GenerateDocumentEmbeddingJob < ApplicationJob
  def perform(document_id)
    document = Document.find(document_id)
    client = OpenAI::Client.new
    response = client.embeddings(
      parameters: {
        model: "text-embedding-3-small",
        input: document.body.truncate(8000)
      }
    )
    embedding = response.dig("data", 0, "embedding")
    document.update_column(:embedding, embedding)
  end
end

Setting Up tsvector for the Keyword Arm

For the keyword arm, raw tsvector columns integrate more cleanly into the hybrid CTE than pg_search does. The pg_search post covers the gem approach if you prefer it.

class AddSearchVectorToDocuments < ActiveRecord::Migration[8.0]
  def up
    add_column :documents, :search_vector, :tsvector
    add_index :documents, :search_vector, using: :gin

    execute <<~SQL
      UPDATE documents
      SET search_vector =
        setweight(to_tsvector('english', coalesce(title, '')), 'A') ||
        setweight(to_tsvector('english', coalesce(body, '')), 'B')
    SQL

    execute <<~SQL
      CREATE FUNCTION documents_search_vector_update() RETURNS trigger AS $$
      BEGIN
        NEW.search_vector :=
          setweight(to_tsvector('english', coalesce(NEW.title, '')), 'A') ||
          setweight(to_tsvector('english', coalesce(NEW.body, '')), 'B');
        RETURN NEW;
      END
      $$ LANGUAGE plpgsql;

      CREATE TRIGGER documents_search_vector_update
      BEFORE INSERT OR UPDATE ON documents
      FOR EACH ROW EXECUTE FUNCTION documents_search_vector_update();
    SQL
  end

  def down
    execute "DROP TRIGGER IF EXISTS documents_search_vector_update ON documents"
    execute "DROP FUNCTION IF EXISTS documents_search_vector_update"
    remove_column :documents, :search_vector
  end
end

setweight with 'A' for title and 'B' for body means title matches rank higher than body-only matches. The GIN index makes the @@ operator fast. Both are non-negotiable for production.

The Hybrid Query: Reciprocal Rank Fusion in SQL

Here is the full service object. It runs both queries in a single SQL CTE and applies RRF in the database, which avoids loading two full result sets into Ruby just to merge them:

class Document::HybridSearch
  K = 60
  DEFAULT_LIMIT = 20

  def initialize(query:, limit: DEFAULT_LIMIT)
    @query = query
    @limit = limit
    @embedding = embed(query)
  end

  def call
    Document.find_by_sql([<<~SQL, embedding: @embedding, tsquery: tsquery, limit: @limit, k: K])
      WITH semantic AS (
        SELECT
          id,
          ROW_NUMBER() OVER (ORDER BY embedding <=> :embedding::vector) AS rank
        FROM documents
        WHERE embedding IS NOT NULL
        ORDER BY embedding <=> :embedding::vector
        LIMIT 60
      ),
      keyword AS (
        SELECT
          id,
          ROW_NUMBER() OVER (ORDER BY ts_rank_cd(search_vector, query) DESC) AS rank
        FROM documents,
             to_tsquery('english', :tsquery) query
        WHERE search_vector @@ query
        ORDER BY ts_rank_cd(search_vector, query) DESC
        LIMIT 60
      ),
      rrf AS (
        SELECT
          COALESCE(s.id, k.id)                            AS id,
          COALESCE(1.0 / (:k + s.rank), 0.0)
            + COALESCE(1.0 / (:k + k.rank), 0.0)         AS score
        FROM semantic s
        FULL OUTER JOIN keyword k ON s.id = k.id
      )
      SELECT documents.*
      FROM documents
      JOIN rrf ON documents.id = rrf.id
      ORDER BY rrf.score DESC
      LIMIT :limit
    SQL
  end

  private

  def embed(text)
    client = OpenAI::Client.new
    response = client.embeddings(
      parameters: { model: "text-embedding-3-small", input: text.truncate(8000) }
    )
    response.dig("data", 0, "embedding")
  end

  def tsquery
    @query
      .gsub(/[^a-z0-9\s]/i, " ")
      .split
      .reject(&:empty?)
      .map { |term| "#{term}:*" }
      .join(" & ")
  end
end

Usage from a controller:

class SearchController < ApplicationController
  def show
    @results = Document::HybridSearch.new(query: params[:q], limit: 10).call
  end
end

A few details worth unpacking:

The LIMIT 60 on each subquery is the candidate pool per arm. Increasing it improves recall at the cost of query time. Sixty candidates is enough for most apps. On a million-document corpus with proper HNSW and GIN indexes, this query runs in 50–100ms cold. Warm, under 20ms.

The tsquery method converts free-form input into a Postgres tsquery. The :* suffix enables prefix matching — “deploy” matches “deployment” and “deploying”. The & between terms requires all terms to match. Change to | for OR semantics if your queries tend to be multi-word with optional terms, but AND gives better precision for most knowledge-base use cases.

The FULL OUTER JOIN in the RRF CTE ensures documents appearing in only one arm still get a score from that arm. A document ranked first in semantic search that does not appear in the keyword results at all still scores 1 / (60 + 1) ≈ 0.016. One that ranks top in both scores 1/61 + 1/61 ≈ 0.033. The combined ordering naturally promotes documents that are both semantically close and keyword-matched.

Hybrid Search Rails: Tuning the Weights

Standard RRF treats both arms equally. In practice you may want to weight them differently — give semantic search more influence on a general knowledge base, give full-text more influence on a technical documentation site where error codes and exact-match queries dominate.

The cleanest approach is a per-arm multiplier, not changing the RRF formula:

SEMANTIC_WEIGHT = 1.0
KEYWORD_WEIGHT  = 1.5

# In the RRF CTE, replace the score calculation with:
# COALESCE(:semantic_weight / (:k + s.rank), 0.0)
#   + COALESCE(:keyword_weight / (:k + k.rank), 0.0)  AS score

Expose these as configuration on the service object and log them alongside query results. After a week of production traffic, look at which arm contributed the relevant result on your worst-performing queries and adjust from there. I rarely tune in practice — equal weights work for most corpora — but having the lever available has saved me more than once.

Performance and Indexing Considerations

Both arms need their own indexes or this degrades to full table scans at scale.

For the vector arm, the HNSW index handles approximate nearest neighbor search. The defaults work for up to a few hundred thousand documents. At millions of rows, tune ef_construction upward (longer index build, better recall) and benchmark with SET hnsw.ef_search = 100 in a session to see the recall/latency tradeoff.

For the keyword arm, the GIN index on search_vector is required. Ensure your Postgres work_mem is high enough for GIN builds on large corpora — 2–4GB per connection is reasonable. Without enough working memory, Postgres degrades to a lossy index and query performance degrades.

One thing that will silently kill query performance: running the vector arm without WHERE embedding IS NOT NULL. Documents waiting for background embedding generation have a NULL embedding. Postgres has to scan those rows, attempt the comparison, return NULL, and filter them out. Always add the filter.

Profile with EXPLAIN (ANALYZE, BUFFERS) before and after any index change. On Postgres 16+ the optimizer is good, but hybrid queries have enough moving parts that the plan can surprise you. The CTE scans especially — FULL OUTER JOIN on two subqueries sometimes gets a bad hash join estimate. If you see that, materialized CTEs (WITH ... AS MATERIALIZED) force the subqueries to run first:

WITH semantic AS MATERIALIZED (...),
     keyword  AS MATERIALIZED (...),
     rrf      AS (...)
SELECT ...

Testing Hybrid Search in Rails

The hard part about testing search quality is that “better” is subjective. The practical approach I use is a golden set: fixed queries paired with expected top results, checked in CI as a regression suite.

RSpec.describe Document::HybridSearch do
  fixtures :documents

  GOLDEN_QUERIES = [
    { query: "404 error not found",       expected_id: :http_404_article },
    { query: "how to deploy rails app",   expected_id: :kamal_deployment_guide },
    { query: "ERR_INVALID_RESPONSE",      expected_id: :chrome_errors_reference },
  ].freeze

  it "returns expected top result for golden queries" do
    GOLDEN_QUERIES.each do |pair|
      results = described_class.new(query: pair[:query], limit: 5).call
      expect(results.first).to eq(documents(pair[:expected_id])),
        "Query '#{pair[:query]}' expected #{documents(pair[:expected_id]).title}, got #{results.first&.title}"
    end
  end
end

Store real embeddings in the fixtures. Generate them once and commit the YAML to the repo. Tests run without hitting any API, and regressions surface when someone changes the schema or the search logic. The RAG with pgvector post covers the embedding fixture strategy in detail.

For benchmarking recall improvement, measure Mean Reciprocal Rank (MRR) on the golden set:

def mean_reciprocal_rank(query_pairs, k: 10)
  scores = query_pairs.map do |query, relevant_id|
    results = Document::HybridSearch.new(query: query, limit: k).call
    rank = results.index { |r| r.id == relevant_id }
    rank ? 1.0 / (rank + 1) : 0.0
  end
  scores.sum / scores.size
end

On every corpus I have measured, hybrid beats semantic-only by 8–25% on MRR and beats keyword-only by 15–40%. The exact numbers depend entirely on your query distribution. Apps with a lot of technical exact-match queries benefit more from the keyword arm. General knowledge bases benefit more from the semantic arm. Hybrid beats both in every case I have seen.

When to Use Hybrid, Semantic Only, or Keyword Only

Use semantic only when queries are almost always natural language, the corpus is homogeneous (no proper nouns or codes), and recall precision matters less than covering vague queries well. Marketing copy discovery. Blog post recommendations. Conceptual similarity search.

Use keyword only when the corpus has many exact-match terms — legal references, part numbers, code identifiers — users are technical and type precise queries, and latency is a hard constraint. Developer portal code search. Legal document lookup by statute number.

Use hybrid for everything else in production. Support knowledge bases. Documentation sites. E-commerce product search combined with attribute filtering. Internal wikis. Anything where your users have mixed query types — some typing precise technical terms, others asking conversational questions. The latency overhead over pure semantic or pure keyword is one additional SQL subquery. The recall improvement is consistent.

After nineteen years of building Rails apps and the last few years adding AI-powered search to many of them, hybrid is the architecture I default to. I stopped asking “do we need hybrid?” and started asking “is there a specific reason to skip one of the arms?”

FAQ

How does Rails hybrid search compare to using Elasticsearch?

Elasticsearch’s multi_match does keyword search well. Its knn query does dense vector search. Their hybrid search support as of 8.x is less cleanly integrated than Postgres and requires a separate infrastructure dependency. For Rails apps already on Postgres with pgvector, the approach in this post avoids that dependency with comparable quality. Elasticsearch makes sense when you need it for other reasons — massive scale beyond what Postgres handles, complex aggregations, cross-index federation. Do not add it just for hybrid search if Postgres already serves you.

For English-language content, OpenAI’s text-embedding-3-small at 1536 dimensions gives excellent results at low cost. For multilingual content, text-embedding-3-large or a multilingual sentence-transformers model. For on-premise or latency-constrained setups, a locally-hosted model via Ollama works with the same code — swap the embedding call. The specific model matters less than consistency: always embed with the same model at index time and at query time, or your similarity scores become meaningless.

Does Reciprocal Rank Fusion work better than linear score combination?

RRF is more robust than linear combination. Linear combination requires normalizing scores from both arms onto the same scale — cosine similarity from pgvector sits in [-1, 1], ts_rank_cd from Postgres full-text sits in [0, 1] with different distributions at that. Normalization is fragile when the corpus distribution changes. RRF uses rank positions, which are scale-independent and require no calibration. The practical quality difference is small but RRF requires zero tuning. I use it by default and switch to weighted linear combination only with a specific reason, such as wanting to heavily boost recency or document freshness.

How do I handle documents without embeddings in the hybrid query?

The CTE handles this automatically. Documents without embeddings are excluded from the semantic arm by WHERE embedding IS NOT NULL but still appear in keyword results if they match. Freshly saved documents that haven’t yet been processed by the embedding job participate in keyword search immediately. Once the background job generates the embedding, they join the semantic arm on the next query. No special handling required — the FULL OUTER JOIN in the RRF CTE covers the asymmetry.


Building search into your Rails app and want it right the first time — hybrid retrieval, proper pgvector indexing, and benchmarked recall? TTB Software builds production-grade search and AI integrations on Rails. We have been doing this for nineteen years.

#rails-hybrid-search #pgvector-rails #semantic-search-rails #reciprocal-rank-fusion #rails-full-text-search #vector-search-rails

Related Articles

Last section. Then please call.

It's a phone call. That's the worst it can get.

No discovery deck. No 45-minute "qualification" call. 30 minutes, your problem, my opinion. If we're a fit, you'll know by minute 12.

Direct line — answered by Roger
+31 6 5123 6132
Mon–Fri, 09:00–18:00 CET · Currently available

OR
info@ttb.software