Rails LLM Conversation History: Persisting Context, Summarization, and Managing the Context Window
LLM conversation history in Rails: persist message turns, manage context windows, implement summarization, and build multi-turn chat that actually scales.
The demo looked great. Streaming responses, a clean chat interface, Claude answering questions about the product in fluent, helpful prose. We shipped it. Three days later, the first real user complaint arrived: “Your chatbot doesn’t remember what I told it two messages ago.”
I looked at the code. The developer had built the streaming UI correctly — Server-Sent Events, Turbo Streams, the whole thing. But the controller action reconstructed the conversation array on every request by loading only the most recent user message from the database. The context array sent to the API was always two entries long: the system prompt and the current input. No history. Every message was a fresh start. Cosmetically a chat. Functionally a single-turn Q&A with a fancy spinner.
This is the mistake I see most often in Rails LLM applications. Developers nail the streaming interface, implement tool use, sometimes even build RAG with pgvector — and then hardcode a context window that holds only the latest exchange, or nothing at all. After nineteen years of Rails and the last three building serious AI-integrated applications, here is the full picture: how to store LLM conversation history correctly, how to manage the context window without blowing your token budget, and how to implement summarization so conversations can run indefinitely without degrading.
The Database Schema: Getting the Fundamentals Right
Before touching any LLM API, you need a schema that can actually represent conversations. The minimum viable design:
# db/migrate/20260929100000_create_conversations_and_messages.rb
class CreateConversationsAndMessages < ActiveRecord::Migration[8.0]
def change
create_table :conversations do |t|
t.references :user, null: false, foreign_key: true
t.string :title
t.text :summary # compressed history once context is trimmed
t.integer :summary_token_count, default: 0
t.integer :total_token_count, default: 0
t.timestamps
end
create_table :messages do |t|
t.references :conversation, null: false, foreign_key: true
t.string :role, null: false # "user" | "assistant" | "tool"
t.text :content, null: false
t.string :model # which model served this turn
t.integer :input_tokens
t.integer :output_tokens
t.integer :position, null: false # explicit ordering within conversation
t.boolean :included_in_summary, default: false
t.jsonb :tool_calls # raw tool-use metadata when relevant
t.timestamps
end
add_index :messages, [:conversation_id, :position]
add_index :messages, [:conversation_id, :included_in_summary]
end
end
A few design decisions worth explaining:
position — an explicit sequence number within a conversation. Do not rely on created_at for ordering. Clock skew, bulk inserts, and transaction isolation can make timestamp ordering unreliable. An explicit counter is not premature optimization; it is correct.
included_in_summary — marks which messages have been rolled up into the summary column. Once you compress old messages into a summary, you need a reliable way to exclude them from the context array without deleting them. Deletion loses your audit trail; this flag keeps the rows while letting the query ignore them.
summary on the conversation — this is where compressed context lives. It starts null and gets populated by the summarizer as the conversation grows.
tool_calls JSONB — if you are using tool use or function calling, store the raw exchange metadata here. You will want to reproduce the exact tool-call/result sequence to the API if you ever reconstruct a conversation or debug a reasoning chain.
Building the Context Array
The LLM API expects a message array: [{role: "user", content: "..."}, {role: "assistant", content: "..."}, ...]. Assembling it correctly from the database is the core operation your LLM conversation history layer must get right.
# app/models/conversation.rb
class Conversation < ApplicationRecord
belongs_to :user
has_many :messages, -> { order(:position) }, dependent: :destroy
MAX_ACTIVE_TOKENS = 80_000 # headroom for system prompt + response
def context_messages
active = messages.where(included_in_summary: false)
if summary.present?
# Prepend a synthetic exchange that establishes compressed context
[
{ role: "user",
content: "Summary of our earlier conversation:\n\n#{summary}" },
{ role: "assistant",
content: "Understood. I have the context from our earlier conversation and will build on it." }
] + active.map(&:to_api_format)
else
active.map(&:to_api_format)
end
end
def next_position
(messages.maximum(:position) || 0) + 1
end
end
# app/models/message.rb
class Message < ApplicationRecord
belongs_to :conversation
def to_api_format
{ role: role, content: content }
end
end
The controller stays clean because the complexity lives in the model:
# app/controllers/chat_controller.rb
class ChatController < ApplicationController
def create
conversation = current_user.conversations.find(params[:conversation_id])
user_input = params[:message].to_s.strip
return head :unprocessable_entity if user_input.blank?
conversation.messages.create!(
role: "user",
content: user_input,
position: conversation.next_position
)
response = ClaudeClient.chat(
system: system_prompt_for(current_user),
messages: conversation.context_messages
)
assistant_text = response.content.first.text
conversation.messages.create!(
role: "assistant",
content: assistant_text,
model: response.model,
input_tokens: response.usage.input_tokens,
output_tokens: response.usage.output_tokens,
position: conversation.next_position
)
conversation.increment!(:total_token_count,
response.usage.input_tokens + response.usage.output_tokens)
ConversationSummarizerJob.perform_later(conversation.id) if summarization_needed?(conversation)
render json: { content: assistant_text }
end
private
def summarization_needed?(conversation)
active_tokens = conversation.messages
.where(included_in_summary: false)
.pick(Arel.sql("COALESCE(SUM(input_tokens), 0) + COALESCE(SUM(output_tokens), 0)"))
.to_i
active_tokens > Conversation::MAX_ACTIVE_TOKENS * 0.8
end
def system_prompt_for(user)
"You are a helpful assistant for #{user.company_name}."
end
end
The Context Window Problem
Here is where most implementations fail. Claude Sonnet has a 200K token context window. GPT-4o has 128K. These sound enormous. They are not unlimited, and the costs compound in ways developers do not immediately see.
At current pricing, 80K tokens of input on every request in a busy conversation costs real money. Multiply that across concurrent users, and the numbers become uncomfortable fast. But cost is not the only problem.
Latency. Sending 80K tokens per request means waiting for the model to process all of them before it produces the first output token. Long contexts are slower than short ones — even with providers who claim near-constant TTFT, there is still measurable latency drag.
Model quality degradation. Counterintuitively, very long contexts can hurt response quality. Relevant information buried early in a 150K-token conversation competes with more recent tokens for the model’s attention. The “lost in the middle” phenomenon — where models underperform on information that is neither at the beginning nor the end of the context — is well documented.
The right approach is a sliding window with summarization: keep the most recent N messages as raw context, compress older messages into a summary, and reconstruct the full context from summary plus recent messages on every request.
Summarization: The LLM Compressing Itself
When the active message window approaches the limit, you roll up the oldest batch into a summary. Future requests start with the summary instead of the raw messages.
# app/jobs/conversation_summarizer_job.rb
class ConversationSummarizerJob < ApplicationJob
queue_as :default
BATCH_SIZE = 20 # compress in batches of this many messages
def perform(conversation_id)
conversation = Conversation.find(conversation_id)
to_compress = conversation.messages
.where(included_in_summary: false)
.order(:position)
.limit(BATCH_SIZE)
return if to_compress.count < BATCH_SIZE
transcript = to_compress.map do |m|
"#{m.role.upcase}: #{m.content}"
end.join("\n\n")
existing_summary = conversation.summary.presence
prompt = if existing_summary
<<~PROMPT
You are compressing a conversation for long-term storage.
EXISTING SUMMARY:
#{existing_summary}
NEW MESSAGES TO INCORPORATE:
#{transcript}
Write a new comprehensive summary combining both. Preserve: key decisions,
user preferences, established facts, open questions, and technical details.
Be specific. Use past tense. Maximum 400 words.
PROMPT
else
<<~PROMPT
Summarize this conversation transcript for later reference.
Preserve: key decisions, user preferences, established facts,
open questions, and technical details. Be specific. Use past tense.
Maximum 300 words.
TRANSCRIPT:
#{transcript}
PROMPT
end
response = ClaudeClient.chat(
system: "You are a concise conversation summarizer. Return only the summary text.",
messages: [{ role: "user", content: prompt }],
max_tokens: 600
)
new_summary = response.content.first.text
Conversation.transaction do
conversation.update!(
summary: new_summary,
summary_token_count: response.usage.input_tokens + response.usage.output_tokens
)
to_compress.update_all(included_in_summary: true)
end
end
end
The transaction wrapping the write is not optional. If update! fails, the messages must stay active so the job can retry safely. Marking messages as summarized before the summary is persisted is the kind of race condition that silently loses context.
The progressive summarization pattern above — incorporating the existing summary rather than re-summarizing from scratch — keeps the cost linear. Each summarization pass costs roughly 1-2K tokens to produce a 300-400 word summary. Compare that to the 15-25K tokens you save on every subsequent request by not including those raw messages.
Per-User Memory That Persists Across Conversations
Conversation-level context management solves one problem. A separate problem: users expect an AI assistant to remember things across sessions. Their preferred programming language. The stack they are running. That they already explained their billing situation last week.
This is user-level memory, not conversation history, and it warrants a separate model:
create_table :user_memories do |t|
t.references :user, null: false
t.string :key, null: false # "preferred_language", "tech_stack", etc.
t.text :value, null: false
t.datetime :last_accessed_at
t.timestamps
end
add_index :user_memories, [:user_id, :key], unique: true
Extract facts from conversations as a background job:
# app/services/memory_extractor.rb
class MemoryExtractor
EXTRACTION_PROMPT = <<~PROMPT
Extract key-value pairs worth remembering about this user for future conversations.
Focus on preferences, facts, and technical context — not conversational details.
Return a JSON array only: [{"key": "snake_case_name", "value": "one sentence fact"}]
Return [] if nothing is worth extracting.
PROMPT
def self.extract(conversation)
recent = conversation.messages
.where(included_in_summary: false)
.order(position: :desc)
.limit(10)
.map(&:to_api_format)
response = ClaudeClient.chat(
system: "You extract user facts for memory storage. Return JSON only.",
messages: recent + [{ role: "user", content: EXTRACTION_PROMPT }],
max_tokens: 400
)
JSON.parse(response.content.first.text)
rescue JSON::ParserError
[]
end
end
Upsert the extracted facts:
MemoryExtractor.extract(conversation).each do |fact|
next unless fact["key"].present? && fact["value"].present?
conversation.user.user_memories.upsert(
{ key: fact["key"].to_s.strip,
value: fact["value"].to_s.strip,
last_accessed_at: Time.current },
unique_by: [:user_id, :key]
)
end
And inject the most relevant memories into the system prompt:
def system_prompt_for(user)
memories = user.user_memories
.order(last_accessed_at: :desc)
.limit(15)
.map { |m| "- #{m.key}: #{m.value}" }
.join("\n")
base = "You are a helpful assistant."
return base if memories.blank?
<<~PROMPT
#{base}
What you know about this user:
#{memories}
Reference this context when relevant. Do not repeat it unprompted.
PROMPT
end
Multi-Tenant Access Control
If you are building a SaaS application, conversation history is personal data, and the access control mistake is easy to make:
# Wrong — broken access control
def show
@conversation = Conversation.find(params[:id])
end
# Correct — always scope to current_user
def show
@conversation = current_user.conversations.find(params[:id])
end
The scoped query generates WHERE user_id = ? AND id = ?. An unauthorized user ID returns 404. Without the scope, any authenticated user can read any conversation by guessing IDs.
If your application also implements PostgreSQL row-level security, conversations become a natural candidate for an RLS policy — the database enforces the tenant boundary rather than relying on application-layer scopes.
For team conversations where multiple users contribute to a shared thread:
create_table :conversation_participants do |t|
t.references :conversation, null: false
t.references :user, null: false
t.string :role, default: "member" # "owner" | "member"
t.timestamps
end
add_index :conversation_participants, [:conversation_id, :user_id], unique: true
The authorization check becomes a join through the participants table rather than a direct user_id match.
Loading History Efficiently
Two distinct queries to optimize separately:
For the API context array — you want only unsummarized messages, loaded in order. The composite index on (conversation_id, included_in_summary) makes this a narrow index scan rather than a table scan. Confirm it is being used with EXPLAIN ANALYZE on a conversation with hundreds of messages; the planner occasionally prefers a sequential scan on small datasets and will switch when the table grows.
For UI display — you want paginated history including summarized messages, for the conversation transcript view. Paginate with a cursor on position rather than OFFSET, especially once conversations run into hundreds of rows:
# cursor-based pagination on position
def messages_before(position:, limit: 30)
messages.where("position < ?", position).order(position: :desc).limit(limit)
end
For the token estimate, avoid loading all message records just to count tokens:
def active_token_estimate
messages
.where(included_in_summary: false)
.pick(Arel.sql(
"COALESCE(SUM(input_tokens), 0) + COALESCE(SUM(output_tokens), 0)"
))
.to_i + summary_token_count.to_i
end
Connecting to Cost Tracking and RAG
If you are already tracking LLM costs per tenant, the input_tokens and output_tokens columns on messages feed directly into that system. A simple GROUP BY gives you per-conversation cost; rolling up by user or organization is one aggregation further.
If you have pgvector set up for RAG retrieval, consider storing a vector embedding of each message. Instead of always loading the N most recent messages, you can retrieve the K most semantically relevant previous exchanges for the current query and then append the last 3-4 turns for recency. This hybrid approach — semantic retrieval plus recency window — is especially effective for support applications where users ask similar questions across multiple separate conversations.
Testing the Persistence Layer
The failure modes worth covering in your test suite:
# test/models/conversation_test.rb
class ConversationTest < ActiveSupport::TestCase
test "context_messages prepends summary preamble when summary is present" do
conversation = conversations(:with_summary)
active_count = conversation.messages.where(included_in_summary: false).count
context = conversation.context_messages
assert_equal active_count + 2, context.length
assert_equal "user", context.first[:role]
assert_includes context.first[:content], conversation.summary
end
test "context_messages with no summary returns all messages in position order" do
conversation = conversations(:fresh)
context = conversation.context_messages
assert_equal conversation.messages.count, context.length
assert_equal conversation.messages.first.content, context.first[:content]
end
test "next_position is always greater than current maximum" do
conversation = conversations(:fresh)
max_before = conversation.messages.maximum(:position)
assert_operator conversation.next_position, :>, max_before
end
end
# test/jobs/conversation_summarizer_job_test.rb
class ConversationSummarizerJobTest < ActiveJob::TestCase
test "marks compressed messages as included_in_summary" do
conversation = conversations(:long)
stub_claude_response("A clear summary of the conversation.")
ConversationSummarizerJob.perform_now(conversation.id)
conversation.reload
assert conversation.summary.present?
assert_operator conversation.messages.where(included_in_summary: true).count, :>, 0
end
test "rolls back if summary update fails" do
conversation = conversations(:long)
Conversation.any_instance.stubs(:update!).raises(ActiveRecord::RecordInvalid)
assert_no_changes -> { conversation.messages.where(included_in_summary: true).count } do
assert_raises(ActiveRecord::RecordInvalid) do
ConversationSummarizerJob.perform_now(conversation.id)
end
end
end
end
FAQ
How many messages should I keep in the active window before summarizing?
A practical baseline: trigger summarization when active messages exceed 20K tokens, and batch-compress the oldest 20 messages at a time. This keeps the active window manageable while summarizing in chunks large enough to produce a coherent summary. If your messages are typically short (one or two sentences), you can use a larger message count. If messages include long code blocks or document pastes, set a lower threshold.
Should I use the LLM to summarize, or can I use extractive techniques?
Use the LLM. Extractive summarization — picking the most “important” sentences — preserves surface wording but loses reasoning chains, decisions, and the causal thread of the conversation. A model-generated abstractive summary captures what was decided and why, not just what was said. The cost is small: generating a 300-word summary typically costs under 1K output tokens, and you save 15-25K input tokens on every subsequent request.
How do I handle conversations that include file attachments or images?
Store attachments separately using Active Storage, and reference them from the message rather than embedding the raw bytes. For images in the active window, pass the vision API format. For images in summarized messages, instruct the summarizer to describe the image in prose — “User shared a screenshot of a Postgres EXPLAIN ANALYZE output showing a sequential scan on the orders table” — so the description carries the relevant information forward without re-attaching the file.
Does this schema work with streaming responses?
Yes. Streaming and persistence are independent concerns. Accumulate the full streamed response in memory, then write a single messages record once the stream completes. The streaming mechanics — SSE, Turbo Streams — do not change. See the post on streaming LLM responses with Action Controller Live for the rendering side.
Building an LLM-integrated Rails application and need it to actually hold a conversation? TTB Software builds and scales production AI features — not demos. We have been doing this for nineteen years, and the last three of them have been increasingly AI-shaped.
Related Articles
Rails API Serialization: Blueprinter, Alba, and JSONAPI-Serializer Compared for Production APIs
Rails API serialization done right: compare Blueprinter, Alba, and jsonapi-serializer with real code, N+1 traps, cach...
Rails Data Migrations: Safe Backfills with data-migrate, Maintenance Tasks, and Batched Updates
Rails data migrations done right: use data-migrate or maintenance_tasks for safe, resumable backfills that don't lock...
Rails Timeouts: Statement, Rack, HTTP Client and Job Timeouts That Prevent Production Cascades
Rails timeouts done right: configure statement_timeout, rack-timeout, HTTP client and background job limits to preven...