yrby
Docs · 3 of 7 View as Markdown

Storage

A channel only needs two hooks, on_load and on_change, and they can talk to any store. yrby comes with an Active Record store because most apps want one. You can also keep documents somewhere else, or nowhere at all.

The bundled models

The models come with the gem, the same way ActionText::RichText comes with Action Text.

Y::Document stores one row per document. Each row has a unique key, which is what channels use. Your app can pick the key, and yrby never parses it. A row can also point at a model attribute through a polymorphic record and a name, such as "body". Rows created by key alone leave those two nil.

You can create a row by key or by record, in either order. Y::Document.for(record, name) finds or creates the row for a record's attribute, with a readable key like post/1/body. If a channel already created a row with that key, for links it to the record, so both end up on the same document.

The row also holds state, the merged snapshot of the document, and nothing else. If you want rendered HTML or search text, compute it yourself, usually in the channel's on_change. By default the channel reads with .load_state(key) and writes with .append(key, update).

Y::DocumentUpdate holds the changes that haven't been compacted yet, one per row. Once there are compact_every of them (64 by default), yrby merges them into state and deletes the rows. Loading a document reads the snapshot plus any remaining rows. A row lock keeps two compactions of the same document from running at once. Rows that belong to an open gap are marked pending and kept until the gap closes. Destroying a document deletes its update rows too.

The migration creates y_documents and y_document_updates. To rename them, edit the generated migration and set Y::Document.table_name and Y::DocumentUpdate.table_name to the new names in an initializer.

Encrypted storage

Y::EncryptedDocument uses the same tables and encrypts state and the updates with Active Record encryption. You turn it on in the model, so no page or client can switch it off:

class Post < ApplicationRecord
  has_collaborative_document :body, encrypted: true
end

Y::DocumentChannel sees that and loads and saves the attribute through the encrypted class. Other attributes use plain Y::Document. In your own channel, point on_load and on_change at Y::EncryptedDocument. Either way, you need Active Record encryption keys set up, and you should always read a document through the same class. Reading an encrypted row through the plain class gives you ciphertext.

Record-backed access

post.collaborative_document(:body) returns a Y::Collaborative::Attribute with load_state, append(update), key, and y_doc. y_doc builds a fresh Y::Doc you can read and render in Ruby. For row-level work like compaction, call post.collaborative_document(:body).document_row.compact!. The channel uses this same object, so encryption works the same way in both places.

An attribute's key is the key its row was saved under. Before a row exists, the key is Y::Document.key_for(record, name). Computing it doesn't create a row.

For a store other than Y::Document, write your own channel with on_load and on_change hooks, as shown below.

Writing your own store

Your store needs to do two things.

Load without losing pending updates, and accept duplicates

on_load has to return state that still includes pending updates. Use encode_state_as_update, or replay the raw log. Don't compact with compacted_state_update while doc.pending? is true. That drops the pending struct, and with it an edit the server already confirmed.

When an ack gets lost, the client resends an update the store already has. Replaying the log still produces the same document, because applying a CRDT update twice has no effect. Deduplicating is optional. If log size matters, deduplicate by content hash:

class DocumentStore
  # Appending a duplicate update is a no-op upsert.
  def append(key, update)
    Revision.upsert({ doc_key: key, update_hash: Digest::SHA256.hexdigest(update), update: update },
                    unique_by: %i[doc_key update_hash])
  end

  # Replay the raw log so a pending struct is kept and can integrate
  # when its dependency arrives.
  def load(key)
    updates = Revision.where(doc_key: key).order(:id).pluck(:update)
    return nil if updates.empty?

    doc = Y::Doc.new
    updates.each { |u| doc.apply_update(u) }
    doc.encode_state_as_update # keeps pending structs
  end

  # Optional. Compact only when no gap is open.
  def compact(key)
    doc = Y::Doc.new
    Revision.where(doc_key: key).order(:id).pluck(:update).each { |u| doc.apply_update(u) }
    return if doc.pending? # a gap is open, and compacting now would drop it
    # ... replace the log with one revision holding doc.compacted_state_update ...
  end
end

Watch for gaps that don't close

An open gap is easy to miss, because the pending edit doesn't appear in the document until its dependency arrives. Usually the gap closes by itself. The sender resends the missing update until the server acknowledges it. Every handshake also asks the client for everything the server is missing. Use the on_gap hook to emit a metric, so you can see a gap that doesn't close.

Pending structs and gap-free state

When a doc gets an update whose dependency is missing, yrs holds it as a pending struct and applies it if the dependency arrives later. Until then, the doc's state vector doesn't include it, and Doc#pending? returns true.

Pending updates are stored and sent like any other state. Don't fold them into a compacted snapshot.

  • Doc#compacted_state_update returns the full state without pending structs. Use it for compaction, since a snapshot that included them would leave them pending forever. It doesn't change the doc.
  • encode_state_as_update includes pending structs. Use it for persistence and for sending state, so a gap can still close.

Ephemeral documents (no database)

Some documents only need to last for one session, like a scratchpad, live form state, or a draft you save on submit. For those, the channel can keep the document on the connection.

class ScratchpadChannel < ApplicationCable::Channel
  include Y::ActionCable

  on_load { |key| @doc_state }

  on_change do |key, update|
    doc = Y::Doc.new
    doc.apply_update(@doc_state) if @doc_state
    doc.apply_update(update)
    @doc_state = doc.encode_state_as_update
  end

  def subscribed    = sync_subscribed(params[:id])
  def receive(data) = sync_receive(data, params[:id])

  private

  # Each connection has its own scratchpad.
  def authorized?(_key) = true
end

On Action Cable the channel instance lasts as long as the connection, so an instance variable is all the store you need. AnyCable builds a new channel instance for each message. There, declare the store as channel state with state_attr_accessor from anycable-rails, and Base64-encode it. AnyCable serializes that state as JSON in every RPC call to anycable-go, so keep these documents small.

Each connection has its own copy, which limits where this fits. With one person editing, you get every delivery guarantee and no database. With several people, one client's update can depend on edits its connection hasn't seen. The server saves it as pending, and the next handshake with that client fills the gap. Everyone still ends up with the same document, but with heavy editing, more updates sit pending between handshakes than they would with a shared store.

The connection and the browsers hold the only copies. After a server restart, a reconnecting client sends its state back through the normal sync handshake. The document survives as long as some client still has it.

The store this site uses

This site's shape demos (spreadsheet, whiteboard, kanban, code, and Tiptap) use the same store this page describes, on SQLite:

class DocumentChannel < ApplicationCable::Channel
  include Y::ActionCable

  on_load   { |key|         Y::Document.load_state(key) }
  on_change { |key, update| Y::Document.append(key, update) }
end

SQLite isn't required. It's just what this site uses. On top of the hooks, the site caps people per room, documents on disk, and bytes per document. A sweeper deletes rooms nobody has edited for a day, since public, anonymous documents shouldn't stick around.

The site runs on AnyCable, where each command gets a new channel instance. So anything that has to last between commands goes in state_attr_accessor, as AnyCable and multi-process describes.

The full site, including every rate and size limit, is in site/. Its README explains the setup.

This page is copied from the repo README. If the two disagree, go by the README.