Skip to content

Retrain a live client in test, then promote it to production

Use this procedure to rebuild an existing production client's knowledge in test (for example with the new website scraping system), check it there, and then make production serve the result. Production keeps serving the client's current knowledge until you finish.

How it works

Test and production share one Supabase project, so they share the client's sources (knowledge.documents, FAQs, configuration). Each environment has its own retrieval stores (MongoDB and Neo4j), which is what the chatbot answers from.

Retraining in test changes the shared sources. Without a hold, production's automatic source sync sees the change within a minute and retrains production from the new sources before you have checked them. A migration hold on the client stops production's automatic sync for all the client's workspaces until you finish or abort. Finishing copies test's MongoDB and Neo4j data to production and records that production is now up to date, so it does not retrain over the copy.

Prerequisites

  • IX-5430 (stop production auto-sync from retraining apptest workspaces) is released to production: the hold table exists in the production project.
  • just target from the repository root prints SHARED. A hold written to a preview branch or local database has no effect on production.
  • Backend credentials for test and production (backend/.env.test, backend/.env.production): finish writes production MongoDB and Neo4j.
  • No other migration hold is open for the client (rose-tenant migration status).

Steps

Start and finish the migration from the command line (from backend/); do the training in between in the test backoffice, apptest.userose.ai.

  1. Hold the client.

    rose-tenant migration start acme.com --env test --reason "new website scraping"
    

    The hold covers every workspace of the client that owns the domain, path workspaces included. The command prints the held workspaces, and any operation production had already started for them: wait for those to finish before training, because they read the shared sources when they start loading.

  2. Switch the client to the new system if needed. Knowledge-system flags are shared configuration and apply to both environments at once. Production does not retrain during the hold, but the client now sees the Website tab in the production backoffice too: tell them not to apply page changes until you finish.

    rose-tenant set-knowledge-migration-flags acme.com --website-migrated \
        --documents-migrated --env test --yes
    
  3. Train in the test backoffice. On apptest.userose.ai, select the client's workspace, open Data Sources → Website (/config-studio/data-sources/website), run Re-scan website, adjust Pages to sync, then Apply selection & ingest. Repeat for each held workspace that needs it. Apptest runs in test, so this trains test's knowledge only. Apptest refuses knowledge changes for any client that is not held: if it answers "Knowledge updates are disabled on the test environment", check rose-tenant migration status.

    Never do this step on app.userose.ai: that ingests into production, and the hold does not block it.

    Command-line alternative
    rose-tenant onboard acme.com --env test --until discover --yes   # prints the run id
    rose-document-loader review RUN --env test --export selection.json
    # edit only the included flags in selection.json, then:
    rose-document-loader review RUN --env test --apply selection.json
    rose-tenant onboard acme.com --env test --resume RUN --until sync --yes
    

    See onboarding and sources. rose-document-loader sync acme.com --env test without --run selects nothing; it only reloads already-published sources into test.

    Scanning and page review stay private. Ingesting publishes the selected pages to the shared sources: from then on, production would retrain from them without the hold.

  4. Check test. Compare the corpus and chat with it until you are satisfied:

    rose-tenant compare-envs acme_com --source-env test --target-env production
    rose-tenant status acme.com --env test
    rose-chat "What do you do?" --site acme.com --env test
    

    For a full answer-quality check, run the client's evaluation against test (rose-eval).

  5. Promote to production.

    rose-tenant migration finish acme.com
    

    For each held workspace, finish replaces production's MongoDB and Neo4j data with test's (copy-tenant), records test's loaded revisions as production's, then releases the hold. It asks for confirmation unless you pass --yes, and refuses while a knowledge operation is still running in test.

Expected result

  • finish prints ✔ Released <storage key> for each workspace.
  • A line ▸ <scope>: test loaded revision … ; production will retrain it means test had not loaded the latest version of that source. Production retrains that source automatically after the release; nothing else needs doing.
  • A failed copy stops finish and keeps that workspace's hold. Fix the cause and run finish again, or abort.

Verification

rose-tenant migration status                  # the client is no longer listed
rose-tenant compare-envs acme_com --source-env test --target-env production  # counts match
rose-chat "What do you do?" --site acme.com --env production

compare-envs takes a tenant id: the domain with dots replaced by underscores (acme_com), or the site:<uuid> storage key that migration start prints for a path workspace. Run it once per held workspace.

Abort

rose-tenant migration abort acme.com

This releases the hold without copying anything. Production then retrains from the current shared sources on its own, which include whatever test changed. To go back to the old knowledge instead, revert the source changes (for example, the knowledge-system flags) before aborting.

During the hold

  • Do not retrain production by hand. The hold does not block manual retrains, and a production retrain reads the new sources. finish still copies test over it, but production would serve the new, unchecked knowledge in the meantime.
  • Client edits reach production only at the end. An FAQ or document change made during the hold reaches production through the copy if test was retrained after it. Otherwise production retrains that source right after finish.
  • Keep holds short. rose-tenant migration status shows each hold's age and reason. A forgotten hold stops production sync for that client.
  • There is no built-in rollback. After finish, production's previous knowledge is gone. To keep a copy, run rose-tenant copy-tenant <tenant> --source-env production --dest-env staging before finishing. That copy overwrites the client's staging data.