Skip to content

HubSpot sync operations

How a client's HubSpot data reaches Rose, how failures are reported, and how to diagnose and repair a stuck or failed sync.

How data is synced

  • Initial full sync. Connecting HubSpot from the backoffice inserts a run in hubspot.crm_hubspot_sync_runs and starts the Cloud Run job hubspot-contacts-sync-<env> for that site (60-minute timeout, no retry).
  • Afterwards, webhooks. HubSpot CRM webhooks keep contacts, companies and leads current through the integrations API and Cloud Tasks. No scheduler runs the job again; it only runs on connect or when an operator starts it.
  • Backoffice badge. The "Initial Sync Status" card shows "In progress" while a queued or running run younger than 2 hours exists. Otherwise it shows the connection's last_sync_status, reported as failed (sync_interrupted) when it is still queued/running with no active run and the connection has not changed for 2 hours. "Ran …" is last_sync_at, the last successful run.

Alerts

Signal Where Fires on
"Job failure in production" Cloud Monitoring → Rose production Slack Any failed execution of a *-production job, including a run closed as job_terminated (Cloud Run timeout or cancellation)
"HubSpot webhooks: nothing processed in 1h (production)" Cloud Monitoring → Rose production Slack No successful task on Cloud Tasks queue hubspot-crm-webhooks-production for 1 hour: HubSpot stopped sending, intake is down, the queue is paused, or the worker rejects every task
"HubSpot webhooks: backlog above 50 for 2h (production)" Cloud Monitoring → Rose production Slack More than 50 tasks queued for 2 hours: the worker is failing or too slow on part of the traffic

Each webhook alert opens once and closes when the flow resumes. Policy definitions and calibration: Webhook alert policies.

An execution killed without SIGTERM (out of memory) also fails, but leaves its run active. After 2 hours the badge shows failed (sync_interrupted), and the next run for that site closes the run with stale_run_never_finished.

Individual webhook events that fail processing (hubspot.crm_hubspot_webhook_events.status = 'failed') are not alerted by design.

Is the webhook flow running?

Count recent events by status. Always filter on a recent created_at window: on this table, a filter on connection_id or status alone times out (57014).

since=$(date -u -v-6H +%Y-%m-%dT%H:%M:%SZ)
for s in processed ignored pending processing failed; do
  curl -s -o /dev/null -D - "$SUPABASE_URL/rest/v1/crm_hubspot_webhook_events?select=id&status=eq.$s&created_at=gte.$since&limit=1" \
    -H "apikey: $SUPABASE_SERVICE_ROLE_KEY" -H "Authorization: Bearer $SUPABASE_SERVICE_ROLE_KEY" \
    -H "Accept-Profile: hubspot" -H "Prefer: count=exact" | grep -i content-range
done

Healthy: thousands of processed events in 6h on a weekday, and pending near 0. ignored (obsolete) events are out-of-order property updates and are expected.

When a webhook alert fires

Check in order; the first failing step is the cause.

  1. Queue state. PAUSED means someone paused it; resume it once the reason is known.

    gcloud tasks queues describe hubspot-crm-webhooks-production \
      --location europe-west1 --project inboundx --format="value(state)"
    
  2. Worker responses. In Logs Explorer, 401 means the Cloud Tasks OIDC token is rejected (service account hubspot-webhook-tasks@inboundx.iam.gserviceaccount.com or audience); 500 carries the traceback in the matching application log.

    resource.type="cloud_run_revision"
    resource.labels.service_name="integrations-production-production"
    httpRequest.requestUrl:"/webhooks/tasks/process"
    httpRequest.status>=400
    
  3. Intake from HubSpot. Same filter with /webhooks/crm and no status clause. No request at all means HubSpot stopped sending: check the Rose app's webhook target URL and subscriptions in the HubSpot developer account. 401 means the signature is rejected (HUBSPOT_CLIENT_SECRET mismatch); 500 means intake failed to store or enqueue events.

Webhook alert policies

Hand-managed with gcloud, like the other Cloud Monitoring policies in the infrastructure README. Both read the native Cloud Tasks metrics of queue hubspot-crm-webhooks-production (europe-west1); no log-based metric. Background: IX-5463. Live since 2026-10-08 as alertPolicies/14381886382622858097 (nothing processed) and alertPolicies/15148253824782898672 (backlog). The commands below create a policy; to change a live one, delete it first or edit it in the Console.

Calibrated on 2026-08-27 → 2026-10-08: production never went more than 10 minutes without a successful task (quietest hour: 27, Saturday 04:00 UTC). A 2,000-task burst on 2026-09-10 drained in 80 minutes; the 2026-09-01 incident kept 8,000–11,000 tasks queued for 12 hours while still processing about 200 tasks an hour.

  • Nothing processed: no successful task attempt in 1 hour. Covers HubSpot no longer sending, intake down, queue paused, and a worker rejecting every task. task_attempt_count writes explicit zero points, so the condition is a threshold with missing data counted as a violation, not conditionAbsent.
  • Backlog: queue depth above 50 for 2 hours. Covers a worker too slow or failing on part of the traffic. A full stop with HubSpot still sending raises both alerts, the backlog one an hour later.
cat > /tmp/hubspot-webhooks-nothing-processed-policy.json <<'EOF'
{
  "displayName": "HubSpot webhooks: nothing processed in 1h (production)",
  "combiner": "OR",
  "severity": "ERROR",
  "enabled": true,
  "notificationChannels": ["projects/inboundx/notificationChannels/864254260479040154"],
  "alertStrategy": {"autoClose": "604800s", "notificationPrompts": ["OPENED", "CLOSED"]},
  "conditions": [{
    "displayName": "No successful hubspot-crm-webhooks-production task in 1h",
    "conditionThreshold": {
      "filter": "resource.type = \"cloud_tasks_queue\" AND resource.labels.queue_id = \"hubspot-crm-webhooks-production\" AND metric.type = \"cloudtasks.googleapis.com/queue/task_attempt_count\" AND metric.labels.response_code = \"ok\"",
      "aggregations": [{"alignmentPeriod": "3600s", "perSeriesAligner": "ALIGN_SUM"}],
      "comparison": "COMPARISON_LT",
      "thresholdValue": 0.5,
      "duration": "60s",
      "trigger": {"count": 1},
      "evaluationMissingData": "EVALUATION_MISSING_DATA_ACTIVE"
    }
  }],
  "documentation": {
    "mimeType": "text/markdown",
    "subject": "HubSpot webhooks: nothing processed in 1h (production)",
    "content": "No HubSpot CRM webhook task succeeded in the last hour: client CRM data in Rose is no longer updated.\n\nCheck in order: queue `hubspot-crm-webhooks-production` state (paused?), worker responses on `/webhooks/tasks/process` (401 = OIDC, 500 = processing), intake requests on `/webhooks/crm` (none = HubSpot stopped sending; 401 = signature rejected).\n\nRunbook: docs/src/backend/hubspot-sync.md, section \"When a webhook alert fires\"."
  }
}
EOF
gcloud alpha monitoring policies create --project=inboundx \
  --policy-from-file=/tmp/hubspot-webhooks-nothing-processed-policy.json

cat > /tmp/hubspot-webhooks-backlog-policy.json <<'EOF'
{
  "displayName": "HubSpot webhooks: backlog above 50 for 2h (production)",
  "combiner": "OR",
  "severity": "ERROR",
  "enabled": true,
  "notificationChannels": ["projects/inboundx/notificationChannels/864254260479040154"],
  "alertStrategy": {"autoClose": "604800s", "notificationPrompts": ["OPENED", "CLOSED"]},
  "conditions": [{
    "displayName": "hubspot-crm-webhooks-production depth > 50 for 2h",
    "conditionThreshold": {
      "filter": "resource.type = \"cloud_tasks_queue\" AND resource.labels.queue_id = \"hubspot-crm-webhooks-production\" AND metric.type = \"cloudtasks.googleapis.com/queue/depth\"",
      "aggregations": [{"alignmentPeriod": "300s", "perSeriesAligner": "ALIGN_MEAN"}],
      "comparison": "COMPARISON_GT",
      "thresholdValue": 50,
      "duration": "7200s",
      "trigger": {"count": 1}
    }
  }],
  "documentation": {
    "mimeType": "text/markdown",
    "subject": "HubSpot webhooks: backlog above 50 for 2h (production)",
    "content": "More than 50 HubSpot CRM webhook tasks have been queued for 2 hours: client CRM data in Rose is delayed.\n\nCheck worker responses on `/webhooks/tasks/process` (status >= 400) and the integrations API latency.\n\nRunbook: docs/src/backend/hubspot-sync.md, section \"When a webhook alert fires\"."
  }
}
EOF
gcloud alpha monitoring policies create --project=inboundx \
  --policy-from-file=/tmp/hubspot-webhooks-backlog-policy.json

Diagnose a site

Recent runs and the connection state (read-only):

select r.id, r.run_type, r.status, r.created_at, r.finished_at, r.error_message
from hubspot.crm_hubspot_sync_runs r
join hubspot.crm_hubspot_connections c on c.id = r.connection_id
where c.site_domain = '<domain>'
order by r.created_at desc
limit 10;

select status, last_sync_at, last_sync_status, last_sync_error
from hubspot.crm_hubspot_connections
where site_domain = '<domain>';

Match a run to its execution by start time:

gcloud run jobs executions list --job hubspot-contacts-sync-production \
  --region europe-west9 --project inboundx --limit 20

Known errors:

Error Meaning
job_terminated: … Cloud Run stopped the execution: 60-minute timeout or a cancellation from the console.
stale_run_never_finished An earlier run never finished (container killed). Closed automatically; check that execution's logs.
lock_not_acquired Another run for the same site was running when an operator-triggered run started.
dense_window_exceeds_hubspot_10k_cap More than 10,000 records share one modification second (bulk import). HubSpot search cannot page past 10,000, so the initial sync cannot finish; webhooks still deliver later changes.

Re-run a sync

just job run hubspot production <domain>

Run from backend/. The job closes stale runs for that site, then runs an initial full or incremental sync according to hubspot.crm_hubspot_sync_state.sync_mode. Disconnecting and reconnecting HubSpot from the backoffice also starts a fresh initial sync, but it deletes the site's synced CRM data first.