HubSpot sync operations¶
How a client's HubSpot data reaches Rose, how failures are reported, and how to diagnose and repair a stuck or failed sync.
How data is synced¶
- Initial full sync. Connecting HubSpot from the backoffice inserts a run in
hubspot.crm_hubspot_sync_runsand starts the Cloud Run jobhubspot-contacts-sync-<env>for that site (60-minute timeout, no retry). - Afterwards, webhooks. HubSpot CRM webhooks keep contacts, companies and leads current through the integrations API and Cloud Tasks. No scheduler runs the job again; it only runs on connect or when an operator starts it.
- Backoffice badge. The "Initial Sync Status" card shows "In progress" while a
queuedorrunningrun younger than 2 hours exists. Otherwise it shows the connection'slast_sync_status, reported asfailed(sync_interrupted) when it is stillqueued/runningwith no active run and the connection has not changed for 2 hours. "Ran …" islast_sync_at, the last successful run.
Alerts¶
| Signal | Where | Fires on |
|---|---|---|
| "Job failure in production" | Cloud Monitoring → Rose production Slack | Any failed execution of a *-production job, including a run closed as job_terminated (Cloud Run timeout or cancellation) |
| "HubSpot webhooks: nothing processed in 1h (production)" | Cloud Monitoring → Rose production Slack | No successful task on Cloud Tasks queue hubspot-crm-webhooks-production for 1 hour: HubSpot stopped sending, intake is down, the queue is paused, or the worker rejects every task |
| "HubSpot webhooks: backlog above 50 for 2h (production)" | Cloud Monitoring → Rose production Slack | More than 50 tasks queued for 2 hours: the worker is failing or too slow on part of the traffic |
Each webhook alert opens once and closes when the flow resumes. Policy definitions and calibration: Webhook alert policies.
An execution killed without SIGTERM (out of memory) also fails, but leaves its run
active. After 2 hours the badge shows failed (sync_interrupted), and the next run for
that site closes the run with stale_run_never_finished.
Individual webhook events that fail processing
(hubspot.crm_hubspot_webhook_events.status = 'failed') are not alerted by design.
Is the webhook flow running?¶
Count recent events by status. Always filter on a recent created_at window: on this
table, a filter on connection_id or status alone times out (57014).
since=$(date -u -v-6H +%Y-%m-%dT%H:%M:%SZ)
for s in processed ignored pending processing failed; do
curl -s -o /dev/null -D - "$SUPABASE_URL/rest/v1/crm_hubspot_webhook_events?select=id&status=eq.$s&created_at=gte.$since&limit=1" \
-H "apikey: $SUPABASE_SERVICE_ROLE_KEY" -H "Authorization: Bearer $SUPABASE_SERVICE_ROLE_KEY" \
-H "Accept-Profile: hubspot" -H "Prefer: count=exact" | grep -i content-range
done
Healthy: thousands of processed events in 6h on a weekday, and pending near 0.
ignored (obsolete) events are out-of-order property updates and are expected.
When a webhook alert fires¶
Check in order; the first failing step is the cause.
-
Queue state.
PAUSEDmeans someone paused it; resume it once the reason is known. -
Worker responses. In Logs Explorer,
401means the Cloud Tasks OIDC token is rejected (service accounthubspot-webhook-tasks@inboundx.iam.gserviceaccount.comor audience);500carries the traceback in the matching application log. -
Intake from HubSpot. Same filter with
/webhooks/crmand no status clause. No request at all means HubSpot stopped sending: check the Rose app's webhook target URL and subscriptions in the HubSpot developer account.401means the signature is rejected (HUBSPOT_CLIENT_SECRETmismatch);500means intake failed to store or enqueue events.
Webhook alert policies¶
Hand-managed with gcloud, like the other Cloud Monitoring policies in the
infrastructure README.
Both read the native Cloud Tasks metrics of queue hubspot-crm-webhooks-production
(europe-west1); no log-based metric. Background:
IX-5463. Live since 2026-10-08 as
alertPolicies/14381886382622858097 (nothing processed) and
alertPolicies/15148253824782898672 (backlog). The commands below create a policy;
to change a live one, delete it first or edit it in the Console.
Calibrated on 2026-08-27 → 2026-10-08: production never went more than 10 minutes without a successful task (quietest hour: 27, Saturday 04:00 UTC). A 2,000-task burst on 2026-09-10 drained in 80 minutes; the 2026-09-01 incident kept 8,000–11,000 tasks queued for 12 hours while still processing about 200 tasks an hour.
- Nothing processed: no successful task attempt in 1 hour. Covers HubSpot no
longer sending, intake down, queue paused, and a worker rejecting every task.
task_attempt_countwrites explicit zero points, so the condition is a threshold with missing data counted as a violation, notconditionAbsent. - Backlog: queue depth above 50 for 2 hours. Covers a worker too slow or failing on part of the traffic. A full stop with HubSpot still sending raises both alerts, the backlog one an hour later.
cat > /tmp/hubspot-webhooks-nothing-processed-policy.json <<'EOF'
{
"displayName": "HubSpot webhooks: nothing processed in 1h (production)",
"combiner": "OR",
"severity": "ERROR",
"enabled": true,
"notificationChannels": ["projects/inboundx/notificationChannels/864254260479040154"],
"alertStrategy": {"autoClose": "604800s", "notificationPrompts": ["OPENED", "CLOSED"]},
"conditions": [{
"displayName": "No successful hubspot-crm-webhooks-production task in 1h",
"conditionThreshold": {
"filter": "resource.type = \"cloud_tasks_queue\" AND resource.labels.queue_id = \"hubspot-crm-webhooks-production\" AND metric.type = \"cloudtasks.googleapis.com/queue/task_attempt_count\" AND metric.labels.response_code = \"ok\"",
"aggregations": [{"alignmentPeriod": "3600s", "perSeriesAligner": "ALIGN_SUM"}],
"comparison": "COMPARISON_LT",
"thresholdValue": 0.5,
"duration": "60s",
"trigger": {"count": 1},
"evaluationMissingData": "EVALUATION_MISSING_DATA_ACTIVE"
}
}],
"documentation": {
"mimeType": "text/markdown",
"subject": "HubSpot webhooks: nothing processed in 1h (production)",
"content": "No HubSpot CRM webhook task succeeded in the last hour: client CRM data in Rose is no longer updated.\n\nCheck in order: queue `hubspot-crm-webhooks-production` state (paused?), worker responses on `/webhooks/tasks/process` (401 = OIDC, 500 = processing), intake requests on `/webhooks/crm` (none = HubSpot stopped sending; 401 = signature rejected).\n\nRunbook: docs/src/backend/hubspot-sync.md, section \"When a webhook alert fires\"."
}
}
EOF
gcloud alpha monitoring policies create --project=inboundx \
--policy-from-file=/tmp/hubspot-webhooks-nothing-processed-policy.json
cat > /tmp/hubspot-webhooks-backlog-policy.json <<'EOF'
{
"displayName": "HubSpot webhooks: backlog above 50 for 2h (production)",
"combiner": "OR",
"severity": "ERROR",
"enabled": true,
"notificationChannels": ["projects/inboundx/notificationChannels/864254260479040154"],
"alertStrategy": {"autoClose": "604800s", "notificationPrompts": ["OPENED", "CLOSED"]},
"conditions": [{
"displayName": "hubspot-crm-webhooks-production depth > 50 for 2h",
"conditionThreshold": {
"filter": "resource.type = \"cloud_tasks_queue\" AND resource.labels.queue_id = \"hubspot-crm-webhooks-production\" AND metric.type = \"cloudtasks.googleapis.com/queue/depth\"",
"aggregations": [{"alignmentPeriod": "300s", "perSeriesAligner": "ALIGN_MEAN"}],
"comparison": "COMPARISON_GT",
"thresholdValue": 50,
"duration": "7200s",
"trigger": {"count": 1}
}
}],
"documentation": {
"mimeType": "text/markdown",
"subject": "HubSpot webhooks: backlog above 50 for 2h (production)",
"content": "More than 50 HubSpot CRM webhook tasks have been queued for 2 hours: client CRM data in Rose is delayed.\n\nCheck worker responses on `/webhooks/tasks/process` (status >= 400) and the integrations API latency.\n\nRunbook: docs/src/backend/hubspot-sync.md, section \"When a webhook alert fires\"."
}
}
EOF
gcloud alpha monitoring policies create --project=inboundx \
--policy-from-file=/tmp/hubspot-webhooks-backlog-policy.json
Diagnose a site¶
Recent runs and the connection state (read-only):
select r.id, r.run_type, r.status, r.created_at, r.finished_at, r.error_message
from hubspot.crm_hubspot_sync_runs r
join hubspot.crm_hubspot_connections c on c.id = r.connection_id
where c.site_domain = '<domain>'
order by r.created_at desc
limit 10;
select status, last_sync_at, last_sync_status, last_sync_error
from hubspot.crm_hubspot_connections
where site_domain = '<domain>';
Match a run to its execution by start time:
gcloud run jobs executions list --job hubspot-contacts-sync-production \
--region europe-west9 --project inboundx --limit 20
Known errors:
| Error | Meaning |
|---|---|
job_terminated: … |
Cloud Run stopped the execution: 60-minute timeout or a cancellation from the console. |
stale_run_never_finished |
An earlier run never finished (container killed). Closed automatically; check that execution's logs. |
lock_not_acquired |
Another run for the same site was running when an operator-triggered run started. |
dense_window_exceeds_hubspot_10k_cap |
More than 10,000 records share one modification second (bulk import). HubSpot search cannot page past 10,000, so the initial sync cannot finish; webhooks still deliver later changes. |
Re-run a sync¶
Run from backend/. The job closes stale runs for that site, then runs an initial full
or incremental sync according to hubspot.crm_hubspot_sync_state.sync_mode.
Disconnecting and reconnecting HubSpot from the backoffice also starts a fresh
initial sync, but it deletes the site's synced CRM data first.