Status: live in production Source: Clay Target: HubSpot CRM Last updated: 2026-04-23

CRM Hygiene Pipeline

Config-driven pipeline that enriches a HubSpot portal from a Clay table, with audit trails, safety gates, and source-scoped writes that never overwrite records the pipeline did not create.

Overview

This pipeline keeps a HubSpot portal clean by syncing a prospect contact list from a Clay table. Every row is de-duplicated, normalized, gated by five safety checks, and written through a source-scoped upsert that refuses to overwrite records it did not originally create.

The integration is intentionally backend-only — there is no mapping UI. Every field transformation lives in a versioned YAML config in the repo, so changes are reviewable like any other code change.

Source

Clay.com

Prospect contact table, per-row HTTP Action.

Orchestration

Trigger.dev v3

Durable executor, per-row retries, rate-limited queue.

Staging

Supabase (Postgres)

Raw audit + normalized rows + run history.

Destination

HubSpot CRM v3/v4

Contacts, Companies, and associations.

System Architecture

Four components. Each one has a single responsibility; nothing owns two roles.

Clay
Source of rows
HTTP Action per row
→
Intake Endpoint
Vercel serverless route
POST /api/clay/crm-hygiene
→
Supabase
clay_import (raw)
sync_rows (normalized)
→
Trigger.dev
Orchestrator + per-row tasks
Rate-limited queue, retries, idempotency keys
→
HubSpot
Scoped upsert by email/domain
Contact ↔ Company association

Every write carries an Idempotency-Key derived from sha256(config_id | external_ref | action). Trigger.dev uses the same key to drop repeat runs, and every write is preceded by a lookup, so a re-sent row updates the existing record instead of creating a second one.

Data Flow — Six Steps

1

Intake

Clay's HTTP Action POSTs one JSON row per contact to the Vercel endpoint. Request is authenticated with a shared-secret header; bodies without the header are rejected before any DB write.

2

Raw audit

The exact Clay payload is inserted into clay_import untouched. This is the forensic record — every transformation downstream can be traced back to a specific raw row.

3

Normalization

Split derived fields (last name from full name), lowercase email and domain, strip whitespace, drop nulls. No HubSpot calls yet — this is pure data cleaning.

4

Staging upsert

Normalized row written to sync_rows, keyed on (config_id, external_ref) with external_ref = Final Email. Duplicate Clay submissions collapse into a single staged row.

5

Safety gate

Five mandatory checks (see below). Gate failures are recorded in sync_runs.counts_json.gate_skipped with a reason — never silently dropped.

6

Scoped HubSpot write

Per-row task looks up the target by unique key (email for contacts, domain for companies), branches on the data_source tag, and writes only to records the pipeline created. Associations are created only after a contact is created or updated, never for a skipped row.

Safety Gate — Five Checks

All five must pass before any HubSpot call. Any failure is logged to the run record with a reason string.

✓API credentials present
HubSpot Private App token resolvable from env or encrypted hubspot_credentials row.
✓Master sync enabled
Environment variable SYNC_MASTER_ENABLED=true. Single global kill-switch.
✓Per-config enabled
sync_configs.enabled=true for the target mapping. Disable a single config without touching the global switch.
✓Row ready to sync
sync_rows.ready_to_sync=true. Lets a reviewer hold specific rows back without deleting them.
✓Segment valid (if declared)
When the config declares a segment whitelist, rows outside the list are skipped. No-op today; in place for future list routing.

Property Mapping

Contact

HubSpot propertySourceNotes
emailFinal EmailUpsert key. Row skipped if empty.
firstnamefirstnameProvided directly by Clay.
lastnameFull NameDerived by stripping the firstname prefix.
jobtitleFinal Title—
companyFinal Company NameDenormalized string on the contact for quick search.
full_name customFull NameStored verbatim for audit.
final_rep customFinal RepsText in v1. Migrates to hubspot_owner_id via email lookup once rep emails are provided.
data_source customconstant crm_hygieneIdentifies records this pipeline created. Read during every re-sync.

Company

HubSpot propertySourceNotes
domainFinal DomainAssociation key.
nameFinal Company Name—
data_source customconstant crm_hygieneSame scoping guarantee as contacts.

Association

Contact → Company matched by Final Domain. Created with HubSpot's default association type. Uses the v4 associations API.

Scoped Upsert — Data Protection

Every HubSpot write is preceded by a lookup on the object's unique key. The write branches on the data_source tag:

Case A

No match

Create a new record. Stamp data_source=crm_hygiene.

Case B

Match with matching tag

Update properties. Tag stays intact. This is the normal re-sync path.

Case C

Match without tag

Do not write properties and do not create associations. Row marked status=skipped with error=scoped:untagged.

Case D

Match with different tag

Treated identically to Case C — another pipeline owns that record, so we do not touch it or its associations. Row marked status=skipped with error=scoped:wrong_tag:<value>.

The result: this pipeline cannot corrupt HubSpot data it did not create, even on a buggy mapping change or an accidental bulk re-send.

Edge Cases Handled

Status

Progress against the major milestones. Updated as each one lands.

Milestone 1

Core engine verified

All HubSpot actions working against a sandbox portal, with retries, rate limiting, and idempotent writes.

Milestone 2

End-to-end pipeline verified

Live intake endpoint → staging database → orchestrated write into HubSpot. Real payload tested; re-sending the same row did not create a duplicate.

Milestone 3

Safety pattern verified

The scoped-write guarantee that prevents overwriting records the pipeline did not create. All four branches (create, update, skip on wrong tag, skip on untagged) verified end-to-end against the sandbox portal.

Milestone 4

Production launch verified

Launched 2026-04-22. Live HubSpot portal wired, Clay connected, scheduled trigger (*/15 * * * *) running. A live contact synced end-to-end (Clay → Vercel intake → Supabase → Trigger.dev → HubSpot) in under 15 seconds from cron tick to CRM write.

All four milestones complete. The pipeline is deployed and runs on the scheduled cron; every run is recorded to sync_runs with queued / synced / skipped counts.