Clean the names first.
Remove formatting differences and common prefixes or suffixes before comparing names.
The Harbor Works, LLC
HARBOR WORKS INC.
harbor worksThe cleaned names match. Two different companies can still share a name.
Different systems. Different names. Five ways to compare the records.
Remove formatting differences and common prefixes or suffixes before comparing names.
The Harbor Works, LLC
HARBOR WORKS INC.
harbor worksThe cleaned names match. Two different companies can still share a name.
Compare unique words, then allow small spelling differences.
Harbor Works
Works Harbr Harbr
harbor → harbrOne missing letter.
Count each word once per record. Rarity depends on the dataset you’re comparing.
Harbor Services
“harbor” appears in
2of 1,000 records
Rare word · More weight
“services” appears in
620of 1,000 records
Common word · Less weight
Compare addresses, phones, and company email domains when names differ.
| Customer | CRM ID | Billing ID |
|---|---|---|
| Harbor Works | C-101 | B-501 |
| Cedar Health North | C-107 | B-107 |
Compare one system with another. Inspect the strongest candidates, resolve ambiguous pairs, and export the proposed mapping with its evidence.
Blocking narrows the pairs before scoring. A token-only block can lose a renamed account even when its phone and address agree. The union also checks qualified phones, addresses, business domains, and shared-namespace IDs. It is independent of the scoring toggles.
Use name similarity with enabled address, phone, business domain, and namespaced reference data. Missing values add no points. Shared values carry less weight.
Only whole phrases at the beginning or end are removed. Broader lists can erase distinguishing words. Separate entries with commas.
The example contains fictional records from CRM, Billing, and Support. Select two systems to begin.
Required columns: id, system, name. Optional: address, unit, city, region, postalCode, country, phone, domains, emails, referenceNamespace, referenceId. Separate multiple domains or emails with semicolons. Quoted CSV cells and multiline names are supported. IDs must be unique within each system.
JSON accepts an array of records, or an object with a records array. Use a structured address and reference: { "namespace": "customer-register", "value": "ORG-110" }. The input id is always a local system key; it is never treated as a cross-system identifier.
The example records and spreadsheet are fictional.
The matching tools compare records in your browser. The reference score is a hand-authored teaching rule: name similarity contributes up to 45 points; a specific address adds 25; a qualified phone adds 20; a business domain adds 20. Two strong reference fields add a 30-point corroboration bonus. A namespaced trusted ID adds 100. Conflicts and shared values have separate checks. The score is clamped to 0–100; it is not a calibrated probability.
Address cleanup uses a small set of US/Canadian street abbreviations and preserves units and countries. Phone cleanup is formatting and country qualification, not phone validation or proof of ownership. Consumer email providers are excluded from company-domain evidence. Domain comparison uses exact hostnames after removing only “www”; no public-suffix or corporate-ownership inference is made. Shared-value detection means more than two records in the current corpus, a deliberately simple warning rule.
The workbench limits imports to 80 records and 128 KB for interactive use. It proposes mappings and exports an auditable report; it does not update source systems or automatically cluster records. Its review decisions live only in the current page session.