Key Takeaways
- Dirty customer data puts one buyer in two segments through name variants and duplicate rows.
- Standardise name plus phone plus code columns before any segment rule runs.
- Fuzzy dedupe at 0.85 with hand review of borderline pairs keeps masters trustworthy.
- Flookup runs standardisation plus matching inside Sheets so exports stay in one place.
Why Segments Fail on Dirty Customer Data
Audience segmentation promises the right message for each group, yet most segments inherit every flaw in the customer database. One record writes Robert Williams while another writes Bob Williams with different email capitalisation, so one person lands in two segments and receives two versions of one campaign.
Data cleansing for segmentation means fixing identity before slicing groups. Standardise the customer data first, remove fuzzy duplicates second and split segments last. Teams that reverse the order rebuild every segment after each cleanup.
Standardise Names Phones and Codes First
Start with text standardisation across name plus phone plus code columns. Lowercase each name plus collapse whitespace plus strip punctuation, then normalise diacritics so spelling variants compare as equal. Standardise phone numbers to one layout and product codes to one pattern with Learn from Examples, as shown in the customer name standardisation workflow .
Paste formulas as values with Ctrl plus Shift plus V so later sorts read plain text rather than live formulas. Keep the raw export on a dated copy so every transformation stays reversible.
Fuzzy Dedupe Before You Split Audience Segments
Run Smart Deduplicate on the standardised name column before any segment rule touches the data. A threshold around 0.85 groups spelling variants for review while keeping distinct people apart. Approve high confidence groups and inspect the borderline band row by row, since a wrong merge pollutes every segment that follows.
Keep one master row per person with source plus consent date plus segment tag beside the address. Future imports then match against a clean reference instead of creating fresh duplicates.
A Worked Example With Fuzzy Match
A fitness brand holds 2,000 rows with name in column A, email in column B and last purchase in column C. Standardisation collapses 60 casing variants while fuzzy matching at 0.85 flags 90 near-duplicate pairs for review in Flookup Data Wrangler. After merging, 1,870 unique customers remain with a note beside each merged row, ready for segment rules on purchase recency.
Send-Ready Segment Checklist
| Check | Action | Target |
|---|---|---|
| Backup | Copy tab with date in name | Original untouched |
| Standardise | Lowercase plus TRIM then paste as values | Variants compare as equal |
| Dedupe | Exact pass then fuzzy pass at 0.85 | One master row per person |
| Segments | Split only after dedupe | Zero cross-segment duplicates |
| Refresh | Monthly check on active lists | New imports cleaned first |
Clean Segments Target Better
Clean customer data makes audience segmentation honest. Each segment holds distinct people, reports count each buyer once and suppressions work as intended. Run the sheet pass on every import and refresh segments that sat unused for 90 days or more. For the wider workflow, see the marketing data cleaning workflow and the CRM data cleaning workflow.
Frequently Asked Questions
How do I clean customer data before building marketing segments?
Standardise name plus phone plus code columns first so variants compare as equal, then run Smart Deduplicate at 0.85 and merge confirmed groups. Split audience segments only after the customer database holds one master row per person.
Why does audience segmentation fail on dirty customer data?
Dirty data places one person in two segments through name variants plus casing differences plus duplicate rows. Reports then double count buyers and suppressions miss, so campaigns send repeats and personalisation reads as error.
What is a good similarity threshold for customer names?
Start at 0.85 for personal names and company names in most lists. Raise the value toward 0.90 when false positives appear and lower it toward 0.80 when expected matches stay hidden. Spot check 200 to 300 rows before locking the rule.
How often should marketing segments be refreshed?
Clean every import before it joins the master and run a full deduplication pass monthly for active lists. Refresh any segment that sat unused for 90 days or more, since contact data decays as people change jobs and addresses.