Data Cleaning Automation in Google Sheets
Introduction
Data Cleaning Automation runs Flookup data cleaning operations in the background. You can schedule tasks at intervals ranging from every 15 minutes to once a day, with intervals up to 7 days.
If you only need a one-off lookup or spot check without scheduling, the Data Cleaning Spreadsheet Formulas may be simpler.
To open the scheduling sidebar, navigate to Extensions > Flookup Data Wrangler > Data Cleaning Automation in your Google Sheets menu.
Data Nova required. Data Cleaning Automation is available on every Data Nova plan and consumes no AI credits. See pricing for details.
How to Schedule a Function
- Select the function mode. Choose the operation you want to automate from the top dropdown. The form updates to show only the relevant options for that function.
- Choose the processing mode. Select whether to process data to the end once or loop continuously. See Processing Modes below.
- Configure the data ranges. Define your data sources. Highlight the range in your sheet and click the corresponding Grab selected range button to populate each field.
- Set the output position. Results are written starting from the active cell at the moment you click Schedule. Select the correct starting cell before scheduling.
- Adjust parameters. Set the threshold, column indexes and operation type. See Function Modes for details on each.
- Choose a frequency mode. Select how often the task should run. See Frequency Modes.
- Click Schedule. A status indicator at the top of the sidebar confirms the schedule has been created.
Tip: Before scheduling, run the function manually on a small sample of your data using Standard Data Cleaning to identify the best parameters.
Function Modes
Four function modes are available for scheduling. Each requires different inputs.
Fuzzy Match by Percentage
Performs percentage-based fuzzy lookups on a schedule. Each lookup value is compared against a table column and the best match is returned.
Required inputs: Lookup values range, Table values range, Lookup column, Return column and a Threshold.
Get Unique Values by Percentage
Extracts unique rows from a data range across scheduled runs by grouping near-duplicates using percentage similarity. Only one representative entry per group is returned.
Required inputs: Data range and a Threshold.
Standardize Text Entries
Cleans text on a schedule using the same operations available in Standard Data Cleaning: remove diacritics, filter stop words, strip punctuation, extract URL domains or paths or learn transformation patterns from examples.
Required inputs: Input range and an Operation type. For text and punctuation operations, provide a Stop array range. For pattern learning, provide dirty and clean example ranges. See Pattern Training for Standardize.
Compare String Similarity
Computes percentage similarity scores between two string columns on a scheduled interval.
Required inputs: Left range, Right range and Comparison mode (by word or by phrase).
Processing Modes
Processing mode determines what happens after a Minutes mode chain finishes a pass over your data.
Hourly and Daily mode schedules repeat the selected function at the configured interval and do not use the processing mode setting.
- Process data to the end: In Minutes mode, the chain processes every row once and then stops. The task is not automatically rescheduled after completion.
- Process data in a loop: In Minutes mode, the chain restarts from the beginning after each full pass. You must stop the chain when it is no longer needed.
Frequency Modes
Three frequency modes control how often your scheduled tasks execute.
- Run every few minutes (Minutes mode): Choose 15, 30, 45 or 60 minutes. Minutes mode processes chunks and schedules the next chain link. See How Auto-Chaining Works.
- Run every few hours (Hourly mode): Set the interval from 1 to 24 hours. The function runs at that interval and repeats the full processing pass each time.
- Run daily at specific time (Daily mode): Set a time of day and the number of days between runs. The interface permits intervals of 1 to 7 days, and the function repeats the full processing pass each time.
How Auto-Chaining Works
Minutes mode uses auto-chaining to handle datasets larger than one execution. Each link resumes from the last completed row, processes rows within the available time window, writes results and schedules the next link. Minutes mode is limited to a 15-minute minimum interval and a 60-minute maximum interval, with 15-minute increments. Each run uses a maximum processing window of about 350 seconds, subject to the remaining daily execution budget.
The chain can stop when all rows have been processed in Process data to the end mode, when you select Stop Chain, when the daily execution quota is exhausted, when Data Nova access becomes inactive or when a run cannot continue because of an input, sheet, service or processing error.
A stopped chain does not resume automatically. You must schedule it again with the required configuration.
How processing modes interact with auto-chaining:
Process to the end + Minutes mode:
The chain runs until every row has been processed once, then stops.
Loop + Minutes mode:
The chain restarts from the beginning after each full pass and continues until you stop it.
Pattern Training for Standardize
When scheduling a
Standardize Text Entries
function with the
pattern
operation, you can teach the system transformation rules by providing example pairs instead of choosing from predefined operations.
Select Learn from examples from the operation dropdown. Two additional fields appear: Dirty values range and Clean values range. These are your example pairs.
Click Test pattern to preview the detected transformation. The sidebar shows a table comparing each dirty value with its expected clean result. Matching rows appear in green; mismatches indicate you may need to provide more diverse examples.
For detailed guidance on preparing example pairs and best practices, see the Learn from Examples
Managing Schedules
- Stop Chain: Halts an active Minutes mode chain while preserving all results written so far. You can resume later by scheduling the same function mode again with the required parameters. The Stop Chain button is shown in the sidebar while a Minutes mode chain is active. Hourly and Daily schedules are removed with Reset.
- Reset: Completely removes a scheduled task and clears all saved progress. Use this when you want to start fresh or stop an Hourly or Daily schedule.
- Updating a schedule: There is no edit function. To change settings, first Reset the existing task, then create a new schedule with your updated parameters.
- Independent schedules: Each function mode maintains its own independent schedule. You can have a Fuzzy Match chain and a Standardize chain running simultaneously on the same sheet. The status bar in the sidebar shows information about the currently selected function mode.
Email Notifications
Minutes mode sends email notifications for key events.
Hourly and Daily mode do not provide the same chain completion notification.
- Task completed: Sent when a Minutes mode chain in "Process data to the end" mode finishes processing all rows. The email includes the number of rows processed and the function type.
- Task stopped: Sent when a Minutes mode chain is stopped manually, its Data Nova access becomes inactive, its daily quota is exhausted or a run reaches the error stop state. The email includes the reason and the number of rows processed so far.
Notifications use the email address associated with the Google account that created the schedule. If an email is missing, check the sidebar status and your email filters.
Processing, input, sheet or service errors can stop a chain without a notification if the failure occurs before notification details are available.
Important Notes
- Daily execution quota: All scheduled tasks share your Google account's daily execution limits. The current limits are 5,400 seconds for consumer accounts and 21,600 seconds for Workspace accounts. Each run leaves a 15-second safety buffer and processes for at most about 350 seconds. If the quota is exhausted, the chain stops and clears its saved processing state. You must schedule it again.
- Sheet renaming: Renaming a sheet after scheduling will not break an existing schedule. The system identifies your target sheet by more than just its name.
- Sheet deletion: Do not delete the target sheet after scheduling. If the sheet is deleted, the task will fail on its next run.
- Authorization: Schedules run under the authority of the user who created them. If you revoke the add-on's authorization, all your scheduled tasks will stop. Re-authorizing the add-on restores functionality.
- Closing the sidebar: The sidebar can be closed safely after scheduling. All processing happens in the background. Reopen the sidebar at any time to check status, stop a chain or reset a schedule.
Frequently Asked Questions
When should I use Data Cleaning Automation instead of Standard Data Cleaning?
Use Data Cleaning Automation for recurring jobs, large datasets and long-running operations.
Use Standard Data Cleaning for quick interactive processing where you want to see results immediately and stay in control.
Can I stop a running schedule without losing progress?
For a Minutes mode chain, click Stop Chain to halt processing while preserving results already written. Schedule the same function mode again with the required parameters to resume. Hourly and Daily schedules are removed with Reset.
How often can I run scheduled tasks?
You can schedule tasks from every 15 minutes up to daily, with a maximum interval of 7 days. More frequent schedules process data faster but consume your daily execution quota more quickly.
What happens during auto-chaining if the system encounters an error?
A processing, input, sheet or service error can stop the run.
The system does not promise automatic recovery or retry. Check the sidebar status and the sheet for the last written results, correct the cause and schedule the task again if it is no longer active. Minutes mode can preserve the results written before the stop.
Can I schedule multiple functions at the same time?
Yes. Each function mode runs independently. You can have Fuzzy Match, Standardize and Compare Similarity all running simultaneously on the same spreadsheet.