data-agent-rules

1

data-agent-rules plugin for Cursor

1 rule

Add to Cursor
# Data agent rules Rules for AI coding agents that touch data: SQL, Spark, dbt, Databricks, Snowflake, BigQuery, DuckDB. Code mistakes show up in review. Data mistakes show up on the bill, in a dashboard, or as missing rows, often days later. ## 1. Look before you scan Read the schema and a small sample before any full read. Use `LIMIT`, `TABLESAMPLE`, `.limit(n)`, or a single partition. Before a heavy query, check the table's size (`DESCRIBE DETAIL`, `INFORMATION_SCHEMA`, row count on one partition) and say what the query will cost. ## 2. Never guess names, and never invent values Look up table and column names (`DESCRIBE`, `INFORMATION_SCHEMA`, the dbt manifest) before you use them. If a column or value doesn't exist, say so. Do not fill a missing value with a plausible default: `NULL`, "unknown", and "not in source" are data. Do not make up thresholds, tiers, categories, or mappings the data doesn't contain. If the user asks for a field that doesn't exist, build the rest, leave that field `NULL`, and ask for the source or the rule. A made-up default looks exactly like a real value to everyone downstream. ## 3. Writes go to dev unless the user names the target By default, write to a scratch or dev schema. Ask before writing to anything named `prod`, `main`, `gold`, or `analytics`, or to any table you did not create in this session, unless the user's request names that table. ## 4. No destructive statement without a count and a yes Before `DROP`, `TRUNCATE`, `DELETE`, `UPDATE`, `CREATE OR REPLACE` on an existing table, `INSERT OVERWRITE`, or `mode("overwrite")`: run a `SELECT COUNT(*)` with the same predicate, show the number, name the undo (time travel, `RESTORE`, a backup table), and wait for confirmation. A `DELETE` or `UPDATE` with no `WHERE` is almost always a bug. A keyed `MERGE` or insert into a dev or staging table the user named is a normal load, not a destructive statement: check the counts and run it. ## 5. Make every write safe to run twice Use `MERGE` on a declared key, and de-duplicate the source first. Use `replaceWhere` or a partition overwrite instead of a full overwrite. A blind `INSERT` or `append` inside a job that can retry duplicates rows. ## 6. Keep data off the driver and out of the chat Never call `.collect()` or `.toPandas()` on an unbounded DataFrame; aggregate or limit first. Don't put raw personal data (emails, names, phones, addresses) in your answer, not even one example: show counts, patterns, and masked samples like `u***@example.com`. Never put credentials in code or notebooks; use secret scopes or environment variables. ## 7. Push the work down Filter on partition or clustering columns, select only the columns you need, prefer built-in functions over Python UDFs, and never write an accidental cross join. For heavy queries, read the plan (`EXPLAIN`) before running. ## 8. Prove the change Every transformation change ships with a check: a dbt test, an assertion, or before/after row counts, null counts, and key uniqueness on the affected table. Report the actual numbers, not "looks good". For dbt, run `dbt build --select state:modified+`, and never run `--full-refresh` without asking.