Controlled beta for Google Ads advertisersRead the update
Use case · diagnose performance

Why did my ROAS drop? Ask it that. In those words.

Mako pulls your live Meta and Google Ads data at the moment you ask, measures it against the return your own gross margin requires, names the campaigns carrying the fall, and shows every read it made on the way. The bar it judges against is arithmetic on your account. We hold no benchmark data and we will not invent a threshold that depends on your business.

No card to connect. Read-only until you grant write access, per account.

Demo account · illustrative figures

Section 01 · what an answer looks like

Here is the answer shape. Reads on top, bar named, campaign named.

Two questions against the sample account, rendered the way the product renders them. Above each answer is the tool trace: which tool ran, the one-line summary it returned, and whether it succeeded. A read that failed shows as failed rather than being dropped so the answer can look cleaner.

YouROAS dropped last week. What happened?

Mako

Account purchase ROAS is 1.84x over the last 7 days and 2.31x over the last 14, so the fall sits in the back half of the fortnight.

Your bar is 2.5x, which is one divided by your 40% gross margin. I did not choose it and it is not a benchmark.

The audit's wasted-spend check fired, and this is its wording rather than mine: "USD 5,190 spent with zero purchases (12% of live spend)". 2 live sales campaigns spent USD 5,190 and recorded no purchases. The worst is "Prospecting-TOF" at USD 4,312. That campaign is past the four-day floor the check applies before it will judge anything.

Retargeting-30d is at 4.16x, above your bar.

Do you want the ad set breakdown on Prospecting-TOF, or should I propose a move?

NOTE · "the recovery band is our Q3 hero product"

YouIs Google doing the same thing?

Mako

No. The Google Ads audit did not raise the under-target finding this run, so nothing there is returning below your bar.

There is a different leak, and again this is the check's wording: "at least USD 612 of search spend on terms that converted nothing". 24 search terms took at least USD 612, about 12 percent of search spend, without a single conversion. The bar is set by your own account: this account converts about once every 11 clicks, so a term needs 33 clicks with nothing to show before it counts as dead, or the term cost more than 2 times what a conversion costs you. The search terms report was capped at the top spenders, so the real total is higher.

I am reporting Google separately from Meta on purpose. Two platforms with two attribution models do not add up to one number.

Product view · demo data · Gymszy sample account

Every number in that thread is a value a read returned, or a sentence a check wrote. None of it is the model doing arithmetic. Sable Mako records what a live read returned into a per-turn ledger, across a fixed list of metrics, and audits the answer against it: a figure claiming one of those metrics that the ledger cannot account for is held back rather than sent to you. The list is where it stops. Frequency, reach, sessions, average order value, conversion rate, customer counts and add-to-carts are not recorded, a column headed something we cannot map to a metric is not checked, and a table cell holding a sentence rather than a value is skipped rather than guessed at. That is why every figure in the example on this page is one a read returned or a sentence an audit check composed.

The percentages you can see are the audit's, not the model's. The share of live spend in the Meta finding and the share of search spend in the Google one are both computed by check code before any model reads them, which is why they can appear in an answer at all.

The figures are the sample account's. They are demo inputs. Every sentence around them is composed by the same formatter the product uses, so the money reads the way our code renders money rather than the way a copywriter would write it.

Section 02 · what it reads, and when

Every read says what kind of read it is. Live, stored history, or our own rules, and never dressed as each other.

A live read runs against your connected account at the moment of the question, and the result is stamped with the moment it was fetched before the model is allowed to see it. Live reads are the only thing that can answer what the account is doing right now. Five of these are not live data on purpose, and they are marked. Three serve the stored daily history our nightly sweep keeps, for trends and period comparisons, with the read dates carried on every payload. The playbook returns our own written rules and the creative analyzer scores an image, so neither may enter the ledger of numbers an answer is checked against.

Table A · every read tool the agent can call
What it readsToolPlatform
Spend, purchase ROAS, click-through rate, cost per click, cost per thousand impressions and purchase actions, at account, campaign, ad set or ad level, over one of eight named windows or any number of days from 1 to 365, each row carrying the identity of the thing it describes, optionally sliced by one breakdown: placement, device, age, gender or country.meta_get_insightsMeta Ads
Campaigns with status, objective, effective status and budgets.meta_get_campaignsMeta Ads
Account name, currency, account status and lifetime amount spent.meta_get_account_summaryMeta Ads
Ads with status and creative: name, thumbnail or image URL, Instagram permalink.meta_get_adsMeta Ads
Per campaign: spend, impressions, clicks, conversions, conversion value and cost per conversion, plus account totals, in the account's currency. Search campaigns also carry impression share and the share lost to budget against the share lost to rank, which demand opposite fixes.google_campaign_performanceGoogle Ads
Campaigns with status, channel type and daily budget.google_list_campaignsGoogle Ads
Ad groups with status, optionally scoped to one campaign.google_ad_groupsGoogle Ads
Keywords with match type, status and live spend, clicks, impressions and conversions.google_keywordsGoogle Ads
The search terms report, what people actually typed, with spend, clicks and conversions. Row capped, and a capped read says so in the payload.google_search_termsGoogle Ads
Responsive search ads: headlines, descriptions, final URLs, status and policy approval status.google_ads_creativesGoogle Ads
The conversion actions: name, status, type, category, and whether each is primary for bidding. The commonest silent misconfiguration on Google is a purchase action that exists but is not primary, so every campaign optimizes to something else.google_conversion_actionsGoogle Ads
Performance Max asset groups with status, optionally within one campaign. PMax campaigns hold asset groups instead of ad groups, so this is the structure view the ad-group read cannot give.google_asset_groupsGoogle Ads
Spend, impressions, clicks, conversions and TikTok's own attributed purchase value over a dated window, day by day for the account or one total per campaign, in the advertiser account's own currency. TikTok drops days with no data out of its response, so an absent day comes back absent rather than as a zero and the payload says how many days it covered.tiktok_performanceTikTok Ads
The advertiser account's name, currency, timezone and status. TikTok reports money in the account's own currency with no conversion, and its report days are that account's timezone rather than yours or the store's, so this is read before any TikTok figure is reasoned about.tiktok_account_summaryTikTok Ads
Sessions, key events, purchases and revenue as the site itself recorded them, broken down by traffic source and medium, day by day over a dated window. These are your own site's numbers and they will differ from what an ad platform claims for the same traffic. That difference is the reason to read both.ga4_traffic_sourcesGoogle Analytics
The same window rolled up by channel group: paid social, paid search, organic search, direct, email and referral. One channel can cover several platforms, so paid social fuses Meta and TikTok together, and the per-source read above is what separates them again.ga4_channel_summaryGoogle Analytics
What people typed into Google to find the site, by query, day by day: clicks, impressions, click-through rate and average position in Google's organic results. No ad platform and no GA4 tool can show this: it is search behaviour before a visit, not what happened after one.gsc_search_queriesSearch Console
The same window by page instead of by query: which pages Google shows in search, and how each one performs there.gsc_search_pagesSearch Console
Order count for a window, and a sampled revenue total with the sample size returned beside it. The read pages through your orders and reports whether it reached the end.shopify_get_orders_summaryShopify
Store name, domain, currency, plan and timezone.shopify_get_shopShopify
The number of products in the store.shopify_get_products_countShopify
Live Meta spend joined to live Shopify revenue over one window: blended MER, Meta's platform-reported return, and the gap between them. Google Ads spend is not in it, and the payload says so.blended_performanceMeta and Shopify
The full deterministic audit: every finding with the numbers behind it, a score out of 100, per-category scores, and the list of what could not be assessed.get_account_auditMeta Ads, Google Ads
A persisted optimiser run over live campaign performance: the one budget move the account's own numbers justify, with the source and destination campaigns, the amount, the figures each side was judged on, and the run id a budget proposal must then cite as the origin of its figure.run_budget_optimiserMeta Ads
How much stored daily history exists per account: the first stored day, settled and provisional day counts, and the longest window a period comparison can honestly cover.Stored historyhistory_overviewStored history
Day-by-day stored figures for one account: spend, impressions, clicks and conversions, or revenue and orders for the store. Missing days are named with the reason, and days the platform can still revise are marked provisional.Stored historyhistory_daily_seriesStored history
A window against the equal-length window before it, computed server-side: totals, blended MER and CAC, and percentage changes. When the two windows do not carry equal reported days, it compares per-day averages and says so; when a comparison would lie, it refuses and names the reason.Stored historyhistory_period_compareStored history
The differential diagnosis: every mechanical cause of a performance change, checked in the order a buyer rules them out, one verdict per cause. Data holes, tracking breaks, store-order reconciliation, spend steps, mix shifts, launches, CPM, CTR, CVR and AOV, with the ruled-out and cannot-assess lists returned beside the confirmed causes.Stored historydiagnose_changeStored history
Our written media-buying rules for a platform and topic. Marked as not live data so a rule's number can never be quoted back at you as something measured in your account.Not live dataget_playbookOur own rules
The ad creatives you have uploaded, newest first, with file name, dimensions, upload date and any score already given. Marked as not live data: a pixel dimension or a byte count is not a measured account metric and can never be quoted back at you as one. File names are your own text and are fenced as untrusted.Not live datalist_brand_assetsYour creative library
A scored creative review of one ad image, on a URL the model was handed by a read rather than one it composed. Marked as not live data, and its output is fenced as untrusted because on-image text is written by whoever made the ad.Not live dataanalyze_creativeAn ad image
  1. Period comparisons come from history_period_compare, which computes both windows server-side from the stored daily history and names the days each side covers. Stored figures are always labelled as stored history with their read dates; they never stand in for a live reading of the account.
  2. Meta reads take the eight named windows, today, yesterday, last 7 days, last 14 days, last 30 days, last 90 days, this month and last month, or any number of days from 1 to 365. Google Ads reads take yesterday, this month, or any number of days up to 365. The audit runs on last 7, 14, 30 or 90 days on either platform.
  3. Blended MER joins live Meta spend with Shopify revenue. The order revenue is a sample and the sample size is shown beside it. Where the read cannot reach the end of your orders, MER is withheld rather than computed on a partial total, because a confident understatement of your flagship number is worse than no number.
  4. Every Google Ads read is row capped. When a cap cuts a read short the payload carries a partial-data note, and a finding built on a capped read states its figure as a floor and says the real total is higher.

Section 03 · where every bar comes from

The bar is your margin. Not an industry average.

Most ads tools tell you a number is bad by comparing it to a benchmark that came out of somebody else's account. We do not have your competitors' data, and neither does anyone selling you a benchmark for your category. So every bar Mako judges against is either arithmetic on your own account or a fixed guard we print here with its value. There is no third kind.

Table B · every bar the audit and the agent judge against
The barWhere the number comes fromSet by
The return a campaign has to clearOne divided by your gross margin. At a 40% margin that is 2.5x.Your account
The same return, once your returns are measurableOne divided by your margin after returns. The ad platform counts revenue at the moment of purchase, so a 12% return rate on that same brand raises the real bar to 2.84x. It uses your Shopify data only with at least 21 days behind it, and a measured rate above 60% is refused as broken data rather than believed.Your account
The return, with no gross margin on fileBreakeven on ad cost, 1.0x. The report records that it judged on ad cost rather than profit, and the audit hands you the question instead of assuming a margin. Our own code calls 1.0x the losing line rather than a profit bar.Your answer, or an honest gap
What counts as an expensive Meta campaign2x the average cost per purchase across your converting campaigns only. Campaigns that bought nothing never dilute that average, and the campaign being judged needs 3 or more purchases behind it.Your account
What counts as a dead Google search term3x the clicks your account normally needs to buy one conversion, bounded 10 to 300. With no conversions to derive from it uses 25 and says in the finding that the bar is a default rather than yours. The other arm is 2x what a conversion costs you.Your account
Enough spend to judge a Google campaign at all2x your own cost per conversion, with a floor for an account too new to have one.Your account
Enough spend for a duplicate keyword to matter5x your own average cost per click, with a floor.Your account
What counts as a winner worth moving budget to2x your target return, with 3 or more purchases behind it.A multiple of your own bar
How old a campaign must be before it is judged4 days. A campaign that young is unstable by design and its numbers forecast nothing, so the audit refuses to call it a loser and says it is too young to read.Fixed, and named as fixed
How many purchases before a return figure is trusted3. One order can swing a return on ad spend far enough to invent a verdict.Fixed, and named as fixed
The frequency that raises a fatigue heads-up2.2 over 7 days, 3 over 14, 4 over 30, 6 beyond. Shipped as a heads-up to confirm against a falling click-through rate, never as a verdict, because we cannot tell prospecting from retargeting without ad set data and retargeting runs hotter.Fixed, and named as fixed
When one campaign holds too much of the account40% of live sales spend, raised at 60%, and only once 3 or more live sales campaigns exist.Fixed, and named as fixed
When under-target spend is worth a finding25% of live sales spend sitting under your target, critical at 50%. The target itself is yours; the share that raises the alarm is fixed.Fixed, and named as fixed
When wasted Google search spend is worth a finding10% of search spend, critical at 25%. A judgment about how much of a budget is leaking, which travels across business types in a way a performance number does not.Fixed, and named as fixed
When a budget is spread too thin to learn on8 or more active campaigns, with 5 or more of them each under 2% of active spend.Fixed, and named as fixed
The money floor before a finding is worth raisingUS$50, converted into your account's billing currency before it is applied. A flat fifty meant about half a US dollar to an Indian advertiser and about 162 to a Kuwaiti one.Fixed, and named as fixed

Every derived bar carries its own basis, and the basis is printed inside the finding. A finding says which bar it applied and how that bar was reached, so you can argue with the right thing. When a bar falls back to a default because your account cannot support the arithmetic, Mako says so and hands you the question instead of using the default in silence.

What you will never see here. An industry average, a category benchmark, a figure somebody else's best accounts produced, or a threshold we picked because it sounded about right for a store like yours. We hold no other advertiser's data, there is no comparison source anywhere in our code, and we will not invent a number that depends on your business.

Connect an ad account and the first audit runs against your own numbers, on the bars above. Nothing is written and no card is asked for.

With an active trial or plan, the same bars run on a schedule on the account audit.

Section 04 · what it can and cannot diagnose

What it can diagnose today, platform by platform.

The deterministic checks behind a diagnosis are eleven on a Meta ad account and eleven on a Google Ads account. Connect both and you get two audits, two histories and two alert streams, because the numbers are not comparable across platforms. There is no combined figure on this page for the same reason.

Table C · what a diagnosis reaches, and what it does not
PlatformWhat it can diagnoseWhat it cannot, today
Meta AdsEleven deterministic checks: account status, delivery problems, broken conversion tracking, wasted spend, under-target spend, spend concentration, reallocation opportunities, cost-per-purchase outliers, frequency fatigue, no active campaigns and fragmentation. Plus live insights at account, campaign, ad set and ad level over every window Meta offers. You can run this one on demand from the audit page.Audience overlap, placement breakdown and creative-level attribution modelling. None of those fields is read, so no check pretends to judge them.
Google AdsEleven deterministic checks: zero conversion tracking, wasted search terms, under-target campaigns, no active campaigns, self-competing duplicate keywords, campaigns past their ramp that bought nothing, profitable campaigns capped by their budget, profitable campaigns losing auctions to rank, a purchase action Google is not bidding toward, bid targets under the profit bar and mature value bidding with no bar at all. Plus live campaign performance, ad groups, keywords, the search terms report and responsive search ad text.Impression share, bidding strategy fit, ad strength, asset coverage, conversion action health, Performance Max brand cannibalisation and wasted campaign spend. Those need reads we have not built, and a check that guesses is worse than a check that says it could not look. The included audit can be run directly for an eligible Google Ads account; repeat audits and the daily watch require active service.
ShopifyOrder counts, a sampled order revenue total with its sample size, daily sales, discounts, returns and net sales. This is what raises your target return by your own measured return rate, and what the blended MER view is built on.There is no commerce audit. Shopify is read-only today: no write permission is requested and no write path exists in our code.
Google Analytics 4It connects, and we read your property list, then sessions, key events, purchases and revenue by traffic source, day by day, nightly and live in chat.No audit check reads GA4, and there is no write path. The site's own numbers are not an ad platform's attributed figures, and no page here blends the two.
Search ConsoleSearch queries and the pages they landed on, with clicks and impressions, day by day, nightly and live in chat.No audit check reads Search Console, and there is no write path. Sitemaps, URL inspection and indexing coverage are not read.
  • A lead-generation advertiser gets less of this than a store does. Every profit check assumes purchases and a gross margin, and campaigns that convert in a chat rather than on a website are excluded from the Meta sales universe on purpose. If you run messaging or form-fill campaigns you still get the account, tracking, creative and structure checks, and the return and wasted-spend checks will have little to judge. We would rather say that here than have you find it out after connecting.
  • The daily watch covers two platforms. Meta Ads and Google Ads are the platforms with an audit engine behind them, so they are the ones a scheduled run can read and judge when the connection and service entitlement are eligible. TikTok Ads, Shopify, GA4 and Search Console can feed a diagnosis in chat but are never audited on a schedule. There is no TikTok audit check anywhere in our code, and the table above is the whole of what TikTok gives you.
  • The prose is written by a model and the money parts are not. A model can misread a pattern, so the parts that touch money were made deterministic instead. Every finding, every severity and the score out of 100 are computed in our code, and no model authors any of them. Treat the prose as the lead and the findings as the record.

Every field we read from every platform, and the ones we have not built, are on the integrations page.

Section 05 · what every answer carries

Every answer carries the reads it was built from.

01

The tool trace, on every turn

Above each answer the app shows what ran: the tool, a one-line summary of what it returned, and whether it succeeded. A read that failed renders as failed. It is never dropped so the answer can look cleaner, because a diagnosis built on four reads when five were attempted is a different diagnosis.

02

Fresh by construction

Every live read stamps the moment it was fetched onto its own result before the model is given it, and the two read tools that are not live data are marked as not live so a written rule or a creative score can never enter the ledger an answer is checked against.

03

Evidence chips, when a diagnosis becomes a proposal

If the answer turns into a proposed change, the approval card carries the metric, the value in your account’s currency, the verified data source, the scope it applies to, the reporting window and the data freshness. A proposal citing a number that was not pulled from the platform that turn is blocked before the card exists, with strict matching, so a hundredfold unit error cannot pass as a rounding difference.

Approval card evidence · demo data
Ad spend
$4,312.00
Meta performanceProspecting-TOFlast 30 daysDemo · not a live read
Return on ad spend
0.00×
Meta performanceProspecting-TOFlast 30 daysDemo · not a live read
Purchases
0
Meta performanceProspecting-TOFlast 30 daysDemo · not a live read

New approval cards carry a server-verified receipt. The source, reporting window and freshness come from the tool call and result that actually grounded the figure, not from the model’s wording. A proposal citing the wrong source is blocked. Cards created before this receipt existed say that the source, window or timing was not recorded instead of guessing.

Section 06 · from a diagnosis to a change

Then you decide whether anything moves.

A diagnosis on its own changes nothing. Write access is off by default on every connection and has to be turned on for that specific account, behind a confirmation. While it is off, the tools that could change your account are never sent to the model, so the agent cannot ask for something it has not been handed.

With it on, Mako can propose fourteen things and no fifteenth. The bounded list covers three Meta campaign and budget operations; nine Google Ads operations for campaign status or name, budget, paused creation, campaign negatives, positive keywords, existing ad-group status or name, and exact existing-ad status; one Merchant Center feed publish; and one paused TikTok campaign creation. Google and TikTok campaign creation remain switched off. Every available operation is a card you approve on its own, showing the current value beside the proposed one. There is no bulk approve, and existing ad text, URLs and creative are not edited.

After a change runs, the platform is read again. The value that came back is compared against what you approved, and a disagreement is recorded as a failed verification rather than as a success.

Availability is narrower than the code, and this page says which is which. Google Ads reads and recommendations are available in the controlled beta. Its bounded write operations require approval of the exact action and remain under supervised validation. Meta is limited to app reviewers and explicitly allowlisted test accounts while Meta review completes. A diagnosis needs read access only, and every connection begins with writes off.

The full permission table, every change operation and what we never collect are on the security page, and the approval card itself, with every limit and every way a change can fail, is on the publish-from-chat page.

Section 07 · straight answers

Straight answers to the questions that decide it.

Does it compare my account to industry benchmarks?

No, and we could not if we wanted to. There is no benchmark table, no external comparison source and no competitor data anywhere in our code. Your target return comes from your gross margin and your own measured returns. The bar for an expensive campaign is a multiple of your own converting cost per purchase. The bar for a dead search term is a multiple of your own clicks per conversion. The numbers that genuinely are fixed are listed in table B with their values, and they are guards on judgment, such as a campaign having to be 4 days old and carry 3 purchases before anything is said about it, rather than performance bars.

What if I have not set a gross margin?

It still runs. The target falls back to breakeven on ad cost, 1.0x, the report records that it judged on ad cost rather than profit, and the audit hands you the question rather than assuming a margin. Breakeven on ad cost is the line where you stop losing money on media, not the line where you make any, and our code calls it the losing line for that reason. Showing you the gap beats filling it with a guess.

Will it change my campaigns while it is diagnosing?

No. A diagnosis uses read tools only. Write access is a separate switch, off by default on every connection, and while it is off the tools that could change your account are not sent to the model at all. Turning it on takes an explicit confirmation, and even then every individual change still waits for your approval on its own card.

How does it compare last week to the week before?

It pulls each window with its own live read and tells you which two it compared. There is no stored period-over-period table and no separate comparison tool, so the windows named in an answer are always the windows it actually fetched.

Can the answer still be wrong?

Yes, and here is exactly where. The prose is written by a model over live data, and a model can misread a pattern or do arithmetic on two figures it pulled. So the parts that touch money were made deterministic instead: every finding, its severity and the score out of 100 are computed in our code, with each finding's headline number written by code and not by a model. When a diagnosis turns into a proposed change, that proposal's cited numbers are checked against that turn's reads before a card can exist, and an unmatched figure blocks the proposal outright. The same check runs over the reply and holds back a figure it cannot trace, in prose and in a table, though it covers a fixed list of metrics, does not check a column headed something it cannot map to one, and skips a cell holding a sentence rather than a value, which is why we tell you to treat the prose as the lead and the findings and the approval card as the record.

Does this work for lead generation?

Partly, and less well than for a store. Every profit check assumes purchases and a gross margin, and messaging-objective campaigns are excluded from the Meta sales universe on purpose, because Meta files them under a sales objective while they record conversations rather than purchases. You still get the account, tracking, creative and structure checks and every live read. The return and wasted-spend checks will have little to judge.

What does a diagnosis cost?

Chat costs credits by depth: one for a quick answer, three for a standard one, nine for a deep dive. One initial Google Ads audit is included without a card. Repeat audits and the daily watch require an active trial or plan; scheduled runs draw no AI credits.

What does it get access to?

Only the accounts you connect, scoped to one workspace. Google Ads, GA4, Merchant Center and Search Console are the ordinary controlled-beta connections. Meta is reviewer-only, Shopify accepts existing installs only, and TikTok Ads is a supervised test path. Google Ads changes remain bounded and require approval of the exact action.

Controlled beta

Bring one active account and inspect the evidence with us.

Apply for the Google Ads beta. Invited partners connect read-only and review the first diagnosis before the agent can change anything.

  • official Google OAuth for the ordinary beta
  • no benchmark data, ours or anyone's
  • every figure in the worked thread: one demo account

Invited partners connect read-only first. Write access is separate and exact approval remains mandatory.

Receipts

  1. We hold no benchmark or competitor data. There is no comparison table, no external benchmark source and no competitor connector anywhere in our code, which is why every bar on this page is either your account's arithmetic or a guard we print with its value.
  2. Every figure in the worked thread is a demo-account input. The sentences around them are composed by the same formatters the product uses, and the finding text is the wording our check code writes.
  3. Every live read stamps the moment it was fetched onto its result before the model is given it, and the two read tools that are not live data carry that flag so a written rule cannot be quoted back as measured performance.
  4. Your target return on ad spend is one divided by your gross margin, raised by your own measured return rate once your store has enough data behind it. A measured rate above the credible ceiling is refused rather than used.
  5. The money floor is written in US dollars and converted into your account's billing currency before it is applied.
  6. A check that cannot judge this run reports itself unavailable with a reason in plain language, and skipped checks are listed in the report rather than dropped.
  7. A trace row carries the tool, its one-line summary and whether it succeeded. An evidence chip carries the metric, the value in your account's currency, the tool that fetched it and the entity it applies to. Neither carries the fetch time today.
  8. Write access is off on every connection until you turn it on for that account, and every individual change waits for your approval on its own card. After it runs, the platform is read again and a disagreement is recorded as a failed verification.