Your Data Stack Is Guessing — Make It Earn Truth
TL;DR: Most marketing data systems don’t understand people — they guess about them with confidence. From 515-page loyalty dossiers to venue surveillance risk scores, operators are building profiles that predict behavior without explaining it. Better data requires participation from the person being described, not just more fields and sharper inference.
More Fields Don’t Mean More Truth
A major data broker recently handed over a 132-field profile when asked. It looked precise. It was mostly wrong. That experience nails something the industry has avoided saying plainly: the volume of data is not a proxy for the quality of understanding. More fields create more places to be wrong.
The person described in that profile had no ability to review it, dispute it, or explain the context behind any of the behaviors that shaped it. A company assembled a version of them, sold access to that version, and used it to influence how other organizations treated them. That is not customer understanding. That is guessing with infrastructure.
For performance operators spending $10K or more per month on paid acquisition, this matters immediately. If your targeting is built on inferred signals layered over inferred signals, you are optimizing toward a model of your customer that may have little relationship to who they actually are. A full marketing audit will surface where inferred data is doing the work that disclosed data should be doing — and where your cost-per-acquisition is inflated because your targeting model is describing a ghost.
McDonald’s Built a 515-Page Dossier from Lunch Orders
A WIRED journalist requested the data McDonald’s collected through its loyalty program. What came back was 515 pages. McDonald’s had tracked years of purchases and used them to predict that this specific customer would visit 2.16 times in the following six weeks, spend an average of $13.49 per visit, and spend $29.15 total. It ranked his most likely orders and calculated his probability of churning.
The customer thought he was collecting points. The company was building a predictive behavioral asset.
McDonald’s privacy policy states it may combine customer-disclosed data with automated collection and third-party information to train algorithms, personalize experiences, support targeted advertising, and build profiles that include preferences, behavior, attitudes, and psychological trends. None of that is unusual for a large loyalty program. The problem is the imbalance: the customer sees a discount on a meal; the company sees a long-term asset it can monetize, score, and sell access to.
Operators running loyalty-style lead capture — common in iGaming acquisition and broker onboarding flows — need to look hard at this. If your lead capture is generating data you cannot act on transparently, you are accumulating liability, not intelligence.
Sweepstakes and Activations Are Data-Collection Campaigns with Prizes Attached
At Fanatics Fest NYC, activations from brands including the NFL, NHL, Honda, American Express, and Qatar Airways followed a consistent pattern. Visitors scanned a QR code for a chance to win tickets or merchandise. The entry forms requested full name, email, phone number, complete birth date, gender, country, and city — sometimes behind a third-party redirect before the form even appeared.
A company needs one reliable contact method to notify a winner. Every additional field is building a marketable identity, not running a contest. The fan sees a prize. The brand sees verified identity fields it can match against existing records.
This model is heavily used in crypto lead generation — airdrop campaigns, whitelist signups, and referral contests that collect wallets, emails, and device data while promising token rewards. The mechanic works. But operators who cannot explain to a participant exactly what is being collected, why, and how it will influence how they’re treated are sitting on a compliance risk that will scale with their audience.
Behavior Does Not Explain Motivation
Ken Beller, co-author of The Consistent Consumer, has spent years making a point the industry keeps ignoring: demographics and past behavior do not explain why people make decisions. Two customers buying the same product may be seeking entirely different outcomes — comfort, status, control, belonging, or relief. Traditional data records what was purchased and assumes the purchase explains the person. It doesn’t.
Research from Beller and J.D. Pincus using their MotivationMetrics methodology found that directly measuring emotional needs predicted consumer behavior two to four times more accurately than standard demographic and behavioral approaches. The implication for operators is direct: you are paying to target people based on what they did, not why they did it, and your conversion rates reflect that gap.
For Forex acquisition, this is acute. A trader who opened a demo account last quarter might have done so out of curiosity, competitive pressure, or a specific financial event in their life. Your retargeting campaign is treating all three the same way. Precise audience segmentation built on disclosed intent signals — surveys, declared preferences, onboarding questions — will outperform behavior-only models consistently.
Where Is the Person in the LUMAscape?
The LUMAscape has mapped adtech for years — data providers, identity companies, customer data platforms, publishers, measurement vendors, activation tools. It is a useful map of the machinery. What it does not make visible is the person whose information keeps that machinery running.
The consumer appears as an audience to identify, a profile to enrich, a customer to score, or an impression to monetize. Nearly every company on that map plays a role in collecting, connecting, interpreting, moving, or activating information about people. There is no category for the infrastructure that lets a person inspect what’s attached to them, correct errors, or participate in the economic value their data creates.
The MSG surveillance story — where a venue reportedly assembled a database of over 39,000 entries including risk scores, social media activity, personal associations, race, and sexual orientation — illustrates where the LUMAscape logic ends up when left unchecked. People were classified as risks for attending the wrong law firm or criticizing the venue’s owner. A separate breach reportedly exposed 10.5 million customer records. The customer thought they bought a concert ticket. The venue saw a face to scan, a record to retain, and a risk to classify.
Operators running paid media at scale need to understand that the ecosystem their budgets fund operates on the same principle. You are buying access to profiles built without the subject’s meaningful input. That is fine as a starting point. It is not fine as a strategy.
What This Means for High-CAC Vertical Operators
Forex, iGaming, crypto, and legal are the four verticals where cost-per-acquisition runs highest and where regulatory scrutiny of data practices is tightening fastest. The convergence is not accidental. High CAC pushes operators toward more aggressive data collection and inference. Regulators follow the money.
Three operational shifts are worth making now:
Distinguish disclosed from inferred. Know which fields in your CRM came from the person directly and which were scored by an algorithm. Build targeting logic that weights disclosed signals more heavily. Your AI-driven lead qualification stack should be capturing declared intent at the point of contact, not reconstructing it from behavioral proxies after the fact.
Build participation into your data model. Preference centers, onboarding surveys, and explicit consent flows are not compliance theater — they are the mechanism by which your data quality improves. A lead who told you they are actively comparing brokers is worth five leads who clicked a banner once.
Audit your inference layer. If you are running lookalike audiences, predictive lead scores, or third-party enrichment, map where those signals originate. For legal marketing campaigns targeting mass tort plaintiffs, inferred medical or financial signals carry specific legal risk that behavioral-click data does not carry on its own.
The McDonald’s dossier predicts your next order. The MSG system decides whether you get through the door. Both begin with ordinary activity. Both create a version of you that the company controls and you cannot inspect. Operators who build data systems that work the other way — where the person is a participant, not just a subject — will build acquisition funnels that convert better and survive regulatory scrutiny longer.
Originally reported by MarTech, August 2026.
Get a playbook for your vertical
Forex lead gen
FTD acquisition, depositor funnels, regulated broker campaigns across Tier 1 & Tier 2 GEOs.
Explore → CryptoCrypto & Web3
Token launches, exchange user acquisition, DeFi protocol growth. Compliant campaigns only.
Explore → LegalLaw firm marketing
Mass tort, personal injury, immigration. High-intent lead gen for US law firms with $50K+/mo budgets.
Explore →