Home / Methodology

Methodology

How Rascasse measures audiences

AI-based analysis of aggregated digital behaviour signals, not surveys. How we combine sources, estimate demographics, check results and where the method has limits.

In one paragraphRascasse measures what people do, not what they say. We analyse aggregated digital behaviour signals with AI: what people search, follow, watch and attend. From them we build audience profiles in 200+ countries, down to postal code where available, updated quarterly with up to ten years of history. No surveys, no cookies, no personal data.

Our approach

Rascasse sits between the two established ways of studying audiences. We are not a survey provider, and we are not a social listening tool. We observe digital behaviour across several independent sources and build audience profiles from what people do.

The reason is the say–do gap: what people report about themselves often differs from what they actually do. Choi and Varian showed that search behaviour contains information about real economic activity, available earlier than official statistics.[1] Kosinski, Stillwell and Graepel showed that digital records of behaviour predict personal attributes with high accuracy.[2]

The research industry discusses the same problem. At IIeX North America 2025, Qrious Insights presented findings suggesting error rates of around 80% in self-reported media consumption.[3] Rep Data's State of Survey Fraud 2025 analysed 4.1 billion survey attempts and classed 33% as fraudulent and 27% as inattentive.[4] Surveys remain the right tool for attitudes. For behaviour, observed signals avoid recall and self-report.

Three ways to study an audience

Survey researchAsks

Asks people what they think, buy and watch. Measures attitudes; subject to recall, social desirability and falling response rates.

Behavioural audience intelligenceObserves · Rascasse

Observes what people do across several digital sources and combines the signals.

Social listeningMonitors

Follows conversations on social channels. Shows what the vocal part of an audience says on a given channel.

The ICC/ESOMAR International Code (5th edition, 2025) recognises the role of the researcher as data curator: deriving insight from existing data rather than from direct contact with participants.[5] In the risk framework for synthetic and imputed data that Heineken presented at ESOMAR Reimagine 2025, this kind of method sits in the lowest-risk step, data imputation, because it draws inferences from real behaviour signals rather than generating synthetic respondents.[6]

Principles

  • Several sources, none dominant. Each profile combines independent sources. No single source decides the result.
  • Aggregated and non-personal. We work with patterns at population level. No individual is tracked or profiled.
  • Behaviour over stated preference. Searches, follows, views and attendance, not answers to questions.
  • Open about uncertainty. Where data is thin, we say so: estimates are marked as modelled, or as insufficient data, rather than shown as measured.

Data sources

Signals come from several independent categories. Each shows a different side of behaviour. Combining independent data streams to produce estimates no single source can deliver is the principle of data fusion described by Ipsos MediaCT.[7]

CategoryWhat we useRole
SearchQuery volumes, seasonal patterns, regional distributionDemand and intent
Social channelsFollowers, engagement, content interactionInterest
Video and streamingViews, channel subscriptions, listeningConsumption
Public recordsTV ratings, charts, award databases, WikipediaValidation
Published researchPublished studies, census data, national statisticsCalibration
PlacesPoints of interest from open geographic databasesLocation

The method does not depend on third-party cookies or browser tracking, and it does not depend on a single provider's interface: if one source changes its access, the others carry the profile. The ESOMAR Guideline on Passive Data Collection, Observation and Recording sets transparency and proportionality as conditions for research on observed data; our design follows both.[8]

Objects

Everything we measure is an object: a brand, a person, a club, an artist, an event, a media title or a topic that leaves measurable digital behaviour. Among them are 50,000+ brands and 100,000+ media channels; together they describe audiences with 500,000+ data points.

Each object is defined by a curated set of keywords, aliases and categories. This matters because one search term can mean different things: "Jaguar" the car brand is not "Jaguar" the animal. Telling them apart needs subject knowledge checked by algorithms.

For each object we calculate a size that combines search and social engagement, so objects can be compared across categories and countries, and a quality score from the agreement between independent sources. A low score triggers review. A new brand can be added in 2 hours.

Audiences

An audience is defined by the objects that describe it. The simplest audience is one object: fans of a club, listeners of an artist, customers of a brand. More complex audiences combine objects with AND, OR and NOT, for example sustainable fashion brands and media, without fast-fashion brands.

Components are weighted by how strongly they describe the audience: for hip-hop fans, artists carry more weight than media titles. Each audience draws on 100,000+ data points. A new audience you define is live in 24 hours, without a new questionnaire or fieldwork. See audience profiling and segmentation and personas.

Demographic modelling

Age, gender and other demographics cannot be read directly from search data. We estimate them from several independent indicators and combine them. Each indicator is a piece of evidence; the profile comes from where they agree.

  1. Channel composition. Each social channel has a documented audience structure; we calibrate against Pew Research Center's studies of social media use.[9] How strong an object is on each channel informs its profile.
  2. Transfer from known audiences. When creators with a known audience profile show affinity to a brand, part of that evidence is transferred by Bayesian updating: the brand's prior profile plus the creator's audience gives a revised estimate.
  3. Published research as calibration. Published studies, national statistics, TV ratings with known age distributions and category data serve as fixed reference points.
  4. Regional patterns. Regions have known demographic structures. When a brand is searched disproportionately in university towns, that points to a younger audience. Bayesian updating combines national priors with these regional patterns.

The Bayesian framework behind indicators 2 and 4 follows established methods in marketing science, as described by Rossi, Allenby and McCulloch,[10] and applied to media mix modelling by Google Research.[11]

Demographic estimates carry uncertainty. Where signals are thin, we flag the estimate rather than present it as measured.

Affinity and psychographics

Affinity measures how strongly an audience is connected to a brand, person or property compared with the market average. The baseline is 1.0. Above 1.0 means above-average interest, below 1.0 below-average. This index form, common in media planning, makes objects and audiences comparable. See affinity index. The calculation uses techniques from collaborative filtering and matrix factorisation described by Koren, Bell and Volinsky:[12] patterns of co-occurrence across behaviour reveal preferences that single objects do not show.

Psychographic traits such as sustainability orientation, technology adoption, luxury affinity or health consciousness are scored through marker objects: brands, people and properties that are strong indicators of a trait. The score shows how much an audience over- or under-indexes on these markers. The approach draws on research predicting traits from digital behaviour,[2] on the Schwartz theory of basic values,[13] and on Boyd et al., who showed that value orientations can be inferred from digital behaviour.[14]

All scores are relative to the market average. A sustainability score of 1.4 means 40% above the market average, not "highly sustainable" in absolute terms.

Location

Location results come from the regional distribution of behaviour signals, combined with points of interest from open geographic databases: venues, shops, cultural institutions and sports facilities. Results go down to region, city and postal code, where available, with up to 10,000 data points per postal code.

Regional affinity measures how strongly a brand or property resonates in one place compared with the national average, building on the regional analysis described by Choi and Varian.[1] Not every local spike is real: a city with unusually high affinity is checked against neighbouring cities and its region. Isolated spikes without regional support are treated as possible artefacts, not reported as findings. See location planning and local targeting and out-of-home.

Share of search is a brand's share of all branded searches in a defined competitive set. Les Binet proposed it formally at IPA EffWorks Global in 2020.[15] The IPA Think Tank, led by James Hankins, analysed 30 studies across 12 categories and 7 countries and found that share of search represents around 83% of market share,[16] and that changes in share of search tend to come before changes in market share.

  • Your competitive set. Defined per question, not taken from a fixed taxonomy. "Premium cars" in Germany can differ from the same idea in the United States.
  • Quarterly, with history. Updated quarterly, with up to ten years of history (4 to 10+ years depending on the market), so trends can be separated from seasonal noise.
  • Each market on its own. Every country is analysed separately, because competition differs by country.
  • Normalised. Volumes are adjusted for seasonality and for growth or decline of the whole category.

See share of search and market share.

Validation

  • Agreement between sources. A result is reported with high confidence only when independent sources agree. Single-source results are flagged.
  • Stability over time. Results are checked across the time series. A sudden shift without support from other signals triggers review, not automatic reporting.
  • Benchmarks against public data. Profiles are compared with external figures where they exist: TV ratings, published sales figures, census demographics and published research.
  • Quality score. Each object carries a score for the consistency and breadth of its data. Low scores are flagged in the Rascasse Solution.

Privacy

We process aggregated, non-personal signals. No individual is identified, tracked or profiled. The method is GDPR compliant, uses no cookies and needs no CRM or first-party data from you, so there is no data integration before work starts and no mixing of your customer data with outside sources.

What we do not do

  • No surveys. We do not ask anyone anything. There is no panel and no questionnaire.
  • No cookies. No tracking pixels, no browser-level tracking, no third-party cookies.
  • No personal data. No names, no e-mail addresses, no individual profiles. Only aggregated patterns.
  • No CRM data needed. You do not need to share customer or first-party data. If you want, your own sales data can sit next to our results.
  • No synthetic respondents. We do not simulate people. Every estimate goes back to observed behaviour.

Limits

A method is only as credible as its stated limits. These are ours.

  • Attitudes are not measured. We see what people do, not what they think or why they say they do it. Stated motivation, perception and purchase intent need a survey. See the comparison with YouGov and GWI.
  • Small audiences and small postal codes are less reliable. The fewer signals behind a result, the wider the uncertainty. Read results for very small audiences as direction, not as a precise figure.
  • Coverage varies by market. Not every source is equally available everywhere. Where dominant services restrict public data, fewer sources contribute and uncertainty grows. History also varies, from 4 to 10+ years depending on the market.
  • Online populations. The data reflects people who are active online. Groups with little digital presence can be under-represented, and we do not extrapolate to offline populations without saying so.
  • Demographics are inferred. They are estimated, not observed, and are more reliable for objects with distinct channel patterns.
  • Quarterly, not live. Data is updated quarterly. This favours stability and checking over speed. For day-to-day conversation, combine Rascasse with a monitoring tool.

Sources

  1. Choi, H. & Varian, H. (2012). Predicting the Present with Google Trends. Economic Record, 88(s1), 2–9.
  2. Kosinski, M., Stillwell, D. & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805.
  3. Moffatt, A. (2025). Rebuilding Trust in Research: Behavioral Data as the Foundation. Presented at IIeX North America / Qrious Insights.
  4. Snell, S. (2025). State of Survey Fraud 2025. Rep Data. Analysis of 4.1 billion survey attempts.
  5. ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics, 5th Edition (2025).
  6. Costella, T. / Heineken (2025). Synthetic Data Risk Framework. Presented at ESOMAR Reimagine 2025.
  7. Ipsos MediaCT (2011). Data Fusion: A White Paper. Ipsos.
  8. ESOMAR Guideline on Passive Data Collection, Observation and Recording. ESOMAR.
  9. Pew Research Center (2025). Americans' Social Media Use. Pew Research Center, Washington, D.C.
  10. Rossi, P.E., Allenby, G.M. & McCulloch, R. (2005). Bayesian Statistics and Marketing. Wiley.
  11. Google Research (2017). Bayesian Methods for Media Mix Modeling. Google AI Blog.
  12. Koren, Y., Bell, R. & Volinsky, C. (2009). Matrix Factorization Techniques for Recommender Systems. IEEE Computer, 42(8), 30–37.
  13. Schwartz, S.H. (1992). Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries. Advances in Experimental Social Psychology, 25, 1–65.
  14. Boyd, R. et al. (2015). Values in Words: Using Language to Evaluate and Understand Personal Values. Proceedings of the International AAAI Conference on Web and Social Media (ICWSM).
  15. Binet, L. (2020). Share of Search as a Predictive Measure. Presented at IPA EffWorks Global.
  16. IPA Think Tank / Hankins, J. (2021). Share of Search represents 83% of market share. IPA Think Tank analysis of 30 studies, 12 categories, 7 countries.

Questions

About the method

Is this survey data?

No. Rascasse uses AI-based analysis of aggregated digital behaviour signals, not surveys. Nobody is asked anything.

How often is the data updated?

Quarterly, with up to ten years of history (4 to 10+ years depending on the market).

Which countries are covered?

200+ countries with one method, down to postal code where available. Coverage and history vary by market.

Is it GDPR compliant?

Yes. The data is aggregated and non-personal, uses no cookies and needs no CRM or first-party data.

What can the method not tell me?

Attitudes and stated motivation. For those you need a survey. Small audiences and small postal codes are also less reliable.

Check it

Test the method on your question

A Proof Case answers one question in one market for up to three audiences, in three weeks at a fixed price, presented to the person who decides on the budget. Or talk to us for 45 minutes about sources, validation and limits.

Know your audience.
Act before others do.

Bring one question. Tell us what you need to decide, and we prepare a 45-minute session with your own data. Use the Rascasse Solution yourself or get the answer as an Insights Report.

  • 45 minutes, online or in person
  • Prepared with your teams, markets and partners
  • A one-page summary afterwards
What are you interested in? (optional)
Audience Intelligence
Markets & Locations
Brand & Market Share
Trends & Content
Media & Activation
Talent & Creators
Sponsorship & Partnerships
Live, Events & Tickets