WEBVTT

00:00:10.000 --> 00:00:16.000
Hello everyone, welcome to the Insights Association Town Hall. We're very glad you're here.

00:00:16.000 --> 00:00:26.000
I see you all coming on in for this really important town hall here as we close the year. As you know, this is Melanie, your CEO of the Insights Association.

00:00:26.000 --> 00:00:41.000
Really glad to see you. I'm super excited about this conversation today. As all of you know, I am a world-class nerd and I am a lover of all things geeky data and so I'm really excited about this town hall data innovation and exploring the promises and pitfalls of synthetic data.

00:00:41.000 --> 00:00:56.000
As you know, Inside Association is all about innovation, evolution, so we're super bullish on the possibilities of AI generative AI synthetic data in our profession in the coming years.

00:00:56.000 --> 00:01:04.000
But we also like to talk about, you know, what should we be thinking about? What should what should we be caring about as we go through this journey and inspiring you?

00:01:04.000 --> 00:01:13.000
Our hope is to inspire you today. So we're gonna wait just about a minute or 2 more, before we get started and let people continue to join.

00:01:13.000 --> 00:01:19.000
If you haven't done so yet, this is a great time to turn off everything around you so you can sort of focus and just give yourself this hour, give yourself permission to focus on this topic.

00:01:19.000 --> 00:01:30.000
For one whole hour, turn off the world around you. You can also bring up your chat window and your Q&A widget.

00:01:30.000 --> 00:01:36.000
Get ready for those. If you have a specific question, we'd love for you to keep those for the speakers.

00:01:36.000 --> 00:01:42.000
We'd love for you to keep those in the Q&A widget because as you know, they can get buried in the chat window pretty fast.

00:01:42.000 --> 00:01:50.000
But we definitely want you to be really active in that chat window. Say hello to everybody. Thanks Chris for saying hello from Portland.

00:01:50.000 --> 00:01:57.000
Tell us where you are if you want to share your LinkedIn link and make some new connections on LinkedIn, feel free to do that as well.

00:01:57.000 --> 00:02:06.000
Hi, Marnie, for Minneapolis. Glad you're here too. I, as you know, I am in the, I'm in the Dallas area.

00:02:06.000 --> 00:02:17.000
A quick note to some, I see using the host and panelist if you want to say hello to everyone, make sure to change that little drop-down where it says host and panelists.

00:02:17.000 --> 00:02:26.000
The little blue drop-down change it to everyone so that everyone can see you saying hello. That's a that's a that's a good trick there.

00:02:26.000 --> 00:02:33.000
Especially if you want to share your LinkedIn link, make sure to choose everyone so that everyone can see your LinkedIn link.

00:02:33.000 --> 00:02:42.000
We are. Just a couple one more minute and then we're gonna get started.

00:02:42.000 --> 00:02:47.000
Again, thank you all. Very much for being here. It's going to be a great session.

00:02:47.000 --> 00:02:52.000
We are, look at that. I think that we have our winner on distance. Hello from Nairobi, Kenya.

00:02:52.000 --> 00:02:53.000
It's great to see you, D'artagnan. Thank you very much for being here.

00:02:53.000 --> 00:03:03.000
I would love to be with you in Carolina Beach, North Carolina, Katie, so glad you are here.

00:03:03.000 --> 00:03:05.000
I would love to be with you in Carolina Beach, North Carolina, Katie, so glad you're here.

00:03:05.000 --> 00:03:06.000
David has his LinkedIn. Looks like he's looking to make some connections.

00:03:06.000 --> 00:03:15.000
From southern New Hampshire, so feel free to people to click on that and make a connection in LinkedIn.

00:03:15.000 --> 00:03:21.000
Montclair, New Jersey via Boston, very cool. Hello from Dallas, Beth. I know Beth.

00:03:21.000 --> 00:03:27.000
Glad you're glad to see you in my chat window, Beth. Great to see you in the area.

00:03:27.000 --> 00:03:33.000
All right, well, it is time to get started. As I said before, I'm super excited about this session today.

00:03:33.000 --> 00:03:39.000
A couple of quick housekeeping points. Just a quick disclaimer, we're going to be talking about synthetic data.

00:03:39.000 --> 00:03:45.000
We're going to be talking about some of the pitfalls. We're going to be talking about some of the legal aspects that you might want to be thinking about.

00:03:45.000 --> 00:03:59.000
This is for your information and for your inspiration. If you need specific advice about a specific incident of AI, about a specific question, and you need a specific legal, tactical, professional advice.

00:03:59.000 --> 00:04:06.000
This won't, this won't replace that if you need to be connected with someone through Insights Association, feel free to send us a note and we'll give you that.

00:04:06.000 --> 00:04:07.000
Space legal advice, but this session won't, replace that specific situational advice.

00:04:07.000 --> 00:04:25.000
That you may need for something. Also, we will be recording the session. We will make the recording and the transcriptions available to you on our town hall library on our website.

00:04:25.000 --> 00:04:26.000
Will, will that will automatically be sent to you on the other side of the session. So thank you all very much for being here.

00:04:26.000 --> 00:04:39.000
Now let me just get started. We have a really exciting panel for you today. Our first guest speaker is Benjamin.

00:04:39.000 --> 00:05:09.000
I don't think I honored his name quite as beautiful as he can say it. So maybe when he pops on, if he wants to correct me, he can do so.

00:05:58.000 --> 00:05:59.000
Okay.

00:05:59.000 --> 00:06:29.000
I don't think I honored his name quite as beautiful as he could say it. So maybe when he pops on, if he wants to correct me, he can do so.

00:07:52.000 --> 00:08:00.000
Are as a profession with AI synthetic data is is still the story is still unfolding. So why don't we start there for everyone if you don't mind and maybe I'll start with Damien and Zack.

00:08:00.000 --> 00:08:10.000
How do you define synthetic data? What's what is it? What what are we creating?

00:08:10.000 --> 00:08:14.000
And what are the different types of data? What is synthetic data?

00:08:14.000 --> 00:08:20.000
I'll start. I think that's a really good question because I don't know that there is a consensus.

00:08:20.000 --> 00:08:28.000
Some people say synthetic data and it's just machine learning. Some people say it and they mean taking training data and having.

00:08:28.000 --> 00:08:40.000
Almost a generative version of what scenarios could be in research and so for the most part what seems to be the through line though is data that's not directly gathered from respondents.

00:08:40.000 --> 00:08:51.000
That seems to be the most general consistent definition. Once you get beyond that, it starts to get a little bit fuzzy and murky, I think.

00:08:51.000 --> 00:08:54.000
Yeah, I agree. Zach and Chris.

00:08:54.000 --> 00:09:09.000
Yeah, I mean, I agree with the general. The general definition that is it's it's functionally made up not from respondents I think how we operate, we consider synthetic data.

00:09:09.000 --> 00:09:22.000
Almost exclusively in like a tabular or structured format. We don't, we don't operate on the assumption right now that anything.

00:09:22.000 --> 00:09:36.000
That's unstructured or that's coming from, like a language model. Or image generation, for us, at least today, that doesn't fit the definition of synthetic data because we're not sure.

00:09:36.000 --> 00:09:49.000
We can trust it to be accurately representative of. The, that we're trying to, you know, garner insights from, I suppose.

00:09:49.000 --> 00:10:02.000
Yeah, that's interesting. Cause I've heard the same sort of, thought process, synthetic data can be anything from statistics because you didn't gather them directly and a stat can now be considered synthetic data.

00:10:02.000 --> 00:10:05.000
Yeah.

00:10:05.000 --> 00:10:06.000
Hmm.

00:10:06.000 --> 00:10:19.000
And I was like, no, no, no, let's not go that far. Let's, let's give it a box, to this concept that synthetic data is an actual data set that's, that's, that is derived from an actual data set that's, that's, that is derived from an underlying data source, but it's actually that is derived from an underlying data source, but it's actually a structured data

00:10:19.000 --> 00:10:33.000
set. But it's actually a structured data set. Chris, you're not in your head, we think.

00:10:33.000 --> 00:10:34.000
Okay.

00:10:34.000 --> 00:10:37.000
Yeah, I mean, I think that, you know, what people are hearing here is we're still trying to get our arms around this and I'm pretty sure that in a year or two's time when we come come together again we'll be saying remember when we were talking about it a synthetic paper and now we're talking about this.

00:10:37.000 --> 00:10:54.000
I think that we're going to see the same kind of evolution as we go through. I mean, I think that Damien's concept of data that doesn't come directly from a respondent, certainly from the point of view of, you know, traditional research.

00:10:54.000 --> 00:11:04.000
That's kind of how a lot of this is being thought about. But then to Zachary's point, anything can be synthetic when we move the next.

00:11:04.000 --> 00:11:11.000
To the next step. So the key thing here is, is that there are new places we can get data from.

00:11:11.000 --> 00:11:19.000
And it can be directly from respondents that it can be from the models. Or it can be this interesting thing where we bring some real data to the models and say to the models, now give us some more of that.

00:11:19.000 --> 00:11:27.000
And so I think we're going to see many different types of synthetic data. I think we're going to see a lot of excitement about it.

00:11:27.000 --> 00:11:35.000
And I think we're going to see a lot of abuses of it as well.

00:11:35.000 --> 00:11:40.000
Yeah, I'm just gonna build upon, Chris point there, obviously working from the kind of panel side of things.

00:11:40.000 --> 00:11:48.000
We are definitely defining synthetic data as data sets that are generated by models that represent the behaviors of those individual respondents.

00:11:48.000 --> 00:11:54.000
And those kind of data sets, you know, created by modeling the response patterns across a range of kind of previous survey types or topic areas.

00:11:54.000 --> 00:12:03.000
We've kind of narrowed down our focus a little bit because I love your point of and then build upon that.

00:12:03.000 --> 00:12:11.000
Perfect. Ben, anything you wanna add before we start talking about where we sing it used?

00:12:11.000 --> 00:12:12.000
Okay.

00:12:12.000 --> 00:12:14.000
I think we've covered it pretty well. I'll hold my comments for some of our further questions.

00:12:14.000 --> 00:12:15.000
Yeah.

00:12:15.000 --> 00:12:16.000
Because

00:12:16.000 --> 00:12:22.000
Well, we do have one question. As we're sort of defining it that I see pop up in the chat, how do you get data from models?

00:12:22.000 --> 00:12:27.000
Can you give an example? What does that look like? How do you how do you get data from a model?

00:12:27.000 --> 00:12:28.000
Chris, this sounds like a question for you. Okay.

00:12:28.000 --> 00:12:45.000
Yeah, it's actually, you can, you can ask a model, you can say, hey, give, give me the responses of 15 respondents, Gen Z respondents talking about such and such.

00:12:45.000 --> 00:13:08.000
I'm going to ask them these 5 questions. Give me the their answers. And you can even say with the more modern models you can say things like, give that to me in Jason form which is a form we can use in programming.

00:13:08.000 --> 00:13:09.000
Okay.

00:13:09.000 --> 00:13:13.000
It's trivially easy. To generate something that plausibly looks like data. And I'm going to stress the plausibly part because these models are trained to be plausible.

00:13:13.000 --> 00:13:20.000
Okay, so generating plausible data is very easy. And in fact, we've been using this to blind studies.

00:13:20.000 --> 00:13:27.000
So when we, you know, that thing where you have to show a new potential client, some work you've done.

00:13:27.000 --> 00:13:38.000
We've been using synthetic data like that to actually generate very plausible data that can be shared and we're not giving away any confidence from another client.

00:13:38.000 --> 00:13:53.000
So generating plausible data. Directly from the model is really easy. Almost trivially, and I think that's going to be part of the dangerous we see because I think a lot of people are going to be part of the dangerous we see because I think a lot of people are going to mistake that plausible data.

00:13:53.000 --> 00:14:05.000
For real data and say hey I can generate a thousand respondents and it's trivially easy and I think that's one of the areas where we have to be very careful what we're doing.

00:14:05.000 --> 00:14:06.000
I will.

00:14:06.000 --> 00:14:23.000
We've actually talked about that in another session where you have to be really careful to as we're creating new kinds of data now to tag, you know, real primary data from created data so that in the future even the model itself knows what the real actual data versus the data it generated.

00:14:23.000 --> 00:14:31.000
It can separate them in case you want to clean them or say, please only use real data or, you know, sorry, I think I interrupted someone.

00:14:31.000 --> 00:14:45.000
I was just gonna add as an extension to Chris comment about triviality that I think Even now from the news, we're used to hearing about huge models that require a vast amount of computing power to.

00:14:45.000 --> 00:14:52.000
To operate like chat GPT, but. With a shift towards open source models that you can pull off.

00:14:52.000 --> 00:15:01.000
Anyone can pull down from GitHub and put into Google. Lab and run. We're, also talking about something that you can literally do from laptop at your house that the, the triviality aspect of it.

00:15:01.000 --> 00:15:11.000
Has. So precipitously changed over the last 2 or 3 years that this wasn't true in recent memory that you could just do this at home.

00:15:11.000 --> 00:15:24.000
And now it is and that's kind of concerning. It's kind of cool, but it's kind of concerning too.

00:15:24.000 --> 00:15:25.000
Right.

00:15:25.000 --> 00:15:28.000
Well, especially given the plausibility of it, right? It's so believable. And so it gives you the ability now to manipulate data to be what you want.

00:15:28.000 --> 00:15:34.000
People do that already, but now it's much easier and it seems plausible. But by the same token, it actually gives you a tool that could be great for something like scenario planning, right?

00:15:34.000 --> 00:15:46.000
Because you want to know what those plausible scenarios are. This isn't the real world, but you're looking at, well, what could potentially happen and when you're planning, that's great.

00:15:46.000 --> 00:15:54.000
So you want it to be plausible. So I think what we'll find in this whole conversation is that there are benefits and there are challenges and both sides, right?

00:15:54.000 --> 00:15:58.000
And sometimes equally sometimes one outweighs the other or more, but it's it's by no means I think a cut and dry and simple.

00:15:58.000 --> 00:16:05.000
This is going to be some dedicated and it's going to set us free or this is going to be synthetic and it's doing.

00:16:05.000 --> 00:16:06.000
Annihilation or something.

00:16:06.000 --> 00:16:07.000
Okay.

00:16:07.000 --> 00:16:13.000
Well, there's a, there's a series of questions here that are all related and, to this, to this, what is synthetic data?

00:16:13.000 --> 00:16:16.000
And they are, is there a connection to predictive modeling? What is, did Bayesian utilities in the choice model start the synthetic data?

00:16:16.000 --> 00:16:24.000
Did synthetic data start with imputation and inference and multiple imputation has been around forever. So I guess the question is, I think the answer is probably yes.

00:16:24.000 --> 00:16:32.000
Some of that is like inferred and imputed data is all synthetic data. But how are we moving into a world that's different?

00:16:32.000 --> 00:16:41.000
Is it that we're now creating entire synthetic data sets where it used to just be in the film, just be to fill it.

00:16:41.000 --> 00:16:43.000
How was it different?

00:16:43.000 --> 00:16:56.000
But before we really dive into that, I do want to split a hair here and say, I do think that that was the start of this, but most importantly perhaps is that that was the start of us being primed to accept.

00:16:56.000 --> 00:17:03.000
Sort of. The conclusions that are inexplicable to a general person looking at something, right?

00:17:03.000 --> 00:17:07.000
Where that Bayesian analysis is obviously very technical that the average reader is not going to necessarily grasp the whole process.

00:17:07.000 --> 00:17:22.000
So that that is definitely the start of priming us to say, well, I don't really know how it works, but it seems to, so let's go with it.

00:17:22.000 --> 00:17:27.000
And then certainly at that point the statisticians know what's happening in the background. Now we're transitioning.

00:17:27.000 --> 00:17:37.000
The average reader is still on the same point, but we're moving towards the statisticians can't go in and say, well, this is how it all works.

00:17:37.000 --> 00:17:47.000
Yeah. What else though? How is, so how is, how is this synthetic data wave different from the things that we've done in the past.

00:17:47.000 --> 00:17:48.000
Or is it?

00:17:48.000 --> 00:17:50.000
I don't know that it is. I think it's a continuation of what we've done in the past, right?

00:17:50.000 --> 00:18:03.000
But there These models are pretty much more complex versions of like spectral kings regressions or whatever sorts of analyses we've done before and now we're using those to.

00:18:03.000 --> 00:18:11.000
And a larger scale to predict plausible outcomes. So it's the same process. And I said, just didn't, you know, just bigger chunks in theory, but.

00:18:11.000 --> 00:18:19.000
How it's being used. I think what makes it different is who is using it. And their understanding of the underlying models and.

00:18:19.000 --> 00:18:29.000
Data that goes into that model. And so you have a person now on their laptop who doesn't have a background on statistics thinking, look, I just did a survey and I didn't have to get any respondents.

00:18:29.000 --> 00:18:34.000
This is great. Now we know what to do versus having someone who's sat back and tried to figure out, oh wow, okay, wait, our squares are, you know, 0 point 5 at the best.

00:18:34.000 --> 00:18:51.000
That means this is not a very strong correlation, but this is a potential outcome, right? So I think the biggest difference now is the access to something that not everyone has a skill set to understand or to navigate well.

00:18:51.000 --> 00:18:55.000
Fair enough. So Go ahead, sorry.

00:18:55.000 --> 00:18:59.000
So I was just gonna key off saying that, Damien said there because I, I agree totally. It's, it's easy access.

00:18:59.000 --> 00:19:09.000
And the fact and the fact that it's plausible when people don't really understand where it comes from and I think we could all Be better educated on how these models work and should be.

00:19:09.000 --> 00:19:20.000
There is this The is this kind of circular thinking and I've had this discussion over beer with friends.

00:19:20.000 --> 00:19:21.000
Hmm.

00:19:21.000 --> 00:19:32.000
Where people have said, you know, now all we need to do is I can create a thousand respondents and I can, you know, and now I can, don't have to go to speak to a real person.

00:19:32.000 --> 00:19:33.000
Yeah.

00:19:33.000 --> 00:19:44.000
I just create a thousand respondents. And then you say Well, what would you do then? And then they say, well, then I analyze the data from the data I find the story of what's going on and when you say So why didn't you just ask that question of the model in the first place?

00:19:44.000 --> 00:20:00.000
Because you're not going to get anything different because the data came from the model. So rather than going through the whole process of analyzing these 1,000 respondents and saying, Gen Z really likes this brand, why not just ask the model, what brand it thinks Gen Z likes.

00:20:00.000 --> 00:20:13.000
And so there's a little bit of circular thinking going on where people have said, well I always run a survey so this is how I run my surveys now and I think that's quite dangerous circular thinking at the moment.

00:20:13.000 --> 00:20:22.000
Yeah, I think when it comes to this, synthetic panel, so it's a great phrase I, was reading in a New York Times article recently about the fear of kind of stochastic parrots.

00:20:22.000 --> 00:20:37.000
So I've heard you can say plausible data, but that's going to stick in my mind, but the kind of stochastic parrot is what I fear from the most kind of from synthetic audiences in particular so Chris maybe you're right it's not about creating a synthetic audience to then question but question the model instead.

00:20:37.000 --> 00:20:54.000
I, I do think that's a super important point though on this, the casting parent one, cause what we're gonna find is Whereas real surveys of the general population are going to capture the most up to date trends.

00:20:54.000 --> 00:21:04.000
These synthetic. Generative AI driven. Results are gonna be limited by.

00:21:04.000 --> 00:21:20.000
Historical information, right? So like when Chat GPT first came out, it was like bench book ended at like end of 2020 or whatever right and so no information you would have gotten up to date since then, 4.5 they say it's open to the whole internet but There's gotta be limits, right?

00:21:20.000 --> 00:21:28.000
For cost reasons and compute reasons and stuff like that. So, Yeah, and most people don't.

00:21:28.000 --> 00:21:40.000
Understand what those limitations are or are cognizant about those as they're using these tools. And so it's a double edged sword and that it introduces productivity and new capabilities, but then.

00:21:40.000 --> 00:21:50.000
How much of it is real and up to date and are you capturing trends or are you just looking at a different way of framing historical information.

00:21:50.000 --> 00:21:53.000
Yeah, it's that, same question that we as really good researchers have been asking ourselves for decades.

00:21:53.000 --> 00:22:05.000
Is this the right? Fit for this particular decision. Right? Is this, is this data the right data to use to generate the answer or do I need to add to it or do I need a whole new data set?

00:22:05.000 --> 00:22:13.000
As you know, so I mean, and so John Bremer is asking that questions and you guys are you're around the topic.

00:22:13.000 --> 00:22:22.000
What are the guardrails? Like what do we think it's good for? What do we, what are the guardrails when thinking about synthetic data.

00:22:22.000 --> 00:22:28.000
Well, for me personally, I'd like it for scenario planning. It's helpful in that sense, right?

00:22:28.000 --> 00:22:32.000
For things that aren't going to change much. So if I'm looking at a business that's pretty.

00:22:32.000 --> 00:22:37.000
Static throughout time. It has this range of motion and they're trying to plan for next year.

00:22:37.000 --> 00:22:41.000
Well, let's figure out what could be the possible rangees for that, right? Hi, mid low.

00:22:41.000 --> 00:22:52.000
Okay, then we could probably fall within these ranges. If it's for some of my other clients where they need to know what's happening today, what trends are going on, or how is culture shifting in the last 2 months.

00:22:52.000 --> 00:23:06.000
Probably not the place where I'm going to use it. I'm going to go with real people and surveys and trending conversations on social media, whatever is more accurate, because those things change so quickly that by the time the model picks it up, it's potentially already.

00:23:06.000 --> 00:23:10.000
Past and we're on to something new.

00:23:10.000 --> 00:23:23.000
One of the things that you mentioned, guardrail's and we've in in in humanate we've put in place internally and we've in in humanate we've put in place internally a number of very specific guardrails that people have to follow in this area, a number of very specific guardrails that people have to follow in this area.

00:23:23.000 --> 00:23:28.000
And one of them is that if you use an LM or a generative AI for anything you have to follow in this area.

00:23:28.000 --> 00:23:32.000
And one of them is that if you use an LM or a generative AI for anything, you have to both, LM or a generative AI for anything, you have to both let the client know and the team know.

00:23:32.000 --> 00:23:50.000
How it's been used and I think that's a guardrail that everybody should have you know it's it's transparency these things are incredibly powerful and they're very useful they are transforming the industry we know that and they're making our, when they're making us more powerful as researchers I believe that very strongly.

00:23:50.000 --> 00:23:54.000
However, we do need those guardrails and one of them is this transparency. You know, we should be saying we have used the large language model in the process of creating this.

00:23:54.000 --> 00:24:07.000
We that should be very clear to other team members. And also to our clients where appropriate.

00:24:07.000 --> 00:24:08.000
In addition to the, oh, I'm sorry, go ahead.

00:24:08.000 --> 00:24:20.000
Thank you. Okay, this is a really important topic that we need to start creating the guide rails, because what we're finding with our clients who work with overcome 500 enterprise clients is that because there are no guardrails.

00:24:20.000 --> 00:24:24.000
They are just not allowing us to use it. So it's actually not the researcher who is very keen to start testing this out, testing those models, etc.

00:24:24.000 --> 00:24:29.000
But she's their legal teams who are basically stating if there's no guardrails, then let's just say no for now.

00:24:29.000 --> 00:24:47.000
Because we're not sure. So I think it's, you know, it's on us in our industry to start creating really come strict guardrails to help allay the fears of any of those kind of legal teams that we're working with so that those enterprise researchers can start start to work with the data.

00:24:47.000 --> 00:24:52.000
And I think one of the things that's really helpful in I spoke about this on another panel is.

00:24:52.000 --> 00:25:01.000
Transparency in terms of what models and what data was used to build those models, right? Because One of the big questions that keeps coming up is, oh, is there a bias in it somehow?

00:25:01.000 --> 00:25:04.000
But if it's a black box and you don't know what data went into it, you you can't.

00:25:04.000 --> 00:25:05.000
Hmm.

00:25:05.000 --> 00:25:10.000
Same that there is or there isn't. Some people say, oh, it's, it's way less biased, but.

00:25:10.000 --> 00:25:16.000
Again, what data went into it? Because to be honest with our traditional methods, we still haven't managed to mitigate bias, right?

00:25:16.000 --> 00:25:23.000
We still have some in there inherently and so how can we say that this based on those things that we're already doing is going to not have bias?

00:25:23.000 --> 00:25:31.000
But at least having transparency into what data fed that model and how it was built, you can at least understand where that bias may exist and how you can then.

00:25:31.000 --> 00:25:35.000
Work with in the confines of that or correct for it.

00:25:35.000 --> 00:25:45.000
Yeah, hopefully most of you saw that the Insights Association just published this updated code of standards. I've put a link to it in the chat and I've put the new section on section 4 data and technology.

00:25:45.000 --> 00:25:53.000
Nothing's really new in the code. It just says AI and data and technology has to use the same rules of standards as everything else that we do.

00:25:53.000 --> 00:26:03.000
Transparency, data governance, fit for purpose, quality, duty of care. It's all wrapped up in the work that we all the work we do including AI.

00:26:03.000 --> 00:26:13.000
So when we use AI, we had we do have to be transparent according to the code. So if you're code abiding company member and a code abiding member, individual member, you're already responsible for transparency for making sure that you can explain the data and where the data came from.

00:26:13.000 --> 00:26:36.000
Make sure you can explain that it fits. So, so there are a lot of elements in it and then, you know, lots of lots of us are also working hard on a very soon delivery of an actual AI standards and best practice.

00:26:36.000 --> 00:26:37.000
Yeah.

00:26:37.000 --> 00:26:42.000
So, so it's a really good conversation, but already you're governed both in law in some places and in our in our own self-governance of some guardrails.

00:26:42.000 --> 00:26:43.000
But what about like you scarred rails? Like where's it, where's it good?

00:26:43.000 --> 00:26:52.000
Where's it not good? Where would you say don't use this yet maybe.

00:26:52.000 --> 00:27:09.000
I think that ties in to a validation question, right? Where are we talking about a question that a reasonably diligent researcher with a reasonably large amount of data could Given time figure out and are we making that a faster process.

00:27:09.000 --> 00:27:29.000
Is it something that, that is believable that that that is static in a way where it's not just credible, but it's actually You can look at and say, I understand how this conclusion was arrived at, even if we're not talking about in the most technical sense, right?

00:27:29.000 --> 00:27:36.000
If you come to me and say we want to, you know, test how this new safety feature in a car is going to work.

00:27:36.000 --> 00:27:44.000
And the model says, well, people with young children are gonna like it. That makes sense given the data that we have.

00:27:44.000 --> 00:27:50.000
If it comes back with something or we're asking a question where they're isn't a solid grounding.

00:27:50.000 --> 00:28:01.000
In existing data, it's still gonna give you an answer. You're just gonna need to trust it last right if you have a lot of data points about the group of us here.

00:28:01.000 --> 00:28:07.000
And then ask us about a topic. You have no data points on. The model is still going to give you an answer.

00:28:07.000 --> 00:28:13.000
It's just Not, you shouldn't trust it. Where's the line on that though?

00:28:13.000 --> 00:28:15.000
I don't really know.

00:28:15.000 --> 00:28:31.000
Well, that's, I mean, that's really interesting in the sense that, you know, one of the things that I think people need to move, move towards is bringing their data to the models, not not just relying on the model, not just asking the model.

00:28:31.000 --> 00:28:39.000
Hey, what does a, you think about such and such? But instead, bringing to the model, we have all this data, we've got tons and tons of data.

00:28:39.000 --> 00:28:48.000
Tell us what our data says about this. And this is this is one of the areas where, you know, I know there's been a lot of developments in in the technology with some of the approaches to working with.

00:28:48.000 --> 00:28:57.000
Bringing data to the model. But I think that that's something that we're going to be doing more and more of.

00:28:57.000 --> 00:29:03.000
You asked Melanie about, you know, where, where a synthetic data being used well.

00:29:03.000 --> 00:29:12.000
One of the areas we're using it very successfully at the moment is in synthetic personas.

00:29:12.000 --> 00:29:13.000
Yeah.

00:29:13.000 --> 00:29:18.000
Personas have always been made up. Let's be very clear about this. We've done a segmentation, we said this is this is the such and such segment.

00:29:18.000 --> 00:29:20.000
They do this, this and this. And then we tell a story to bring that segment to life. We tell a story about a person.

00:29:20.000 --> 00:29:39.000
Well, you know, one of the things we're doing is building synthetic personas where you can actually have a chat with one of those people where you can not just say, hey, this is Becky the soccer mom and she does XY, and C, you can actually have a discussion with Beckett.

00:29:39.000 --> 00:29:47.000
You can actually say what about this but that's based on our data that's not based on just generally asking the model what a soccer moms like.

00:29:47.000 --> 00:29:53.000
It's about saying this is our data, this is this is what's trained the model and we're now having a discussion with that data.

00:29:53.000 --> 00:30:06.000
So this idea of bringing the data to the model is very, very important and there are lots of new techniques around that where we can actually bring our own data in.

00:30:06.000 --> 00:30:14.000
And then use the large language model to allow us to bring that to life to expand on it.

00:30:14.000 --> 00:30:17.000
Okay.

00:30:17.000 --> 00:30:18.000
Yeah.

00:30:18.000 --> 00:30:19.000
Yeah, because you've developed those personas as synthetic personas and they wouldn't apply to other data sets.

00:30:19.000 --> 00:30:39.000
They wouldn't be tagged the same. They would be categorized the same. So you'd probably get different answer if you asked chat GPT to tell you what soccer moms think versus your data set which and you and and that piece of being able to say I can explain the data because it's my data versus I can't explain the data because it's this open source data is is an important point, right?

00:30:39.000 --> 00:30:41.000
Sorry, Katie, go ahead.

00:30:41.000 --> 00:30:48.000
Yes, I think that's a great use case, Chris and to build up on kind of your use case with one that probably should not apply is if we were to ask that synthetic soccer mom for example.

00:30:48.000 --> 00:30:58.000
About kind of disruptive innovation, so new flavor trends, etc. Probably unlikely to answer. I was recently purchasing.

00:30:58.000 --> 00:31:16.000
Cold brew, caramel flavored M and M's and of course Col Bruno would even heard of maybe a year or 2 ago so it'd be tricky to then question that synthetic soccer mom on new flavor trends or something that's maybe you know in the cultural like guys today and maybe in the future we'll get there but I think today that may be where we start to fall down on any kind

00:31:16.000 --> 00:31:20.000
of disruptive innovation testing.

00:31:20.000 --> 00:31:22.000
Yeah. Fair. Damien, were you going to add something?

00:31:22.000 --> 00:31:27.000
No, I, I think both Chris and Katie got it.

00:31:27.000 --> 00:31:34.000
What does bias look like? I mean, how should people be thinking about bias? It's a big important topic and and think to yourselves about the audience that we have.

00:31:34.000 --> 00:31:45.000
Some of them have their own data sets. And so how do you think about bias? Would in a known data set and how should they be thinking about bias in a large unknown unexplainable data set.

00:31:45.000 --> 00:31:46.000
Yeah.

00:31:46.000 --> 00:31:52.000
I think for for me and in the interactions that I've had and the times when we were looking to employ a couple of agencies that did almost exclusively synthetic data.

00:31:52.000 --> 00:32:02.000
The biggest issue is that we couldn't identify the bias. We couldn't identify the data that went into the model.

00:32:02.000 --> 00:32:11.000
So without that, we were running blind. And so we we'd be likely to. Make decisions, make correlations that.

00:32:11.000 --> 00:32:24.000
May or may not have existed previously. I think that's one of the problems is that we have a lot of sort of fly by night people who've heard of this new thing AI and synthetic get it so they want to jump in but they're not actually.

00:32:24.000 --> 00:32:35.000
Necessarily the rigor to it or sharing the insights behind it and sometimes they may not even have it they may be using an open source model so they have no way of understanding what went into building it.

00:32:35.000 --> 00:32:48.000
And so. We find ourselves in this place of, well, is it amplifying a bias that could potentially lead us down a path that's not beneficial to us or the audience that we're trying to reach or not.

00:32:48.000 --> 00:32:52.000
That's where I think Chris's point of bringing your own data sets in and making sure that you're.

00:32:52.000 --> 00:33:09.000
Intimately familiar with how how the sample was collected, who is in this data set, what were the parameters around those respondents being put into it is so crucial because then we can at least understand well yeah we know we were short in this population or we how to bias towards X.

00:33:09.000 --> 00:33:12.000
And when we see that in the subthetic data, we can speak to, hey, but we should take that with a grain of salt because of these biases.

00:33:12.000 --> 00:33:35.000
Right now I don't know that we have that because it's so new and it's so fragmented and it's sometimes open source and sometimes not transparent that we're all just kind of going in different directions and you can't tell who the reputable and trustworthy sources versus the one who just discovered a new model.

00:33:35.000 --> 00:33:44.000
I would add to that though that there's a significant opportunity on this as well. You know, thinking about data sets, for example, that we know are biased.

00:33:44.000 --> 00:33:55.000
Synthetic data could be the way to unbiased those data sets. So, you know, for example, We know that generally speaking, historical financial data is deeply biased data.

00:33:55.000 --> 00:34:03.000
So either let's say you're a bank and you're trying to make a model you could go through your real data.

00:34:03.000 --> 00:34:06.000
That and correct it for that bias by saying, well, this should have been accepted. This should have been accepted, but it was rejected.

00:34:06.000 --> 00:34:24.000
And re unscrew it back towards an unbiased data set. Or, I mean, frankly, if you have enough data, that could be the opportunity to say, well, we're going to augment this fully with synthetic data to to make an unbiased data set.

00:34:24.000 --> 00:34:28.000
So there's a significant, it's not, it's not all bad news, right?

00:34:28.000 --> 00:34:37.000
There's a significant opportunity there to. To be better than natural data. But that, that's a gargantuan lift, even though we're talking about this being accessible and, you know, easy.

00:34:37.000 --> 00:34:38.000
Okay.

00:34:38.000 --> 00:34:45.000
You can do it on a laptop. That that level of complexity is still really significant.

00:34:45.000 --> 00:34:51.000
Yeah, I would agree with that and we've got some examples at McDonald's where because we over We're a global organization.

00:34:51.000 --> 00:35:06.000
We have data and restaurants and customers all over and the perception and the behavior of those customers at McDonald's is different based on not just the geography within globally, but even within the United States.

00:35:06.000 --> 00:35:20.000
Because there's some markets like the United States that are so large. It creates a lot of noise within our existing datasets and we are able to use synthetic.

00:35:20.000 --> 00:35:35.000
Data to kind of expand the influence or the weight that's given to other behavior patterns or other markets so that we can do the same type of

00:35:35.000 --> 00:35:47.000
Data modeling that requires kind of these massive data sets, mass amounts of data for smaller markets without them getting washed out in the noise coming from the United States.

00:35:47.000 --> 00:35:54.000
So that does offer us just kind of a tangible example of using the strategy that Ben was talking about.

00:35:54.000 --> 00:36:00.000
That's actually it brings us kind of full circle back to one of the things we discussed earlier, which is where did this come from?

00:36:00.000 --> 00:36:09.000
And that sort of feels like the data waiting, right? You don't have enough of a specific population in the sample, so you build out waiting to get yourself to where you need to be.

00:36:09.000 --> 00:36:18.000
And so. Does that make sense and it allows us to do that that smoothing or to focus on a population that might be smaller or that is

00:36:18.000 --> 00:36:39.000
Maybe not creating much of an action within the data set. But it, we need to have the understanding of what the data is going into it and what the drivers are behind it before we can actually even get to that point again.

00:36:39.000 --> 00:36:40.000
Okay.

00:36:40.000 --> 00:36:44.000
So for for the other statisticians out there. If, if bootstrapping blew your mind when you first started using it, this is like boot wrapping on steroids.

00:36:44.000 --> 00:36:54.000
Yeah. Well, Chris, there's a few questions here about your personas. I think it boils down to if you're using, primary data.

00:36:54.000 --> 00:37:05.000
To create these personas is a synthetic data or or is that just analytical segmentation? I think I think we still are just having like what is synthetic data in the chat a little bit for me.

00:37:05.000 --> 00:37:06.000
Yeah.

00:37:06.000 --> 00:37:16.000
Synthetic data and so let me have you react to this. I'm gonna try to I'm gonna try to give what we what at the association what we're referring because there's also a question from Kamal about is it primary data collected by your own org is anything that's not primary data click by your own organ.

00:37:16.000 --> 00:37:35.000
So it's more complicated than that because you've got primary data which you've collected and then you've got first party second party third party 0 party you know and then you've got secondary data which is data that was collected by another org that you're using and then got third-party data.

00:37:35.000 --> 00:37:44.000
Synthetic data is none of that. It is generated data that's data generated from those primary data sets, right?

00:37:44.000 --> 00:37:46.000
Whether it's yours or and secondary data sets that are real data from real people. Synthetic data is data not from real.

00:37:46.000 --> 00:37:55.000
Real people. It's the data generated from real people, but not for, right? Am I saying that right?

00:37:55.000 --> 00:38:17.000
You are and it's basically you know and of course you know the the models themselves are trained on data you know we as when we talk about bias with my response to biases is always to say a large part of the training of the these things is based on Reddit.

00:38:17.000 --> 00:38:18.000
Yes.

00:38:18.000 --> 00:38:20.000
Have you ever been to Reddit? You know, so I mean that's part of what's in the model.

00:38:20.000 --> 00:38:31.000
So you're absolutely right. It's what those beyond that. So the question about is it still synthetic if we've built a synthetic persona based on the based on the real data?

00:38:31.000 --> 00:38:33.000
The answer is what's happening is and this is why we restrict it to things like personas where we believe this is okay and ethical.

00:38:33.000 --> 00:38:52.000
So for instance, you know, we have a person where I was working with a client on a particular persona and the person said, So does this person have a pet?

00:38:52.000 --> 00:39:01.000
And of course that's not in the data, that was but we asked and the persona said yeah I've got fluffy I got fluffy from a from a from a shelter a couple of years ago.

00:39:01.000 --> 00:39:10.000
Now what's happening is this is the plausibility kicking in. It's synthetic in the sense that we didn't bring that part of the data to it.

00:39:10.000 --> 00:39:16.000
But the, but the model is saying it's plausible that this person, this persona, would actually have a pair.

00:39:16.000 --> 00:39:22.000
And if they did have a pair, they probably have got it from the shelter. Because of their other, because of their other values.

00:39:22.000 --> 00:39:45.000
Because we program in the values. So that's the point where it becomes synthetic and you're exactly right, it becomes synthetic and you're exactly right, Melanie, it's it's when we go beyond what the data says, it's when we go beyond what the data says that it becomes synthetic, it's when we go beyond what the data says that it becomes synthetic that we're getting it from

00:39:45.000 --> 00:39:47.000
the data says that it becomes synthetic that we're getting it from the models assumptions that it's becomes synthetic that we're getting it from the models assumptions that it's layering on top.

00:39:47.000 --> 00:39:59.000
Yeah, and I'm trying to think about governmental regulation too, right? As your association. So what I'm trying to do is give a ring fence around a definition because it's going to be regulated.

00:39:59.000 --> 00:40:11.000
And so I want them to regulate the right thing and not the whole thing. And so I'm, you know, We've been collecting primary data for years and years and years and years and years and not had to have that regulated.

00:40:11.000 --> 00:40:16.000
So, I, but, but John's making a point that, John's making a point that, you know, having a single term we are pointing back to is likely going to trip us up and and that may be true.

00:40:16.000 --> 00:40:30.000
Especially right now when it's so evolutionary and changing so rapidly. But I'm hoping we can come to a term that we all can live with and that also is, puts a ring fence around what gets regulated.

00:40:30.000 --> 00:40:52.000
That brings us to regulation. And then maybe the, and Zack too in your role of watching things that are regulated at McDonald's, like what should people be thinking about in terms of coming regulation, legal boundaries, Ben, you're already leaning forward, jump in.

00:40:52.000 --> 00:41:01.000
Well, I think as far as regulation a lot, a lot remains to be seen, particularly in how effective it is, whether it is.

00:41:01.000 --> 00:41:17.000
Behind the times. I mean, there's this much touted push by the EU to get AI regulations in place and the just the debate the negotiation is kind of broken down because it's already there the premise is already out of date.

00:41:17.000 --> 00:41:39.000
So the, this is, this is a field that is moving so fast. That you know. Recently in a Supreme Court decision they sort of famously said like we're not the the greatest experts on the internet that ever lived and that's doubly true for legislators who are trying to get their heads around a technology that is Not at all a part of their their wheelhouse.

00:41:39.000 --> 00:41:51.000
You know most most politicians started as lawyers. This is not really the traditional domain of the lawyer. So there's gonna be a speed issue of can they regulate.

00:41:51.000 --> 00:41:59.000
Fast enough that it is effective. Or are they gonna be sort of regulating yesterday's danger?

00:41:59.000 --> 00:42:04.000
I don't know what that is gonna look like when it shakes out. I'm not even confident necessarily that it will shake out.

00:42:04.000 --> 00:42:12.000
We may still be having these debates 5 years from now when the the baseline of what we can do with.

00:42:12.000 --> 00:42:18.000
AI how we're using AI and synthetic data it may be a totally different place and we're still, you know.

00:42:18.000 --> 00:42:26.000
Trying to get a handle on it. But as for

00:42:26.000 --> 00:42:27.000
Okay.

00:42:27.000 --> 00:42:29.000
That's also the importance of the Insights Association and like we're having those conversations with them to fill in their knowledge gaps and to say, no, don't do that because that would harm a, you know, a hundred 50 billion dollar profession.

00:42:29.000 --> 00:42:37.000
So that's, yeah, that really important process.

00:42:37.000 --> 00:42:42.000
One, as to the legal component, it's sort of the same issue and that we're still contextualizing.

00:42:42.000 --> 00:42:50.000
This new thing in sort of the old way, right? So our concerns are copyright, privacy, things that we've always been concerned about, discrimination and bias, but in this totally new format.

00:42:50.000 --> 00:43:05.000
And to Katie's point earlier, I think part of the reason why legal departments, compliance departments generally sort of Put a stop to this is that.

00:43:05.000 --> 00:43:12.000
It's not really clear. What needs to be done about it, right? Or what to what degree?

00:43:12.000 --> 00:43:20.000
You know an MSA with a new vendor is gonna get into the specifics of the technology underlying this thing that they're providing, right?

00:43:20.000 --> 00:43:34.000
It's not even necessarily clear how you contract for that. So I mean, I think part of that hesitancy is that we as lawyers are on the back foot of not totally knowing what.

00:43:34.000 --> 00:43:46.000
We should be signing off on and so instead just saying, well, none of this then. But But I think those main concerns of what's going in, where did it come from?

00:43:46.000 --> 00:43:56.000
What did people? Consent to, what did you pay for, what do you have the rights to, all of that kind of It's it's a perfect storm there

00:43:56.000 --> 00:43:59.000
Yeah. And if you guys didn't notice, Howard Feinberg is, is in the chat.

00:43:59.000 --> 00:44:09.000
He's on the session with us today. He's your advocacy lead and and your lobbyist and he was saying that he was just discussing synthetic data with Senator Schumer this week.

00:44:09.000 --> 00:44:17.000
So definitely our job to help them understand how it's being used and and what should be regulated and what shouldn't.

00:44:17.000 --> 00:44:27.000
So, and I, I agree with what you said, Ben, you know, those providence ownership, John Bremmer put who owns the data, you know, do you have the right to it?

00:44:27.000 --> 00:44:34.000
Is it your own data? Is it a walled garden that you have all the rights to and so then therefore do you own all the outputs or is it a shared data set somehow and then do you own all the outputs?

00:44:34.000 --> 00:44:42.000
Do you have the rights to use it? It's a really big questions we need to be asking ourselves and was that consent when it was collected, you know.

00:44:42.000 --> 00:44:43.000
A lot of questions.

00:44:43.000 --> 00:44:50.000
Well, and some of these things we're going to endlessly talk about them until eventually someone gets sued and it goes up the ladder.

00:44:50.000 --> 00:45:01.000
I mean, right now we're in a world where unless you're in China as of last month, something that wasn't created by people can't be copyrighted.

00:45:01.000 --> 00:45:02.000
Hmm.

00:45:02.000 --> 00:45:05.000
Can't be copyrighted, you can't protect it. Can't protect your right to it.

00:45:05.000 --> 00:45:11.000
Anyone can use it. So I would say that's a powerful disincentive for enterprise use.

00:45:11.000 --> 00:45:32.000
Where, for example, you don't wanna use. A tool like Dolly or Mid Journey to brandstorm your, corporate logo and then use one of them and then you can't protect that right that's a no one's gonna do that that's not a great example but that sort of thing of Well, if I can't then make it exclusive to me, why would I use it?

00:45:32.000 --> 00:45:40.000
I'm still going to go with the human. Maybe they've brainstormed using using these tools, but the end result can't look like that.

00:45:40.000 --> 00:45:43.000
And

00:45:43.000 --> 00:45:50.000
But changing that is also gonna require that everyone get on board with doing it, which would then sort of force the needle.

00:45:50.000 --> 00:45:55.000
And that doesn't really seem to have happened yet. So that's another. Maybe we'll still be having this conversation in 2 years.

00:45:55.000 --> 00:45:59.000
Maybe it'll be a totally different environment.

00:45:59.000 --> 00:46:05.000
Yeah. Well, so there's a couple questions here I see about platforms and tools. Are there platforms and tools that researchers are using, are there platforms and tools to create personas?

00:46:05.000 --> 00:46:20.000
Are there platforms and tools to implement vector databases like what tools exist today that you think people should be playing with and and investigating.

00:46:20.000 --> 00:46:25.000
So I'll, I'll jump in on this as you know that this is something that I care about.

00:46:25.000 --> 00:46:36.000
I've been saying for many years that the the insights industry I don't know the only person saying this, but the Insights industry.

00:46:36.000 --> 00:46:37.000
Yeah.

00:46:37.000 --> 00:46:38.000
Okay.

00:46:38.000 --> 00:46:39.000
Has to get hip with Python. If you could, if you could be hip with Python, nerdy with Python, let's say that.

00:46:39.000 --> 00:46:48.000
The reality is, is that to get anything beyond the very basics, you've got to be writing code.

00:46:48.000 --> 00:47:00.000
Either you have to be writing code or your partners have to be writing code. I'm not saying that everybody has to but either somebody in your company has to be or your partners have to be doing this.

00:47:00.000 --> 00:47:07.000
So, you know, the question about this synthetic personas, that's a system we built ourselves from scratch.

00:47:07.000 --> 00:47:16.000
In terms of the actual tools that are a number of great tools out there there's this one called lang chain which is a little controversial in some quarters but it actually makes building these kind of workflows a lot easier.

00:47:16.000 --> 00:47:33.000
There are a number of other tools a lot easier. There are a number of other tools all within the Python ecosystem pretty much all roads, all within the Python ecosystem pretty much all road to Python when it comes to this. There are other ways to Python when it comes to this.

00:47:33.000 --> 00:47:36.000
There are other ways to lead to Python when it comes to this. There are other ways of doing it, but that's pretty much it. Somebody talked about vector day spaces.

00:47:36.000 --> 00:47:39.000
There are a number of if you're in, if you're in memory, you can use, Meta's face database where Churchill is very good and I use a lot.

00:47:39.000 --> 00:47:54.000
But then there's commercial ones like Pinecone and actually a whole slow of others that are that are coming up and vector data databases are going to be huge in insights because we work with a lot of data and we should be working with our own data.

00:47:54.000 --> 00:48:21.000
So, vector databases are going to become very important. But basically I would say you gotta write code don't you or one of your partners has to write code at some point if you really want to get the best out of this you're not going to plug something into chat GPT.

00:48:21.000 --> 00:48:22.000
Ever ever.

00:48:22.000 --> 00:48:32.000
And you must never put your clients data into check GPT. Never ever. And so you should be building your own systems or working with companies that have their own more closed systems.

00:48:32.000 --> 00:48:33.000
I think that's, oh, go, go ahead, Jack.

00:48:33.000 --> 00:49:01.000
I just. I just wanted to echo some of what Chris points for. So we, I lead the data, the global data sense org at McDonald's and about a year ago we merged with the more traditional insights group, consumer insights and business insights and foresight and we all became one team and I think my whole team writes code and is very familiar with Python and, building vector databases

00:49:01.000 --> 00:49:06.000
and building vector databases and other similar capabilities from scratch. Building vector databases and other similar capabilities from scratch.

00:49:06.000 --> 00:49:18.000
And a huge gains, I think, in the speed and complexity of the insights we were able to deliver by partnering these 2 organizations.

00:49:18.000 --> 00:49:35.000
Because the ability to write. Your own kind of custom code either to join on to a vendors platform or to build a simple solution that you can turn around relatively quickly.

00:49:35.000 --> 00:49:42.000
Unlocks a lot of a lot of capabilities and breaks down quite a few bottlenecks.

00:49:42.000 --> 00:49:48.000
And then back to day basis, which I'm in general a big fan of. I also agree are gonna be important.

00:49:48.000 --> 00:50:05.000
We talk about the difference between purely open source or open frameworks versus like a walled garden environment. And I do think to get the most and the most security out of a lot of these LLMs and other synthetic data generation systems.

00:50:05.000 --> 00:50:24.000
You do need to have. Your own walled garden of data. That prevents sharing of customer data or private information and satisfies all the legal requirements, but then still unlocks the power of these kind of models that are trained on large large.

00:50:24.000 --> 00:50:39.000
Datasets and the best way to do that is through kind of homegrown or bespoke vector database style systems that can take a bit like at McDonald's, for example, tons and tons of operations manuals, right?

00:50:39.000 --> 00:50:52.000
For the restaurants and how to do stuff and a big pain to search. We don't necessarily want those out there, not that it would create a lot of problems, but you know, just there's no reason for them to be out in the public.

00:50:52.000 --> 00:50:54.000
And available, but if we can create a vector database that contains all of that documentation and the relationships and then have a language model that can access it.

00:50:54.000 --> 00:51:08.000
You can do a lot with that same with like the custom, segmentation type work that Chris was talking about earlier.

00:51:08.000 --> 00:51:11.000
So.

00:51:11.000 --> 00:51:18.000
Alright, anyone else? Well, so, there's one more question, then we'll move on to sort of our last questions.

00:51:18.000 --> 00:51:24.000
There's a quick, there's a few questions here about how brands are or are not using what brands.

00:51:24.000 --> 00:51:37.000
How brands are or are not using synthetic data and I think they're talking more about like both the kind of synthetic data that you've been talking about but also this question about panels that are starting to offer synthetic data.

00:51:37.000 --> 00:51:52.000
How are how are brands thinking about that? Specifically, Zack, there's a question for you that says would McDonald's consider using synthetic data generated by panel to make a decision, how and where would you consider using that.

00:51:52.000 --> 00:51:53.000
Okay.

00:51:53.000 --> 00:51:56.000
So, so maybe put Zack on the spot a little bit, but everybody, how are we seeing it used?

00:51:56.000 --> 00:52:12.000
Yeah, I mean, I think we recently have been talking to some vendors about what they can There's some really interesting stuff around, product creation for us.

00:52:12.000 --> 00:52:31.000
It's very expensive to create and test and receive panel feedback for. Food items. A lot of times if we're if we're doing it for a specific market.

00:52:31.000 --> 00:52:38.000
We will physically ship basically an entire kitchen to like Nielsen's New York office, right?

00:52:38.000 --> 00:53:00.000
And they'll set it up and they'll do build a panel and do all this. Do all this stuff if we could If we could streamline that process by developing a virtual panel with virtual flavor profiles of our products, which theoretically we've had discussions about it as possible.

00:53:00.000 --> 00:53:20.000
We would definitely be interested in exploring more. I just, I'm not confident or the place where we're running up against now is what Again, how Is it really getting you still need to build kind of a source panel for anything new and then kind of extrapolate the information or enhance the the bulk of the information from there.

00:53:20.000 --> 00:53:37.000
That still costs costly. So what is the value of kind of the second half of the panel being almost exclusively virtual and then how accurate is it relative to, you know, a real-world experiment.

00:53:37.000 --> 00:53:51.000
There's just not a lot of benchmarks available to say this is achieving a certain level of accuracy relative to, you know, the normal panels or the in-person panels.

00:53:51.000 --> 00:53:53.000
You're on mute.

00:53:53.000 --> 00:53:56.000
Yep. Katie, you're in the panel world. What do you think?

00:53:56.000 --> 00:54:08.000
Yeah, I think. I think it's going to be a key area of growth over the next couple of years but it is it is that as you mentioned the benchmarks the accuracy.

00:54:08.000 --> 00:54:16.000
Flavor profiling in particular, or we are still going to need every single day, every single week every single month, real consumers, tasting their changing profiles, the changing palette of consumers in particular.

00:54:16.000 --> 00:54:23.000
So I'm glad you touched off on player testing for this for this topic. It's going to take.

00:54:23.000 --> 00:54:25.000
A long time I think for us to gain that trust but it's definitely going to happen. I think we need it to happen.

00:54:25.000 --> 00:54:31.000
In the industry, Melanie and I were chatting about the impact of kind of quality concerns in the panel industry right now.

00:54:31.000 --> 00:54:39.000
Kind of spiking. And so since the synthetic panels potentially being more accurate given the some of the quality challenges that the parliamentary faces today.

00:54:39.000 --> 00:54:48.000
There's going to be the way forward.

00:54:48.000 --> 00:54:51.000
Damien, anybody else wanna add anything?

00:54:51.000 --> 00:55:00.000
I don't, at least in what I've seen, people are asking about it right now, but I've not seen a lot of people actually using.

00:55:00.000 --> 00:55:09.000
Employing synthetic data for decisions. Yeah, they're every now and then. Something comes up, but that's the exception rather than being the norm.

00:55:09.000 --> 00:55:12.000
And so to Katie's point, I think it's it. We're a long way off before that happens.

00:55:12.000 --> 00:55:23.000
There are lots of things that have to fall into place first. We have to have. Models in place we have to have enough data in the mode that I can be able to use something.

00:55:23.000 --> 00:55:33.000
However, I think we'll. Start to see it happening in those places where we do have data where like I was talking about earlier for scenario planning.

00:55:33.000 --> 00:55:34.000
I've seen those sorts of things already start to happen because like, hey, we have sales here from the last 20 years.

00:55:34.000 --> 00:55:43.000
Well, let's plop it in and see what happens, right? And so I think we'll start seeing it for those sorts of.

00:55:43.000 --> 00:55:49.000
Things first because those are the places where I've seen people actually, hey, yeah, let's let's try it.

00:55:49.000 --> 00:55:53.000
To that point though, it's not. Chat GPT that they're using to do it.

00:55:53.000 --> 00:56:08.000
It's they may start with an open source. Base of some model and then it's internalized and walled off and they have their own instance of it with their own data and it becomes something that only the organization itself has access to so that it's not this.

00:56:08.000 --> 00:56:21.000
Public. Anyone can check it out sort of. Sort of model. So yeah, I think we're a little bit off and it's not there's not really an option yet, but considering how fast AI is moving people are.

00:56:21.000 --> 00:56:28.000
Are much more sensitive to it right now than you would. Normally see it for the growth of a new technology.

00:56:28.000 --> 00:56:34.000
Great. Well, so okay, we got 4 min left and so for each of you, one final opportunity to share what's on your mind regarding this topic.

00:56:34.000 --> 00:56:48.000
What's one thing that everyone listening today should either know or do when it comes to synthetic data so as not to be either misinformed or behind the curve.

00:56:48.000 --> 00:56:59.000
Who wants to go first? One thing everyone should know or do. Chris we haven't heard from you in a minute how about you go first and then we'll do Katie

00:56:59.000 --> 00:57:07.000
Okay, I think the most important thing. Is to not just rely on the model but to bring your own data.

00:57:07.000 --> 00:57:17.000
I think using using it to augment data. Or to better understand data rather than to generate just Just brand new data.

00:57:17.000 --> 00:57:20.000
I think that bring your own data is what I would say to everybody.

00:57:20.000 --> 00:57:22.000
So BYOD.

00:57:22.000 --> 00:57:26.000
If, exactly.

00:57:26.000 --> 00:57:29.000
For me.

00:57:29.000 --> 00:57:30.000
Yeah.

00:57:30.000 --> 00:57:31.000
Right. But don't put it into an open platform like Chatty. For your potato to a walled garden or some sort of private world.

00:57:31.000 --> 00:57:36.000
But, Okay.

00:57:36.000 --> 00:57:46.000
For me, I think it's about what is the right use case for this tool. So it's not, you know, somebody you've mentioned Max Differ, we don't need Max different, absolutely everything. It's for the right use case.

00:57:46.000 --> 00:57:50.000
I think when we think about synthetic data now and in the future, what is the right use case?

00:57:50.000 --> 00:57:55.000
It will not, certainly not be for all use cases.

00:57:55.000 --> 00:58:01.000
Right. How about how about next we do Ben and then. Damien and then Zack.

00:58:01.000 --> 00:58:09.000
I would say the most obvious is this is changing so fast that if you want to keep a finger on the pulse, you need to be constantly reading about this constantly playing with new tools.

00:58:09.000 --> 00:58:16.000
Constantly seeing what the the freshest newest thing is. Or it will rapidly get away from all of us.

00:58:16.000 --> 00:58:22.000
And the other. Thing that sort of ties into what Katie just said is we should be setting a very high bar.

00:58:22.000 --> 00:58:40.000
You know, I think a basic tenant of this business is that If you talk to enough people. Patterns will emerge that you would otherwise not have known about, but which are observably true and real in the real world.

00:58:40.000 --> 00:58:47.000
And that should be the bar for this that the level of validation that we expect should be very high.

00:58:47.000 --> 00:58:56.000
And is that gonna put a damp on things? Maybe, but I think, you know, the priority is getting quality insights and that's the only way to do it.

00:58:56.000 --> 00:59:06.000
Yeah, I remember when we moved to online we did a lot of parallel testing and a lot of if I asked it this way and this way and this way and that's the world we're gonna live in for a little while.

00:59:06.000 --> 00:59:07.000
Damien and then Zack.

00:59:07.000 --> 00:59:10.000
That's where I was gonna go, validate, validate, validate, validate, right?

00:59:10.000 --> 00:59:22.000
Compared to current data sets, if you're using it for predictive modeling compared to the actual outcomes and base your decisions on what the synthetic data is providing you on outcomes rather than.

00:59:22.000 --> 00:59:27.000
The inputs or the actual analytics behind it.

00:59:27.000 --> 00:59:30.000
Wonderful. And Zach, you're up.

00:59:30.000 --> 00:59:31.000
I don't, I don't know if I have anything new to add, but just don't.

00:59:31.000 --> 00:59:32.000
Okay.

00:59:32.000 --> 00:59:39.000
Yeah, test, test, don't be afraid to start small. It doesn't have to be big.

00:59:39.000 --> 00:59:47.000
There's lots of patient. There's room for a lot of patients because things are changing quickly and fairly dramatically.

00:59:47.000 --> 00:59:53.000
Who knows what's gonna come in terms of legislation or if tomorrow Amazon is gonna decide they have the best model in the world.

00:59:53.000 --> 00:59:54.000
Hmm.

00:59:54.000 --> 01:00:01.000
So. Yeah, test, test and be patient.

01:00:01.000 --> 01:00:02.000
Yeah, I love that. And then John Brimmer. Thank you for being with us and being so busy in the chat.

01:00:02.000 --> 01:00:12.000
He says we should also think about variation. A single draw from a probabilistic model could be made better with multiple draws.

01:00:12.000 --> 01:00:21.000
I love that. Yeah, sample and then sample again and then sample again. So I love that. So thank you all very much.

01:00:21.000 --> 01:00:26.000
If you are, listening, let me tell you a couple of things. If you are listening, let me tell you a couple of things really, really quick.

01:00:26.000 --> 01:00:30.000
I'm gonna tell you a couple of things really, really quick. I'm gonna pop my screen back into my slides really really quick.

01:00:30.000 --> 01:00:35.000
I'm gonna pop my screen back into my slides and then pop my screen back into my, slides and then just we have some really important, my screen back into my, slides and then just we have some really important, town halls coming up and webinars coming up in January.

01:00:35.000 --> 01:00:41.000
How to ask race and ethnicity in international research will have those findings out on January, the eighteenth as well as how and when to ask.

01:00:41.000 --> 01:01:01.000
A gender identity and sexual orientation and research and how quickly that whole world is changing. So really, really important sessions for you and we'll be doing 2 sessions one and with a longer one for the second one so that we can field some Q&A.

01:01:01.000 --> 01:01:10.000
And then, finally, our upcoming 2024 events annual conference in Atlanta leadership event in Charleston and our CRC in New York City.

01:01:10.000 --> 01:01:16.000
Thank you. All very, very much. I'm going to stop sharing so you can see your wonderful panelists again.

01:01:16.000 --> 01:01:32.000
Thank you all so much. You made a difference, some great comments here and even some people saying that that the this kind of conversation that makes C in size association valuable so thank you to the people in the in the chat saying that I appreciate you sincerely and we'll do this again.

01:01:32.000 --> 01:01:39.000
Let's do this again in 6, 8 months and see how it's changed. That will be a really cool way to to do a follow up on this.

01:01:39.000 --> 01:01:43.000
Thank you all very much. Happy holidays to everyone if we don't talk sooner. Enjoy.

01:01:43.000 --> 01:01:49.000
Make sure you take some time to yourself and to your family over the next couple of weeks. And take care of yourself.

01:01:49.000 --> 01:01:53.000
Thank you all very much and see you on the next town hall.

01:01:53.000 --> 01:01:54.000
Bye.

01:01:54.000 --> 01:01:55.000
Bye.

01:01:55.000 --> 01:02:05.000
But

