SSE / Datopian weekly sync - June 16 VIEW RECORDING - 61 mins (No highlights):


0:02 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Just wondering, since I didn't get any confirmation from Jeff, we are going to have a guest on our meeting today, right?

0:10 - Rich Baker Yeah, so Vijay is going to pop on, but he says he's going to be about 15 minutes late, he's just in another meeting, but I think he just wants to come on and just get an update on where we're at from the pen testing point of view and the actions that are open and all that sort of stuff.

0:24 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) I wasn't actually referring to that, I was referring to the guest interested in querulous.

0:31 - Rich Baker Okay, yeah, has not, has Jeff not responded on that? Yeah, it was, because we, think, last week wasn't he and then, yeah, he couldn't make it and it was, was it going to be this week?

0:44 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Yeah, yeah.

0:47 - Rich Baker Yeah, I don't know where that is at the moment, let me make a note of that, that might be, might be, that might need to do that next week. And, yeah, remind me of the name, was it, was it, oh, was Somebody called Francis, was it?

1:02 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Oh, I'm not sure, to be honest, not familiar, but yeah. Okay, no worries. It's Fraser, actually. Fraser. And he's trying to join. Hello there. Hello, Fraser. Am I pronouncing your name, actually?

2:10 - Fraser Macintyre Hi, sorry. Yeah, that's right. Sorry, was on mute. Oh, no worries.

2:15 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Welcome. Thank you for joining the call. Cheers. Rich, do you have any info? Will Joe join us? Jeff will not be joining us.

2:29 - Rich Baker Jeff had a baby yesterday. Oh, did he have it yesterday?

2:33 - Fraser Macintyre Yeah, yeah, it was yesterday. It ended up being a scheduled birth.

2:37 - Rich Baker So, yeah, that's why he's not in. he was trying. Initially, they thought it was Wednesday. So, was trying to, like, squeeze a few bits in yesterday and today. And then the hospital said, actually, no, it's Monday. So, yeah, it was all a bit of a mad rush back in the last week. But, no, I assume he's had it. From what Zoe was saying, it sounds like he has. But, no, it's it's The updates, but yeah, no, if all have gone so well, he's now a father, so. That's brilliant.

3:08 - Fraser Macintyre So yeah, I think we could forgive him for not being on this call today.

3:12 - Rich Baker He's on paternity leave now for about four weeks. I it's about mid-July he comes back. Great.

3:21 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Okay. Okay. Well, I mean, there's no better reason to be out of office for sure, so. Yeah, that's a bad reason, is it? Great stuff, great stuff. Okay. In that case, guess, for the sake of time, since we have a guest on today's call, we can immediately jump in to the main topic, which is Queryless. First of all, before Zhao takes over and start with the demonstration, I'm just wondering, Fraser, do you have a context around? Yeah. Yeah? Okay. Yeah. And so our couldn't.

4:03 - Fraser Macintyre We work across innovation stimulus projects, which are either NIA or strategic innovation fund projects. So I'd done a couple of data-related projects under NIA. So I did one called NERDA, which was near real-time data access, in which we made our real-time HVNLV data available via portals and APIs. So that was delivered under NIA. Laterally, we did a CIF Alpha project. I don't know if you're familiar with any of that, but we do. We get CIF Discovery, CIF Alpha, and CIF Beta. So your discoveries are about $150,000, like, three months worth of effort. Your alphas are about $600,000, six months worth of effort, multiple partners. And then your betas are up to, I think, 10 or 12 million multi-year projects. Thanks. So we did an alpha project called Data to Insights with ourselves, University of Strathclyde, Energy System Catapult, and a couple of others to look at how we could use AI to drive insights from our data sets. So looking at open data and stuff like NERDA as well so that our stakeholders could pose natural language questions to an AI to deliver insights rather than having to churn through like multiple different data sets in order to, you know, because if you're like doing a connection, you've got capacity related question, you know, you need to take some real time data to understand what's happening. You need to understand what constraints are in that local area. You need to understand what the future development plans are for that substation as well. You could also look at cable ratings. So you're having to take a lot of different data sets in order to understand what are, you know, doing a… Connection there is feasible or placing a place, putting a place, or where's best place to place some flexibility. So we did the alpha project to look at how we could use a sort of natural language AI interface in order to deliver insights rather than just the raw data that our stakeholders are using currently. So the alpha project was quite successful. And then we looked to move to beta, but one of the things we discovered when we were doing the alpha project was one of the main things that we were lacking was like the orchestration governance for our data sets going into an AI model. So telling that AI agent, giving that AI agent some orchestration information about the data set that it was looking at. So telling it how it could understand it, telling it how it could interpret it, telling it what What guardrails it needed to put in place. So when we moved into Beta, we sort of pivoted, and instead of just delivering the insights and moving on to bigger insights, more data sets, we kind of focused a lot more on the orchestration and developing some governance for how energy-related data sets could be orchestrated within, sort of natural language, large-language AI agents, and we didn't get the Beta funding, however, we got a lot of positive feedback on the project, and OffGem and Innovate UK would like to see it again, which is something we're working on. They'd like to see another DNO or a transmission operator join us, which we're looking at. And… And… but the actual problem and the need for orchestration and governance.. down. In this area, they understood and recognized, and that was sort of what led on to some of the conversations with Jeff and some of the stuff that you guys are presenting around, like sort of insight-driven use of AI, and it would be good to understand, like, you've obviously got a platform that would do that, and you take in, like, open data in order to do that. It would be good to understand what your requirements are around, like, orchestration and governance for those data sets that you're consuming. Yeah, and just to sort see where you are, and if you sort of also recognize that need for orchestration.

8:41 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Yeah, yeah. First of all, thank you very much for the intro. That's definitely really valuable, and it's really good to understand what you're actually looking for and where you are. So, I guess, Jao, you can take over.

8:59 - Fraser Macintyre And…

9:00 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) And possibly answer the questions. Yeah.

9:04 - João Demenech Hello, everyone. Nice to meet you, Fraser.

9:08 - Fraser Macintyre So I'm João Demenech, I'm a software engineer at Datopian.

9:11 - João Demenech Just an introduction. And I've been working on Quarrelous AI.

9:19 - Fraser Macintyre So just wondering, have you seen a demo of it? Are you familiar? No? No, your platform, no. no. Okay. Yeah.

9:28 - João Demenech So I think it would be good if I do a quick demo. Yeah. Yeah. So we have this Porto.js demo portal here in which we integrated Quarrelous AI. It lives in this floating action button currently. You can see Ask AI here. It's always accessible regardless of the page we're browsing. And even if I open that up, it keeps the context. Also, you can notice that… that… that… Yeah, it says here, viewing search. And if I go back to the home page, it should say, yeah, viewing home. So we are also feeding information about the page that the user is browsing as context to the AI, so that the AI can figure out exactly what the user is looking at. And with that, I'm going to start, like, the usual user journey here. I know for sure there is this world happiness data set. I'm going to ask it, are there data sets about happiness? And this is one of the use cases that we have implemented for the AI to assist with discovery of data sets. So what it's going to do behind the scenes now is it's going to figure out the intent and then it's going to use the CKAN APIs to do a data set search. And then it's going to return results. And since it's aware of the portal itself, even though this is not CKAN, this is the portal.js frontend, which is using the CKAN APIs, but it's a decapital frontend, you can see that it can provide relative links like it did here. So it's pointing out to the happiness dataset that currently lives in the portal. And just before I click on that, I can also show you that if I ask, for example, are there datasets about transit, it shouldn't find anything just because it's grounding its responses. Yeah, notice it's about transit. It's grounding its responses on the CKAN API. Okay, so now I'm going to actually navigate to the dataset. You can see here, now it's updated saying that the user is viewing the dataset world happiness dataset. The user would likely now explore a bit the metadata, review the data.

12:09 - Fraser Macintyre Here we can take a look at the columns.

12:11 - João Demenech In this dataset, for example, it's basically the happiness perception across 153 countries. And now I can continue the chat. You can see it remains with the same conversation history as before. And I can start asking questions about the data itself. For example, we could ask, what's the average happiness perception across all countries? And in this case, what the AI is going to do is, it's going to figure out the metadata. About this data. So it's going to figure out the columns that are available, the column types and so on.

13:07 - Fraser Macintyre And then it's going to create a SQL query, which we can inspect on this, how this was calculated.

13:16 - João Demenech You can see it created this SQL query and then it runs it against the data and formats the response. So in this case, we are getting 5.47. This is a result that I have already verified before.

13:35 - Fraser Macintyre And we can also evaluate the query, which looks right. So it's the average letter score, which is the happiness perception call. By the way, feel free to interrupt me anytime if you have any questions or comments. Now I'm going to ask you, for example, example, generate.

14:04 - João Demenech This is another capability that we implemented, which is for generating charts. So in this case, the AI will generate a Vega spec and the chart will produce the specs.

14:25 - Fraser Macintyre So here we just see the interactive chart, sorry.

14:31 - João Demenech And yeah, we also get the, how this was calculated, getting country name and letter score. And I believe it's applying the limit. Yes. And ordering by letter score. I could also, another use case would be to actually, what if I want to explore this data, but I'm not sure what to ask.

15:00 - Fraser Macintyre For example, what other indicators we could calculate based on this data, and hopefully to give us some creative ideas. So, let's try this one.

15:27 - João Demenech Which factors contribute the most to happiness? Because there are some columns available here related to that.

15:37 - Fraser Macintyre So there's like social support, health-life expectancy, and so on. Yeah, so we got another chart.

15:51 - João Demenech Yeah.

15:53 - Fraser Macintyre So it's saying that social support seems to be the factor that better correlates to…

16:00 - João Demenech Average perception, followed by GDP, life expectancy, and so on.

16:10 - Fraser Macintyre Can you query across two datasets at one time? Yes.

16:17 - João Demenech I can show you this maybe using this fuel emissions and global temperature. These are the set of datasets that we've been using to demonstrate this. So I could come here, for example, and ask you to, or I don't even have to go there. I can just do it straight from here. So can you join the global fossil fuel emissions dataset with the global temperature time series? Yeah, but it should be, it should be. Capable of creating joint queries against the data.

17:06 - Fraser Macintyre And what if those data sets had competing answers for the same field? I'm thinking if you think about like the SSE and data portal and you think something like transformer rating. So we're going to have multiple data sets that have got transformer ratings for a specific geographical location in it and some of those, some of them might be different because of, you know, flaws in our data set. How would the AI deal with something like that if it was getting competing information from different data sets?

17:50 - João Demenech Yeah, so I think the best would probably be to provide enough context about the data sets so that the AI can figure out. what it has to do. So in this case, for example, it asked me whether what it was suggesting makes sense. So yeah, there are the schemas. It found out that it can join by the year. But to be honest, I'm not sure I'm getting exactly your question. Is that something that you could show us? Uh, the, uh, the distance you have in mind and so on?

18:38 - Rich Baker Um, yeah, Fraser would be like, you know, so if we, if we did follow, following the, the, the use case, if we gave like a, an annual view of transformers and that was their rating, there's like a, like an annual report. This is our network type, type report. Then we added another dataset onto the portal that was like updated monthly, which was, you know, what are our transformers and substations and the associated, you know, ratings and capacities and all that sort of stuff. From a query, if you were to query it, it would find the same transformer or substation or something else within the two datasets. But to Fraser's point, it might have two different answers on it because one would be that point in time, which might be your annual one, and that's what it was at the time, whereas the monthly ones might be more fluid and they might step up and step down a little bit and the ratings might change. Would, if you had just those two datasets, would the AI come back and say, I've found, I've found the same, the same transformer or the same asset in two datasets, but I'm getting two conflicting reports, on report one it says X and And on report 2, it says, why? Or would it try and be cleverer and say, I found two different responses, and then I just averaged it out and said, I think it's this? It's almost a, at what point would it stop presenting facts and start, not necessarily hallucinating, but start to create its own data? Does that make sense?

20:21 - João Demenech Yeah, I see what you mean. So you're seeing that there's a bit of a ambiguity there, because it could just come up and use the annual data, or it could use the monthly data, and it might make a decision on behalf of the user, right? Yeah, I believe that by the full, it would just make that decision. To be honest, I'm not 100% sure, but that's what I'm assuming, based on what I have seen. It could be the case that, for example, here, I asked for the data sets, and it did. The search. So it was asking like, is that really what you want? But I'm not sure. It might not be always this way. But I think the best probably would be to give it some sort of bias in terms of context. Probably this way we could make it more aligned with this use case. Or even I think like the actual user experience doesn't have to be exactly what we have here. So for example, if you if you often join these data sets, there could be, for example, custom sort of user experience just for these data sets and so on.

21:53 - Fraser Macintyre If that makes sense. Yeah, so it sounds like that.

22:00 - Rich Baker That's something we would have to declare and direct what it would do rather than letting it assume, yeah, yeah, okay.

22:09 - Fraser Macintyre So, mean, that's where we were getting to with the orchestration that for each, you know, data set on our, that we allow an AI agent to look at, we need to provide it with some context for, you know, what it can do with that data set. Maybe if there's, like, there's maybe another data set that might take priority, if, because it could be that, you know, going back to the Transformer rating one now, that, like, the annual one that Rich described has that value, and then the one that we're updating monthly for 99% of Transformers, that's getting updated there. But there could be, there's one missing, so we'd want the AI in that place to take the, the annual one and ignore, because there's not available in the monthly one. But then, if monthly was available, all. Take the monthly, there's these, do you see what I mean about the sort of rules in which we want to tell it to interpret that data and what it can do with it and what it can't do with it so that it sort of avoids hallucination and we don't want, because if we don't tell it the rules, we could have, you know, stakeholders come back and knocking at our door and saying, well, the AI agent has told me I could do X here when actual fact that's not what, like, an SSE engineer looking at the same data would have told that person.

23:40 - João Demenech Yeah.

23:48 - Rich Baker Yeah, it sounds like it's almost the, it's how you train the agent on the sort of rules of engagement on the data sets, isn't it? And is it, you know, always bring the most recent one back, fracking? It's, you'd want it to declare that as well. Yeah. And, you know, if it did have to go and get two data sets, 99% of it was from one, but as you say, there was a record missing, so it went and got the other, you'd probably want it to declare that and say, you know, majority of the data came from here, I did find missing records, so I've, you know, populated it with another data set. And that was another thing we were sort of potentially saying as well, Fraser, is, you know, we'd probably need, like, a health warning in there somewhere that says, you know, this is for insights and personal usage. We knew we'd need to talk to, like, maybe the legal team or something that says, you know, how you use the data is up to you, but, you know, if you're going make decisions on it, don't come back to us, I think.

24:41 - Fraser Macintyre So, yeah, that was part of the thing with data to insights that we were wanting to understand is, like, how much of a warranty or a guarantee we could give with that insight. insight as well, because, like, what you described there, like, and that's what we do at the moment, but how much value of that, how much value does that actually create for If what we're saying is you can't really rely on it, whereas if we were able to maybe put more of a warranty or a guarantee on what they could do with that, it would be creating more value for that end stakeholder. I mean, certainly what you're showing is quite powerful and the actual way it's presented is a lot more advanced than where we got to with our D2I project. But where we got to was like, actually, we need to take a step back and think about the rules in which we are allowing that AI agent to interpret that data once it's put into something like this.

25:49 - João Demenech Yeah, so on our current implementation here, yeah, we talk of the how. This was calculated as some sort of way for the user to verify what's going on. Although it looks a bit technical now, we could make it more friendly. It has this explanation, like, okay, this example here, for example, it's I calculate the average value of each explained by column across the countries. I think, like, the user could use this to, you know, sort of figure out how it went. And also, it has a bias to try to provide any assumptions that it made. I'm not sure if I have a dataset here that I could use to show that. But, for example, if there were two column names that look very similar, it would have a bias to explain it here.

26:55 - Fraser Macintyre Oh, I made the assumption that this is the right column for this. could use the And so on.

27:01 - João Demenech And also we added this warning here, saying that Parallels AI can make mistakes and we have this AI Terms of Use. Yeah. So these are the ways that we have tried to address this as of now.

27:30 - Fraser Macintyre Okay. Any other questions or comments? Not from me, I don't think. No, no, not from me.

27:46 - Rich Baker And obviously this is the second time that I've sort of seen this and it gets a bit more impressive every time I see it actually. But one of the things we were potentially talking about as well, you know, Notwithstanding the token element, but one of things we were talking around, I think it was myself, Jeff, and Michael, was around, actually, we could almost use this internally as well to do stuff like identification of conflicting data sets and that sort of thing. You know, I had some key queries that you'd almost throw at it on a monthly basis that says, is there anything in here that contradicts itself? You know, there was a good use case for internal usage.

28:28 - Fraser Macintyre Yeah, and if we were to develop rules or governance for how it could interpret our data sets, would this tool be able to consume that? Sorry, didn't understand the question. So, like, at the moment, you know, like, if you take, like, an API from one of our open datasets, and within that API, if we had some rules for an AI agent on how it could interpret and understand it, would that tool, would the current AI assistant be able to follow those rules? Ah, yes.

29:24 - João Demenech I see what you mean. Yeah, it should be able to follow the rules. Currently, we're feeding it with the dataset metadata, and you could have some sort of instruction there for it to read some, like, AI guidelines field or something in that sense.

29:52 - Fraser Macintyre Okay.

30:03 - João Demenech Yeah, so just since you're talking about internal use case, we also implemented a sort of like data quality audit instructions. So now if you ask it to audit the quality of the dataset, it will do some standard procedures. So yeah, here for example, it's saying that overall it's good, but with one minor issue. Um, it's providing an overview. There's one blank cell out of 1,300 rows. So one is missing the value for inflation. And it will, it will try to find any duplicate rows. It will try to find blanks and news. Uh, it will try to find values that look like anomalies. overlook happens 증 Thank Let's say that you have a column whose values go from one to ten, and then suddenly there is a hundred, it reflects that as a potential anomaly. Yeah, so this is an example of how it could be used internally as well.

31:18 - Fraser Macintyre And you can have other sorts of procedures like this one.

31:27 - Rich Baker So essentially what could do with that, without going into solution mode, is we could, you know, almost pull through some of the, you know, data quality aspects that, like, Donald Frazier's, Donald Camel, sorry, team do. And almost get it to do some of that as well, which is like a bit of a data quality aspect, you know, what we do on the internal data. Yeah. Again, there's a layer on, on that published data.

31:53 - Fraser Macintyre Yeah. Yeah.

32:00 - Rich Baker In theory, the logic and the measures would all be the same, wouldn't they? But actually, rather than doing it at a data set level, this might be more powerful, and does it, you know, your data set, two data sets, you know, both scored in the 90th percentile of accuracy in data and completed some art sort of stuff. But actually, when put the two together, does the score go up or down?

32:23 - Fraser Macintyre Yeah. Sounds good.

32:31 - João Demenech Sorry, there's a bit of noise. I couldn't understand the last part.

32:37 - Rich Baker Yeah, no, what I was saying is internally, when we do some of our data lake integration and then, you know, prep the data before we publish, so it's like the open data portal, that sort of thing. We've got a data quality team that do a number of measures of how they measure the quality of the data. So it's stuff like accuracy of the data, completeness of the data. Timeliness of the data, you know those types of things, multiple sort of measures. What I was saying is actually you could use those same measures and actually apply them to this as well so that either if you queried it it could tell you the quality of the data as you've done now or actually depending on how we wanted to set up the agent you could you know it could almost do that I have you know queried the data I have found the following you know but at the bottom I say oh by the way uh you know the data quality on this you know there are three rows missing bear that in mind if you are using this data for decision making or something um you know you know that sort of thing and then what we were while saying as well is where we are scoring our data and publishing it at a data set level what we potentially could do with this as well is say well if you put two data sets on that have got quite good um data quality scores If you publish them both, does that increase the quality of the score because they become complementary, or actually does it reduce the overall quality of the score because actually they start to conflict with each other? I suspect it would be the former, but that would be, again, useful insight for both us and how our data consumers would be using the data.

34:23 - Fraser Macintyre Yeah, so I get what you mean in general. Yeah, I think absolutely.

34:31 - João Demenech I'm not sure, yeah, whether, like, whether quality, in the example that you provided about multiple data sets, yeah, I would have to think about this. It's just not something that I, I have any thoughts right now.

34:52 - Rich Baker Yeah, I'm just sort of thinking, you know, again, without going to solution mode, so like this, the one that we see on the screen, you know, if I was the user, I'd be like, yeah, okay, I get it. There's one cell that's missing a value. In the grand scheme of things, it doesn't feel like that's a major impact in the output. And then even the recommendation says it's negligible.

35:11 - Fraser Macintyre It's unlikely to affect the analysis. That's great.

35:13 - Rich Baker If you were doing multiple data sets and suddenly you're pulling in multiple missing data or the formats don't look right or something like that, it might want to pull that and flag that that says, you know, most of my cells are at four decimal places.

35:31 - Fraser Macintyre However, the big hitters that are influencing the insight are only at one decimal place.

35:37 - Rich Baker Is that a problem or not?

35:42 - João Demenech Yes, I see. Yes, Yeah, I think it could flag this sort of, of, Lotus.

35:52 - Fraser Macintyre Yeah, so it's one, yeah, as I say, I don't know, don't necessarily know if it's a, you know, an essential for us, but it's certainly one to consider.

35:59 - Rich Baker Oh, yes, Do we do that as a rounded output, or is the fact that we already do the internal data quality enough that we've got confidence in the products, but then making it available to the public and externally?

36:19 - Fraser Macintyre I guess the same for unintended consequences of, you know, could you identify any sensitive sites or GDPR related issues as well? You you could be quitting it on that as well. Yeah, I was thinking that from a critical infrastructure point of view.

36:41 - Rich Baker Using all these data sets, can you tell me where there are obvious holes in the data? Yeah.

36:47 - Fraser Macintyre What could it be? Yeah. And just following that thread, is the agent purely ring-fenced just to this environment?

37:02 - Rich Baker I'm just thinking, could it cross-reference a map and then go, hmm, that seems to be a donut-shaped building in the middle of the country, cross-referencing that.

37:13 - Fraser Macintyre I can tell that's GCHQ, blah, blah, blah, blah, blah.

37:16 - Rich Baker Or would it just go, look, you know, there's a hole there. That's all the data tells me. Hmm.

37:27 - João Demenech Someone raise your hand. Okay. Yeah. So it's definitely capable of crossing external sources. So if you provided a link, for example, it should be able to read that link. I think it might be disabled right now. I can't remember, but I tried it a while ago. And it can definitely, like, uh, mix the dataset you're looking into with external. Data. But yeah, I think like the example you provided in terms of the building, maybe if we tweak it, I'm not sure. Because maybe it knows that the shape might be, the shape in a certain area might be a specific building. It could search that. Not sure how reliable it would be and so on. But yeah. Okay.

38:35 - Rich Baker What's the underlying model for it?

38:37 - Fraser Macintyre it OpenAI? Yeah, so we're using Gemma 4 currently.

38:44 - João Demenech Yeah, but we also tried, if the model is flexible, we can switch it to pretty much any other model, I would say. Mm Mm Mm-hmm.-hmm.-hmm.-hmm.

39:02 - Fraser Macintyre Yeah, I think it's really interesting. I think just some of the stuff that me and Rich have been saying, I think we need to try and think through some of that a little bit more, maybe as a business before we made it, you know, like that, the national security thing's a good, a good one to try and pick out of it.

39:26 - Rich Baker Yeah, think it's a bit of an internal quandary we need to have a think on, isn't it, around, you know, not just what's the benefits of it, but also what's the unintended, you know, I'm thinking a little bit like, Nick, you might not remember this, when we were doing some of the cable data and we published that and we thought this had been really helpful, and actually then there was a DES-NES response that came back saying, pull the data immediately, you're, you know, inadvertently, you know, giving away, you know, critical infrastructure, and it's because we were doing overhead lines. And on the ground lines, that's why we had to pull the data. Completely unintentional, but yeah, if you've got the right skill set, know, putting some of these data sets together, you know, you could give yourself quite a powerful view. So yeah, we may need to take some of that away. Yeah, and I suppose, I think we said on the last showcase, this is quite sort of new, uncharted ground for us. So we haven't really got this, well, anywhere. So it does feel like it's also a centralised, like, policy, security, legal, you know, all those sorts of groups. So I just now sort of need to sit and have a think about, you know, what is and isn't acceptable and that sort of thing. So I just feel like we just need to internally create some sort of guardrails. And I'm wondering if we just need to do some scenarios as well, how far can you take? This, you know, that sort of thing. Almost put a baddie hat on for the day and try and not use it as a responsible user, almost do the other that says, you know, it's a little bit like when you're doing systems, isn't it? You know, you love those people that do user acceptance testing, but then there's always that one person that goes in and says, right, my job today is to break it. ACTION ITEM: Email Zoe + Michael re: AI guardrails/governance; then run bad-actor scenarios - WATCH: https://fathom.video/calls/710631805?timestamp=2485.9999 Oh, OK. So, yeah, I think I'll have that. Sorry, I'll have that. So I think next steps for us is probably to link in with, you know, Zoe and Michael and others and just say, OK. What's the next steps internally? I think we'd said on the last, last bit, you know, we need to do some internal wrangling with, you know, what we, what exactly are we trying to get and build and use and, you know, what's the sort of value that we want to give to our data? But yeah, it's certainly, it's certainly impressive, think my head, it's sort of positive feedback, I think, is it's impressive, but it could be, yeah, it's almost so impressive, we need to think, yeah, let's not get excited, let's just make sure we're really careful with it in the first place. So we might need to think about data sets and just what we're, what we're putting in that space. But that's a, you know, that's a, that's a positive for you guys, which is like, yeah, it looks, it looks powerful.

42:33 - João Demenech Thank you. Yeah, and I think now I'm much clearer on the sorts of work rails that you would like to have in place. This was interesting, thank you. Yeah.

42:44 - Rich Baker I suppose another question, again, without getting solution mode, but this might be a bit of a, a bit of a bridge between the two. So, so if we, if we had like a sandbox version of the portal, so there was only. available to internal SSE people. wondering if that would be a way of doing it so we could at least test it. It would be like a non-prod environment or it would be like a beta version where it's like it's real to all intents and purposes from an SSE point of view, but it's not publicly facing. So at least if we did find anything, the risk is reduced.

43:24 - João Demenech Yeah, maybe the current test environment could be used for that, but I think you would want to have the same data as on protection.

43:41 - Rich Baker Yeah, that's what I'm is almost, could you have like an instance on an environment that's like the portal, so the landing page will look the same, but it's got the additional attributes of having the AI agent in it. It would point to the backing. It of all the data sets that are available on our portal at the moment, so it could pull the data through and do it all, but actually the insight and the views that it was saying aren't publicly available and wouldn't leave an imprint of query public level, and then almost once you've finished your session, that's it, you know, you can't, you you might be able to download it to an SSE asset, but you couldn't do anything else with it. Just to give that separation of public versus internal testing type piece in the short term. Yeah, I see what you mean.

44:35 - João Demenech I think this would be straightforward. I assume it would be straightforward. Yeah.

44:43 - Rich Baker Okay. So I think in our world, that's probably where we would want our data owners, data stewards to also have a play with it, because they're the ones that are most intimate with their data, and they're the ones that certainly when we triage the data say, that's got, you know, that's That's got the intentions of this, or when we triage it, we need to make sure that's behind it, you know, password protected, or that can't be open data, but it can be shared data. I think it's just about, you know, is there anything on that at moment we've deemed as open, but actually if you stick it all together and do some analysis and insight on it, suddenly it becomes recategorized as shared or, you know, it shows you a view that we never realized it could show you before type thing. So, one, just one, but, you know, as we get down the road, you know, what, you know, how would we, how would we give ourselves confidence before we unleashed it publicly? Cool. Okay. Now, Fraser, have you got any more sort of questions on that or, or anything?

45:51 - Fraser Macintyre Yeah, no, I think that, that sandbox environment could be really interesting. As well, and I think, yeah, if we could record the sort of queries that are going to, yeah, do you get any analytics on the queries?

46:13 - João Demenech You mean like, for example, most asked questions and so on?

46:18 - Fraser Macintyre Yeah, I mean, like, I think that could be quite useful because it could inform, like, if there's a lot of interest in a specific geographical location or something like that, or a particular data set that's getting queried the most as well. That could be quite interesting for us to see.

46:38 - João Demenech Yeah, this should be possible, and it's something that we want to deliver, along with the product. Yeah, I don't have right now, like, a dashboard that I could show you. Yeah, yeah. This is, yeah, this is part of our vision.

47:00 - Fraser Macintyre Okay.

47:01 - Rich Baker Yeah, because one of the things we were saying before was if we, if you could, if you could almost see like a, you know, view of all the queries and see the most frequent ones where, you know, if everyone's coming onto and using it and saying, right, stick data set A with data set B, and then just doing that was almost a, rather than burning through tokens and everyone, you know, doing that same thing over and over again, you could use those queries and almost sort say, well, people keep asking for it. So why don't we just provide that view? Yeah. I mean, to keep creating it themselves. And as I say, you know, you can imagine like some, some data sets that are monthly, people are coming on every single month, multiple users doing that. You know, I might, there's a view of what we should prioritize and make available rather than people having to create themselves. Yeah, exactly.

47:56 - Fraser Macintyre Okay.

47:56 - Rich Baker It's an internal one that we need to, to just work through. But I think we were starting to land on last time we showcased was around, maybe we shouldn't have users using this to create data, but the data should always be the data. This is more about the insights and how do you use this so that the data is telling a story. So again, it feels a little bit like if people are using it to go make this data, give me a bespoke bit and using it as a really blunt analysis tool. You know, I don't know how to use Excel, so I'll just get AI to do it myself. That's probably not the greatest value, but, you know, if there's a need for it, we shouldn't look it out. Cool. Okay. So yeah, I think next steps is probably with us to sort of determine… you know the what and the why and the how and the you know as i say like what what does it what does it require it feels like it doesn't need a bit of security lens on it a bit of a legal lens on it um you know how do we have sort of confidence to it are we looking at very specific data sets in the first instance and therefore we might lock some of them out that are on available on the portal just because of what they are or actually are we saying anything on the portal is okay because we've already been through triage don't don't know um think we need to have that debate it's great okay no thanks again as i say good good showcasing him yeah thank you cheers and thank you for joining and for being here um okay um what

50:00 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) We have a small amount of time left, 10 minutes, so Rich, what we can actually discuss further. The pen testing, pen testing fixes are in progress, and as we agreed, we do expect that most of the items are going to be… I will jump off, thanks. Thanks again for your time, and most of the items we do expect to be fixed by the beginning of the next week. ACTION ITEM: Fix DataPusher for file uploads; notify Rich to retest w/ Oliver - WATCH: https://fathom.video/calls/710631805?timestamp=3028.9999 Okay. Yeah, yeah. So, Leo?

50:39 - Leonardo Farias Yeah, just I think that I found the reason why the file wasn't uploading on the issue shared on the last week. I'm working a fix for it right now, so I'm basically testing it now, but… What I found is the file was being uploaded, but the DataStore wasn't pushing the new changes. So the file is updated, but the DataPusher didn't push these changes. So that's the thing that we need to fix for now. So I believe it's a fix that we can do it today. And tomorrow you can test it. Okay. Okay.

51:35 - Rich Baker Yeah, just let us know when you think you've done it, and then we can feed that back to, I think it was Oliver that was struggling. ACTION ITEM: Review hero-section link; approve for prod - WATCH: https://fathom.video/calls/710631805?timestamp=3096.9999 So yeah, we can feed that back and say, have a go now. Yeah.

51:47 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) And also there is an update regarding that link on the hero section, if you remember, on the homepage. Okay. Yeah, yeah. Yeah. That's, that's, that's, that's That's It's ready for your review. If it works as expected from your side, you should get approved. We will push you to production. Okay, great stuff. Yeah. So that's ready as well. And Leo, do you have any comments regarding the pen testing fixes?

52:19 - Leonardo Farias No, I don't have any new updates regarding that. Yeah. ACTION ITEM: Schedule Vijay pentest closure review next week - WATCH: https://fathom.video/calls/710631805?timestamp=3149.9999

52:25 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) But then we said we will stick to the timeline and the roadmap we actually set. And yeah, we will manage to, we're still really highly confident that everything will be fixed due to every time. Okay, okay.

52:40 - Rich Baker So just thinking forward then, if we think all the items on the pen testing list will be done by beginning of next week, does, does, this time next week feel a little bit too early to run through and almost close it with, with Vijay? Or are you? guy's sort of confident that we could do that. I'm just sort of thinking, can just get a VJ-ing for one last time, show him that we've done everything, and then he goes, yeah, great, okay, I don't need to, he said, yeah, okay, should we try that?

53:14 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) He sort of joined, he, and then I had to… Yeah, we were in the middle of it, yeah. Yeah, okay. Cool, we'll do that. Yeah, great stuff.

53:22 - Rich Baker Cool. And then on the subject of next week's session… I actually saw your email, but I wanted to discuss with you now, this point.

53:34 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) So what's your actual suggestion? Because you mentioned that you asked us to move it a little bit, or rescheduled. ACTION ITEM: Propose next week's sync time to Rich (1h earlier or after 3 pm UK) - WATCH: https://fathom.video/calls/710631805?timestamp=3219.9999 Is it about time slots only, or the whole day is not working for you? No, yeah, it's just literally the time slot.

53:50 - Rich Baker Of all the meetings I've got in the morning, the one that's come in that I can't move is literally the same time as our regular. So if we could either do it… yeah. You An hour before, or any time after, any time after 3 o'clock UK time, I appreciate you guys a bit further ahead, so if you want to do it any time.

54:15 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) one hour earlier or after 3 p.m. your time? Yeah. Okay, let me check the calendars for everybody and I'll make sure to propose the solution that works for all. Okay. Okay, always. Cool, cool. Nice. Thank you for flagging this. Good to know. So if I understood you well for the next four weeks, we will not hear a lot from, or at least we will not see Jeff, right? Okay, cool. So if you any concern or question, we will definitely feel Yeah, just come to me.

54:49 - Rich Baker Yeah, so I'm Jeff until Jeff's back type thing. Okay. I'm both Jeff and Rich, so yeah, a little bit busier, but it's okay. ACTION ITEM: Log minor portal UI issues in GitHub for Datopian - WATCH: https://fathom.video/calls/710631805?timestamp=3295.9999

54:59 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Cool stuff. stuff. Cool Cool Thank Okay, cool. Do you have any other questions for us?

55:06 - Rich Baker No, I don't think I do, to be honest. There's a couple of, just as a heads up, they're not major. So when I do get five minutes, I'll just put them on the board to have a look at when we get timed. There's a couple of little bits and bobs I've found on the portal that we just need to have a look at.

55:23 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) But they're not major.

55:25 - Rich Baker Everything works. It's all just like either, can we just make it look a little neater? So one of them's like where a button's got a word on it, but the word's so long, it's almost like stretching over the button. And then something I was looking at this morning when I was doing some persona work, I think when we've loaded them in, say I'll put it on the board, but when we've loaded them in, I think we've done it as a table with white text, but then because it's on a white background, you can't see the text on the headers. So it probably just needs either the text changing color or like the cells just put in as a non-white color so we can see the white font on the white bits. But again, nothing major, but it's just about needing help. And if anyone wants to use it, they might not realize it's a table with headers.

56:19 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Okay, just feel free to log them in the GitHub and we'll make sure to take a look at this. We will definitely orient our focus mostly on the pentesting fixes so we can meet the agreed timeline. And as you could hear, we are preparing these fixes for the Oliver's issue. ACTION ITEM: Email Michael re: SOW status; expedite approval - WATCH: https://fathom.video/calls/710631805?timestamp=3404.9999 Yeah. Yeah, that's a good thing. Okay. For the rest of time, I think we are steady. Any particular updates on the statement of work? How that's progressing so far? Good question.

56:55 - Rich Baker I'll flag with Michael and just see where we're at. It feels like it's the usual, and while there's no Geoff on the call, getting things through our internal governance at SSE is painful. Hence why we still haven't got a replacement for Ethan either. It's just it has to go through so many different people and get signed off. The distribution has to go to like central and has to come back down. It's over elaborate in my opinion, but process is the process. I'm not going to argue with it, but it is painfully slow. But yeah, I'll have a check with Michael and just see if there's anything that we can hurry along or not something. But yeah, we're all keen to do it. Okay, please don't thank you in advance.

57:43 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) I'm asking this because, you know, also we internally do have some kind of the monitoring process and we are trying to forecast our allocation through the months and to understand who will be actually delegated to which portion of work. It's for us really important as well to have some kind of rough estimation so we can make the correct calculations around people's work. that's the reason why I'm asking. But however, thank you very much in advance for checking this out.

58:15 - Rich Baker Yeah, I absolutely get it. I was having exactly the same conversation with Jeff on Friday afternoon as we did a quick handle and I was saying, know, what we need to be better at with your guys is forecasting the what and the when. You know, so obviously some, some, you know, sort of capacity will need to be spent on, you know, just maintenance and that sort of thing. So like queries like, you know, the Oliver one that's like, you know, got a bit of a thing. So there's a bit of like maintenance in there. Some of it will be, okay, well, you know, what's the resource needed for like this AI bit when we start to do that and all that sort of stuff. What's the bit around, you know, we talked around the LTDS and making it more intuitive and are we working towards November for that? So that will need resource and, you know, research and development and all sorts of stuff. So we've got to be better at stacking all the asks for you guys. So you can put the. It's not people just solely working on SSE, which I think a lot sometimes forget, because, oh, yeah, Datopian can do everything, you know, just send it across, and it's like, yeah, but, you know, there's only so many of them, they've got a small army to do all this stuff.

59:19 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) We do definitely have a people, but for the sake of, let's say, efficient and effective as well, developing and moving forward, for us, it's, as you know, really important to understand when and what should be reserved for. So that's the case. But cool. Once again, thank you very much. No, that's okay. Yeah, looking forward to updating this site. Brilliant. Cool.

59:44 - Rich Baker We've done a full hour this time, haven't we? uh… Yeah.

59:48 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) Although, this was not the case for a long time, so, yeah, for a change. Yeah, I'll complain to it.

59:54 - Rich Baker It's all right then, so, um, no, thanks for, thanks for continued efforts, and, uh, yeah. So, yeah, if you need to, Nick, just bounce some messages back to me if we need to work out when we do this session next week. Yeah, we'll do.

1:00:10 - Nikola Mladenovic (Viderum Inc. (Trading as Datopian)) And, of course, feel free to reach out to us if anything pops up as an open question. Also, if Vijay has anything to ask, I mean, he can definitely feel free to reach out to us and that will be available. Cool. Smashing. Have a good week, Sol. You, too, man. Have a great day. See ya. Thank you. Thank you.

1:00:33 - João Demenech Bye.

Built with LogoFlowershow