May 6, 2026 TDC / Datopian - Queryless demo and AI discussion - Transcript 00:00:00

Anuar Ustayev: um you know about the AI integration possible like AI integration or just share ideas about way AI can be integrated into the uh TDC portal um and also we are interested in just hearing your thoughts like you know what do you think what you know because we already read your email um which raises two like points I would say and you know we'd like to discuss uh them as well. Um I don't know if it will be useful if we start with a quick sort of demo of what we are doing. Uh would it be useful? Nicolas Becker: uh you mean the the application you're working on that's uh very less Anuar Ustayev: Yeah, Nicolas Becker: or um yeah maybe to start a Anuar Ustayev: exactly. Nicolas Becker: conversation it's it doesn't have to be long but maybe I don't know two three minutes of what you you're clicking on maybe also because I'm not sure if Paul already um is aware of this. Um you sent the some documentation to me but uh I haven't forwarded at Anuar Ustayev: Yeah. Okay, then I think it will be useful.

00:01:10

Nicolas Becker: all. Anuar Ustayev: Um, yeah. Dominic, do you do you have the it set up? Okay, great. If you can João Demenech: Yes. Anuar Ustayev: share. João Demenech: Can you see my screen? Anuar Ustayev: Yep. Nicolas Becker: Yep. João Demenech: Yeah. So, this is our demo port JS portal. And here we have querless enabled. You can see on the right hand side this SKI floating button. It's always visible uh regardless of the page I'm looking at. And if I click on that, it will open up the chat where I can talk to the AI assistant. Um, you can see it's aware of the page I'm looking at. So it can better can provide provide better guided responses based on this context. And we can use it use it both for searching for data and also for querying data. So to get started, I'm going to ask it about um data sets in this instance. So I'm going to ask are there any data sets about happiness? Uh I know for for sure that there is a data set but just I just wanted to retrieve this data set for me and behind the scenes um is using the CCAN API to uh retrieve this uh information.

00:02:33

João Demenech: So it responded with saying that there is one data set which is the world happiness data set 2020 and as you can see it uh provided a link to the data set. So it knows the structure of the website so that it can provide a proper links. So yeah now I have navigated to uh the data set. I could look into the metadata uh preview the data here and then I can continue my conversation from where I stopped. As you can see it it already knows that I'm browsing this uh resource. So now I'm going to start asking questions about the data itself. We can start with a very simple one. Uh what are the top 10 happiest countries? And we can verify this answer easily if I sort by the letter score column. Uh so yeah it's fetching the data and yeah so Finland, Denmark, Switzerland. Yeah, as you can see it provided the right answer first as a table here and then it followed with a bar chart as well which is interactive.

00:03:48

João Demenech: Um next I could ask something more complex. So maybe uh what is the average happiness perception across all countries? So now it has to do uh an actual aggregation uh by the letter score and have 5.47. Uh we we have some columns here about um factors that contribute to happiness perception. So I could ask for example what factors contribute the most to happiness across all countries. Yeah we have social support healthy life expectancy freedom generosity and so on. Yeah. Uh also responded with this table here and a bar chart. Um, and yeah, it can generate all sorts of visualizations. Uh, it uses Vega behind the uh behind the scenes. So it could generate pie charts, line charts and so on. Um and it it also provides insights along the the charts. So yeah, for example here, uh having strong personal relationships and community support matters more than wealth. Um, I could ask it to generate sharable reports as well because these are just in the the chat window, but I could ask for uh can you create a sharable report?

00:05:27

João Demenech: Um, yeah. And then it should generate a report that is actually hosted somewhere and I can use it to share with other people. Let's see. Um yeah now let me generate an interactive churable report takes a few seconds and yeah well while it's generating uh just to mention it's capable of doing cross data set analysis so it can join data from different data sets to provide answers Um it also supports different languages. Uh I could ask questions in other languages. Uh okay. So now we have the sharable report. As you can see it's a link. If I click on that um there we go. So it generated this bar chart uh with the countries, this scatter plot with regions and yeah some that table here and so on. So I could also ask it to tweak this report. It would regenerate it with new charts or changing charts that are already there. Um, yeah. And yeah, I think I think it's that's it for a brief demo.

00:06:56

João Demenech: Do you have any Nicolas Becker: I think Paul had one in the chat but got al already answered right so which kind of tool João Demenech: questions? Nicolas Becker: okay uh thanks a lot thanks a lot um I think it's very interesting Anuar Ustayev: Yeah. Nicolas Becker: um one one question I have is it um possible to trace the process behind the the the generation of these charts or there is there any transparency layer in João Demenech: Yes. Um so I could ask for example how did you calculate the um Nicolas Becker: there. João Demenech: average happiness per reception um because behind the scenes it's actually creating a SQL query and running that against the data. Um so here it can explain uh what it did. So it's uh doing an average operation and using the world happiness data set.c CSV data. Um so yeah it has this transparency. You can ask it how things were calculated. Nicolas Becker: Okay. But we need to actively prompt the model to to explain what it did. It's not that it's automatically uh generated the back end.

00:08:13

Nicolas Becker: Okay. Anuar Ustayev: Yeah, I think uh the point here is that this is uh uh this is for demo purposes, but uh obviously we could like automatically show this right or we could hint that you know you can expand and see what it's doing. I mean that that is doable. Um just um just to be um clear here what we are doing here uh we're not developing LLM or anything like that. In in this particular demo we are using JAMA 4 uh which is a open source uh model from Google uh but it can use any LLM um depending on you know what is your preference what is the project's preference and so on. Nicolas Becker: Thank you. And another question I would have is at the beginning you you started your prompt with a question or something you know that the database already contains this this happiness state table. How's the mechanism when you ask it something that you know is not contained in the portal? Will it hint you to something similar or will it just tell you that this is something you feel happy with

00:09:24

João Demenech: Yeah. So it it should say um well the process behind the scenes is it's creating Nicolas Becker: or João Demenech: this API query with relevant keywords uh and then it's it's providing a summary of the results. So if I ask for something that doesn't exist, it should just say that there isn't anything like that in this portal. We could try it with um for example I know for sure that we don't have transport related data. So I could ask um are there any data sets about transport and then it should say that there isn't any data set related to this topic except if there actually is but I don't think so let's just uh see it here. Uh so no transport data sets here. I'm afraid this portal has seven data sets focused on these topics and provided uh links to the groups as well. Yeah. So it it got it right Nicolas Becker: Okay. Anuar Ustayev: Yeah, Nicolas Becker: Yeah. Anuar Ustayev: but I think it's not it's not like 100% guaranteed. João Demenech: here.

00:10:31

Anuar Ustayev: Never. Uh but we can I mean it's it's possible to tune this right so that it it you know for example it only uses the existing catalog. Um but um the I the idea is that you know all it does I mean what we want at least is that it helps the user to use the catalog existing catalog and the data sets um like you know in a better way. So we know that um for example search is not always returning you know right answers and so on and we want to use LLM to improve the discoverability of the data sets. Um and then also improved how the uh how people use the data sets like how they understand the columns the schema uh or how they even do some data analysis and so on. So it's just more like a you know assisting them to build queries um or you know to generate SQL queries and then get some subset of the data and so Nicolas Becker: Yeah, I think I think that's one of the points I see clear value in in integrating doesn't have to be a full AI solution, but I don't know at least semantic search is something that can enhance the the findability of data sets because I know there are cases where users looking for vehicle data but for some reason they're called car and and the description and I'm not sure if in our current solution the user would be directly led to these kind of entries um in the case that of course the the

00:12:11

Nicolas Becker: standardized uh um terms for would want to promote aren't used and for the upload of the data Um but then yeah again thanks a lot for this demo. I think this is cool what you're doing. Um and it's somehow something that's evolving uh naturally. So it's not the first time that uh that we've got approached asking um if there's any kind of AI we want to any kind of AI integration we want to use into or we want to add on to our current portals. So um yeah, I think uh it's very close to what we've uh what we've been discussing before. So what um Paul just mentioned in the chat is that actually one of the the fun uh funders of the portal development two years ago is currently working on a similar solution that tries to um implement an AI agentic solution that covers the entire sector space in terms of data and information that's available out there and it sees TDC as one of the João Demenech: Oops. Nicolas Becker: possible um sources of transport data in and the standardized format or at least for the um uh structured data sets.

00:13:32

Nicolas Becker: Um yeah, and this is why we've been already confronted with these kind of questions. What is what is feasible? What is what makes sense uh in this area? Um, so I think this this is something we could discuss about uh what is currently portal able to do, what what are possible expansions in this area, what makes sense and what is really um something we can um do without concerns about the credibility of the of the of the responses and of the outputs. But uh I see the Paul also has his hand up. So maybe you want to add to this. Uh you're currently muted. João Demenech: Anyone going to say Nicolas Becker: I think Paul wants to see something but maybe the browser is not allowing activation of the João Demenech: something? Nicolas Becker: microphone. Sometime have sometimes have these issues also when using Google. Anuar Ustayev: Yeah, I mean in the interface it's Yeah. Okay, cool. Yeah, while Paul is rejoining um based on your email, uh just to flag that the TDC portal already has things like you know API access, right?

00:15:03

Anuar Ustayev: It also has Okay, pause back. Paul Natsuo Kishimoto: Sorry, is my audio working now? Nicolas Becker: Yes. Paul Natsuo Kishimoto: Great. Yeah. Anuar Ustayev: Yep. Paul Natsuo Kishimoto: Um, so, uh, yeah. Yeah. Yeah. Um, so, so yeah, I think the context is we probably are not going to focus resource or I would argue strongly that we don't focus our resources for TDC proper on on AI tools. Um but this this UK government funded TGIS project is very much set on that, right? And so I think what we one thing we can convey to them is that if they want to use TDC and its contents which are backed by SECAN as an input for whatever tool they're building, it probably makes sense that they go through this queryless uh uh tool, right? And the reason is that I'm sure that you guys as as like you're a CCAN shop, right? So I'm sure that you've made sure that it is correctly interacting with the CECAN API and and so on, right?

00:16:14

Paul Natsuo Kishimoto: So if if whoever is contracted by that TG TGIS project to uh if they were just to like vibe code something that tries to interact with the CECAN API, it would probably do a bad job of it, right? So rather than waste time and resources on that, I think it probably makes sense that they they just go through that route, right? Um so I think that's something we should suggest to them. Anuar Ustayev: Yep. Paul Natsuo Kishimoto: Um the uh so but the the model that I would prefer is that they do it on a separate site, right? So as I I imagine that like I noticed that that you like based on the text that was coming out of the system that it is aware of you browsing to different pages on the data set. So it uses that as a context window or something like that, right? Um but I imagine this can also be embedded elsewhere. Um like not like if we're not hosting it on our transportdata.org, it could be also just sat in some other web page.

00:17:19

Paul Natsuo Kishimoto: The other thing I would also be concerned with is um the the responsiveness of the TDC portal is already not the best. Um so I would wonder you know I guess it's possible to like snapshot or make a mirror or something like that of the of our SECAN content. So even if they are cuz like I know I know when you use AI tools with APIs they can hammer those APIs because they they have no sense of like I mean they they may just decide to make thousands of requests because that you know that's that's what they choose to do. So um I would not want that to degrade the performance of the TDC portal as experienced by like actual people. Um, but I guess it would be possible like to say how like let's have a mirror or a duplicate or something of the API and then their query list instance or whatever could talk to that or talk like look work off of a dump or something like that so that the ones that our users are actually using uh is not uh is not affected.

00:18:30

Paul Natsuo Kishimoto: Um but is is that is that roughly those are things in the realm of possibility or Anuar Ustayev: Yeah, Paul Natsuo Kishimoto: uh yeah okay um that's good to know. Anuar Ustayev: definitely. Paul Natsuo Kishimoto: I think the other thing would be about uh pricing. So I think if we were going to recommend that to the TGIS people that they take that route. Um like I I understand it's connected to um whatever whatever LLM they would choose to use. So I guess they would be paying the bill for that directly. Um but if you had information on what is the like cost structure for um using that that tool and service uh we could pass that on to them, right? Um and uh yeah, I think those are those are my only reactions to uh to seeing it so far. Yeah. Anuar Ustayev: Okay. Yeah, we can definitely share the model uh the pricing model um because there is few options. One is yeah as you said bring your own key and another one is we could provide some open source models um hosted versions if needed but it's yeah it is flexible.

00:19:45

Anuar Ustayev: Okay. Any any other points or questions Nicolas Becker: Yeah, I was wondering in general how how do you think is the current see Anuar Ustayev: because Nicolas Becker: API that we have implemented prepared for these kind of uh requests by agents or I don't know doesn't have to be this TJS that Paul mentioned, but in general any I guess if anyone using CHPD or any other solution asks for transport data at some point they will be directed to to transport data comments. Um yeah do you see any any kind of support there on on from seem side centrally that's the requests that are sent to the API are actually validated somehow that's models get some sort of guidance how to use it there's this MCP uh conversation going on or at least emerging um as a technical solutions of how to manage this this kind of conversation between portals or tools and and models Uh, Anuar Ustayev: Yeah. Nicolas Becker: what do you see Anuar Ustayev: Yeah. Generally speaking, um you know, all these AI um or AI agents, Nicolas Becker: there? Anuar Ustayev: what they're doing is they're generating code, right?

00:21:01

Anuar Ustayev: It can be Python code, it can be some bash script and then um you know, it's about just you know executing that that code and as Paul just wrote in the chat, it's just like another Python script or another some you know, some programming language script that then runs somewhere. uh and hits the seeant API uh and I think these LLMs are getting better in terms of understanding how to use see API right and uh you know we don't even know which one is actually being executed by a human or by AI agent and so on you know we can only see that um the seeant API is being called um yeah and that's that's all that we can we can see But we're definitely noticing increase in general like not not on TDC particularly but in general on all data portals that the traffic is growing. We're seeing that more data is being uh sent like egress for example is growing um in our cloud and so on and uh we don't have a um you know a data to prove but we can we we believe that this is this is you know related to um all this AI um related you know code that is you know now it's very easy and quick to generate and then um our APIs are open to the entire internet right that they can they can call it without any authentication.

00:22:32

Anuar Ustayev: It's publicly available. So Nicolas Becker: Thank you. Anuar Ustayev: yeah. Nicolas Becker: Um I think Paul just posted a related question if if there's any safety mechanism to prevent if yeah an API call is executed so much it's Anuar Ustayev: Yeah. Nicolas Becker: um has an effect on the general um function of the Anuar Ustayev: Yeah. So on like that's not really seeant specific generally we obviously we do Nicolas Becker: portal Anuar Ustayev: have I uh on a proxy level on a uh load B we have we use load balancers and we have rate limits per IP uh it doesn't really matter if you know what who's the because in most cases we don't know who's the user you know it's anonymous calls and we just make some limits per IP and it's I think right now it's quite generous uh But when we see some issues like you know we start noticing when for example current resources are not enough and the you know our system starts scaling up automatically but we receive notifications that you know some sudden spike in in demand for example and then we start looking and usually when we do some analysis we can see that here's like five IP addresses that are really abusing the service and then you know we discuss like normally we will let you for example that we want to block these IPs and so on.

00:23:55

Anuar Ustayev: This is how we approach Paul Natsuo Kishimoto: Okay. Yeah, that that's really good to know. Thanks for that answer. So yeah, I I guess I again I would be concerned as these TGIS people or their contractors start to work that they would not be careful about that and and I would want to avoid that they like in in in theory they're collaborating with us but if they inadvertently DDOS us in in building something that's supposed to be complimentary to the TDC that's not good. So, it's good to know that you guys have those practices in place and that we would get some notification uh before the portal Anuar Ustayev: Sure. Paul Natsuo Kishimoto: went down. Um yeah. Um yeah, guys, I'm sorry. I really have to I have to run to another meeting, but uh I appreciate the the answers to all our questions and I know Nicholas can follow up with you. So, ciao. Anuar Ustayev: Thank you. Nicolas Becker: Okay, Anuar Ustayev: as just as a quick follow you know as a next steps Nicholas I think uh we need to send you an

00:24:45

Nicolas Becker: Paul. Anuar Ustayev: email with some pricing model for this so that we can share further right um Nicolas Becker: Yeah. Yeah, that would be great. Yeah. As Paul said, currently um yeah, from CDC side, we do not have neither the the mandate nor the funding to to work on any of these questions. Um it's all just really peace work. And as I said, the only thing I was thinking that could be maybe within our reach would be to enhance the search function somehow. Um but there are these initiatives and people willing to to invest in in these kinds of applications. Um yeah, so I can I can happily set up a contact there in case they they decide to go with the TDC as one of their uh sources of data. Um and yeah, all all sort of information that you sent me, Anuar Ustayev: Yeah. Nicolas Becker: I can follow up and maybe can also set up a call so we can aware of what they're planning to do. I think they're sending a first prototype these days of um of the text that they're planning and yeah we're trying to to to position CDC somehow as a one centralized uh portal to access all of this data and to also manage access to the data and I think that's that would be an

00:26:02

Nicolas Becker: opportunity there. Um but in general the scope is a bit broader. They also want to integrate the geospatial data sources from different um yeah locations. Um but yeah thanks a lot for the demo for the information. I think this is very as I said very interesting very very cool what you're doing and yeah somehow natural next step I feel as well although um yeah from TC side we naturally can't we unfortunately can't uh can't go on this on this end Anuar Ustayev: Yeah, to totally understandable. Nicolas Becker: currently Anuar Ustayev: Yeah, the the main reason why we're doing uh this is to share like so that you are up to date of you know what we're doing uh around seeen and data portals in general. Uh as you can see we're just trying to adapt to you know to all the changes in the world and uh yeah feel free to reach out to us if any need arises. Um yeah, Nicolas Becker: Okay, thanks a lot. Yeah, Anuar Ustayev: thank Nicolas Becker: I think that the conversation is more about uh what implications do these uh do these Anuar Ustayev: you. Nicolas Becker: developments have for the portal from the user side from from our side as as providers. So I think it's good to keep this channel up. Anuar Ustayev: Yeah. Awesome. Nicolas Becker: Thanks. Anuar Ustayev: Thanks, Nicholas. Byebye. João Demenech: Thank you. Nicolas Becker: See? João Demenech: Bye.

Transcription ended after 00:27:30

This editable transcript was computer generated and might contain errors. People can also change the text after it was created.

Built with LogoFlowershow