LLMs: What Data Viz Teaches Us – Ep.557
A model can hand you a pixel-perfect chart and a cheerful note that the work is done. Tommy Puglia and Mike Carlo spend this Explicit Measures episode on Eva Sibinga’s Nightingale essay, and on what data visualization already knows about trusting that chart.
News & Announcements
-
Using AI Harnesses to Harness Microsoft Fabric — The Chicago Fabric / Power BI User Group meets Thursday, September 24, 2026, from 3:00 to 5:00 PM Central at the Microsoft Technology Center, 200 East Randolph Drive. Tommy Puglia will demo the two-harness approach: a second brain that holds the instructions, plus MCP servers that execute against Microsoft Fabric, including which models to use and how the two loops talk. The rest of 2026 is already on Meetup if you want to register ahead.
-
Power BI August 2026 feature summary — Tommy Puglia framed the August 21 summary as the month of things Power BI should have shipped already. Doughnut charts can now place a number, text, and an image in the center, the layered native-visual workaround Mike Carlo published on October 8, 2019, and matrices can freeze row headers and expand or collapse column headers, the item he put on the gold podium after Zoe Douglas posted it. The same update adds single-date slicers, mobile landscape rotate and export to Excel, and SharePoint embeds that pick a workspace or one visual, while modern visual defaults go generally available with a gray page, white visuals, and rounded corners that Mike likes, separate from the on-canvas title boxes he still dislikes.
-
Bringing agentic warehouse development to Fabric Data Warehouse — Notebooks, semantic models, and report design already had an agent path; full warehouse work — understanding objects, running queries, diagnosing failures, judging performance — did not. The Fabric Data Warehouse MCP server gives agents a governed way to run T-SQL through a global endpoint with Microsoft Entra ID, with no local server to install, plus skills for authoring, consumption, and operations. Mike Carlo, passing along a count Matias had agents research, put Microsoft’s Fabric and Power BI MCP servers at about eight, and the new problem is knowing which one to turn on: inside VS Code the Power BI modeling MCP shows up as an ordinary extension, so the agent window misses it and you add it by JSON.
Main Discussion
Topic: Critically evaluating large language models, and what a chart will not forgive
Tommy Puglia walks Eva Sibinga’s Nightingale essay as a data visualization problem, not a Fabric product review. Sibinga’s two charges are that hallucination is built into generative models, and that Claude, Copilot, and ChatGPT are products whose design favors appearing helpful over being accurate. Her questions sort into accuracy, security, and sovereignty: is it true, is the data safe, and am I using the tool the way I mean to. The hosts accept that diagnosis. They spend the rest of the episode on the remedy that matters once you have a semantic model, an MCP server, and a harness.
-
Cheerful completion is the failure mode. Sibinga asked a model to compile a public table. It returned a polished spreadsheet, dropped columns without a flag, and filled in values that were wrong more often than they were right, in the same pleased tone it uses when the work is real. Tommy connects that to earlier conversations with Nicola about code that looks finished and is not. Mike agrees the risk never leaves. Every harness he trusts — memories, skills, the pre-prompt he never sees — exists to shrink it. His line is that a person who cannot test the answer will see more of those failures.
-
A TMDL file got him most of the way, then the engineering started. Someone handed Mike a single TMDL and no data. He pointed an agent he calls Luna at it, with no MCP servers and no skills attached, and asked for tables, columns, measures, and relationships. The first pass landed in 30 to 40 minutes at about 60 to 70 percent. He then spent about three hours moving tables and refining. He still wants someone in the loop who can read the SQL and ask why it was written that way.
-
Name the framework, and build the table before the chart. Tommy’s pushback is that reading a model file is a different job from designing a report, and Sibinga’s example is a model that invents data and then charts it. Mike’s sequence is the one he already uses in Power BI: drag fields into a table, vet the numbers, then turn them into a visual and a story. He builds the design template before real data arrives, and he refuses to throw an agent at a blank canvas and say “build the report.” Tommy changes one word in that claim. Frameworks the model can know become frameworks it must be given — Vega-Lite, Power BI Desktop Bridge, or a visual spec aimed at agents — as part of the prompt, with the skill that goes with them.
-
Tommy’s three legs are MCP, skills, and context. For work that costs money he wants an MCP server when one exists, even when setup is a headache. Skills are the second leg: his Skill Vault pushes them to Claude Desktop and Claude Code and syncs them to Notion, and on a large project he tells the agent to list the relevant skills before it starts. The third leg is context with the schema and the tests attached before exploration, so he is pointing at instructions in Notion instead of writing a giant prompt. Miss any leg and, as he put it, the table wobbles.
-
Mike hears a harness in those three legs. Skills and MCP servers are pre-baked context, so you stop re-explaining the API on every call. A report can still come out without them. It will cost more tokens and more checking, because you are reinventing tests other people already wrote. Skills for Fabric changed report authoring for him in a way he could feel with them and without them. Some of those skills run past the 400-line length he treats as guidance, which he reads as a team that is not living inside a customer token budget: Microsoft can point staff away from Anthropic and still have them in GitHub Copilot all day, because the company owns the hardware.
-
A aligned chart still needs a click. Tommy’s point, via a meme of a chatbot agreeing that Bilbo should keep the ring, is sycophancy on a semantic model. Filter context, evaluation context, and many columns sit behind one visual. The layout can be pixel-perfect and wrong the moment you interact with it, so validation is the last step of the design, not a polish pass. With Power BI Desktop and Fabric skills in the loop, Mike is getting builds that stay on the brief. He treats a missed one-shot the way he treats search: if the answer is not on the first page, the query was his, and the variable that changes between a strong result and a weak one is the person wielding the same model.
Prototyping is where the idea has to survive
Sibinga cites a prototyping essay whose line Tommy reads aloud: prototyping is where you test the intellectual rigor of an idea. AI applied at the idea stage can make a shaky concept look sleek. The better opening is the zone where the idea has already proved itself and the missing piece is skill or capacity. Mike maps that chart onto how he builds apps now. One agent refines the idea. Another produces UI only as components, in Fluent 2, from references he already likes. He spends one or two turns in low fidelity and many turns polishing high fidelity. He is acting as a program manager. The engineering is still his. He is not typing the code.
Tommy’s close puts the same three legs under Power BI work: Desktop Bridge is a strong road, and it still needs an MCP server, Microsoft’s skills or ones you write for your own organization, and real context. Mike’s close is plan mode. Spend more tokens up front on a stronger model, let a cheaper model execute the plan, and when the output misses, improve the harness and the prompt.
Looking Forward
On the next visual, name the framework, point the agent at the MCP server and the skills for that job, and hand it the schema and the test before it explores. Spend the plan-mode tokens on a stronger model, then let a cheaper one execute. If the first pass misses, treat that as a harness and a prompt to fix, and keep the engineer in the seat through low, medium, and high fidelity before anything is called production.
Episode Transcript
0:02 the mix. Fabric and AI get your fix. Explicit measures. Drop the beat now. Kings feel the crowd. Hello and welcome back to the explicit measures podcast with Tommy and Mike. Good morning, Tommy. How you doing? I am doing excellent, Mike. The weather is amazing here now. Finally, that crisp
0:32 morning. So, it’s it’s good here. It’s It is good. I will take I’ll take the weather as we can get it. We’ve been getting a lot more rain in our area. We’ve been getting a bit more we never have tornadoes in our area. there are people in the country that have this all the time. So, we can start feeling for you. usually there’s another area around us called Walworth County. Wworth County seems to get tornadoes or warning tornado warnings all the time. We never get anything and for whatever reason this year we’ve gotten two tornadoes
1:02 coming through our little hometown here. we’ve had a couple buildings knocked off, some roofs ripped up, ripped apart. thankfully everyone’s been safe so far, but man, it has been wild and the weather has been crazy up here. So anyways, storms are just rolling through. It’s been a wild year for us already. Yeah, I I we’ve July was just a bad month for summer in Chicago. It’s finally getting better where it was either 95 degrees, which was crazy for July. It was either tornadoes, which we had to go down to
1:33 the basement a few times, or the smoke that we had because of the Canada. It was one that was the three f actual outcomes that we could have. So nice to have the good weather. Things are picking up with work. Things are good. A lot of updates, man. and get to talk to you. Yeah. Let’s let’s go through the let’s go through so before we get into our news for today, let’s do a quick little recap here of the main topic today. The main topic will be an article off of I guess
2:04 it’s like Nightingale Dev data visualization pieces. This is made by Eva, I believe, is the author here. And we’re going to talk about critically evaluating large language models, what data visualization can teach us. Okay, so just a kind of a deeper topic here, large language models and how AI and large language models teach us about our visualization process. All right. Anyways, we’ll get into this one in a little bit here and we’ll unpack this article which I think this is a pretty good article and
2:34 some there’s some I think interesting concepts that we can draw out of this and talk about today. All right, with that being said, Tommy, what news do you have? All right, so our first one is thanks for all who are able to come to our Chicago user group in August. Just letting again we have all of 2026 planned out. That’s actually and they’re all published on Meetup right now which is really cool. We have September, October, November, and December. You can register for all those now. Our next one is going to be September 24th. That’s
3:05 going to be downtown 3M. And the title of it is using AR harnesses to harness Microsoft Fabric. So, this is a little inspiration from you, Mike, when we talk about what a harness is. And it’s off of an article I also wrote just from our conversations as well where really how to actually effectively run AI LMCs in Microsoft Fabric in the best way possible rather than just trying to prompt it. We’re going to go over workflows, which models to use, and even
3:37 choosing a harness to really have, I would say, the most successful way of using MCP servers, getting your data accurate, and really choosing the best solution. And so, we’re going to go over what I call the two harness approach, using a second brain, and then going through with an MCP server and just how those both interact with each other and how to set it up for yourself. Awesome. That’s going to be really neat. this is a really interesting project because Alex put a lot of the fundamentals behind this. He did
4:07 a lot of like CLI and code to help build the structure and then Tommy you skinned it with a lot of well that was our that was attached one. Yeah. Now this this is going to be using MCP servers. so just MPC servers this time. Yeah, MCP servers, but also how do you build context in your second brain to effectively say, “Hey, we want to write a notebook or we want to build a lakehouse actually coming from some set information that you can already set.” Memories or having agent hold on to memories now seems like it’s the next
4:38 wave of development. I think I hear a lot of people discussing in my circles of friends that are also very much into large language models. They’re all like, “Well, the agent doesn’t remember things.” And we wanted to remember everything, but some things. And some things are more important than others. And there’s there’s also this idea, I think you Tommy, you’ve probably stepped in this, too. When you tell an agent to do something and then you walk away and you come back and it still doesn’t remember what you just asked it to do or, “Hey, I told you not to do it that way.” That’s really annoying for a user of a large language
5:10 model to tell it something and then have it just go do the exact opposite thing that it thought it needed to be doing before it didn’t actually like learn what it needed to know how to do, right? And, so we’re going to talk about skills too, but again, I completely agree, Mike. That’s so important where honestly Notion is my second brain and my AI brain. Anything I’m going to do that’s actually going to be implementing or executionwise is coming from instructions and context. Honestly, that is in that second that harness brain. Well, I think there’s a better name. Interesting.
5:40 In my case, notion. So, I am never writing more than two paragraphs in a clawed MCP or an MCP prompt because it’s all I’m just really pointing it to instructions built through notion instructions built with the context rail into a milestone or a project or what we worked on or what to explore and then that writes back. So, it allows that memory to stick. And there’s been MCP servers that do memory, but again, the great part is how do you find something that’s both for the user to have the
6:12 memory and organization, but also for an agent? There’s a gentleman, oh gosh, Stean, I think it’s his name. He said you could say Steven too. I think it’s his butan out of I think he’s out of Prague, I believe. Data Brothers, I think, is the company that he’s part of. he was talking about this as well and he gave a little bit of a word of caution on LinkedIn. people were using Obsidian and giving Obsidian like
6:42 full access to their agents and they had to re recently rework much of their Obsidian vault because the agent was putting in weird and odd stuff into it. And so while I agree with you Tommy, I’m not I’m not debating you. I have I have no I have no skin in this game. Like I’m still trying to figure this part out. I’m also trying to figure out how does how do you carry memories between agents? How do you build a system around your second brain? I think second brain’s a
7:12 really interesting topic here in general. Like how do you have something else think like you think? and how do you train some AI system to help you think or or build things in a system that works really well with you? so that that’s also interesting as well. Okay. Awesome. Well, that being said, let’s go into the next topic. What other news items do we have here, Tommy? All right, we got to go to update one. We got a feature summary. So, the August 2026 PowerBI updates came out four days ago, which I believe was August 21st. And there’s some major updates
7:43 across reporting, modeling, data connectivity, mobile embedded analytics, and developer experiences. So, Mike, I don’t know if you had a chance to take a look at these. There are some good good nuggets here. I didn’t take a look at this one, but I believe Zoe made a note. , so I want to I want to I’m going to I wish I had I don’t have the dream symbol anymore. I restarted my computer so I have the dreamy thing. But imagine a little dreamy sound right now because this goes Yeah, this pulls me way back., one of
8:17 the article posts here was inside the doughnut chart, you are now able to add a data point into the middle of the doughut chart. You can add an image, you can add some text, a number, and an image all together inside the doughut So, I think this is really funny because this was an idea I had blogged about a long time ago at PowerBI. Tips. I had talked about a tutorial that was talking about building doughut charts. I’m going to see if I can pull it up here and tell you like if I can find when it actually
8:47 happened., okay., a filled doughut chart. Okay, you ready for this? We today we just got the release of doughnut charts with numbers in the middle. My filled doughut chart blog post came out October 8th, 2019. Wow. I was building filled doughut charts back in 2019. That’s when I was explaining how to do this and and it was it was complex. I had to have a bar
9:17 chart. I had to have a number. I had to have the donut. I had a whole a bunch of things to like make a bit more of an interesting graphical show about what was occurring with this doughut chart. What was the name of the visual tool that you were investing a lot of time on that essentially you have something built in PowerBI tips too? charticulator. Yeah. It’s still It’s still alive. It’s still around. It’s still alive. But man, I’m going to bring it back. It’s coming
9:47 back. I will know how long listeners we’ve had if they even know what the heck that is who are listening right now in 2026. Charticator. You could build this with charticulator, but to be honest, what charticulator is? Oh, what is it? It’s No, I’m saying Yeah. Oh, well, people wouldn’t even know what it means. What is it? You could build it if what articulator is. yeah. Yeah. No no no I built so the blog post here to be very clear the blog post of this is an actual blog post and it is a real
10:21 tutorial on and a demo that has all native visuals from PowerBI just it’s just a layering of native visuals from PowerBI all in the same thing. But you needed like four visuals to do what I was doing here or three visuals to do it and it was just not very easy. So, ,, not easy to do., now it’s built into the product. Anyway, I just find that funny that now here we are. How many years ago that, Tommy? That was 19.
10:51 Seven years ago. Seven years ago, I was building things that we are now just getting some visual style updates for. Well, one thing I will note though, Tommy, in this blog post, the amount of really rich visual and visual reporting things that are coming out, a lot of GA items in the Great. I love the fact that we’re kicking out so many GA items around the reports and the visuals and features and things that we should have just been h like these are just should have been happening. Yeah, this is this is their so that that
11:22 one was well. So that’s that’s one of my
11:24 main feedbacks here. What do you think, Tommy? In this list, what did you find that was relevant? Speaking of things that should have been happening, Mike, let me run down a slight list here or a small list of things we’ve already should have had and you let me you tell me what are the most important here. Give I’m going to tell you things of things that probably should have been out there. You give me a podium, bronze, silver, or gold here on how important they are. First, date picker for slicers supports single date mode and clear selection controls.
11:54 Again, that’s actually available now. You mentioned the doughnut chart comments and or apps are now actually available. So, that was not previously edible available before. Okay. Matrix improvements. You can now freeze the freeze state for row headers and expand collapse column headers. Okay. Gold podium. I think I think I think someone I don’t care what other features you I should have known. Yeah. Whenever you throw I actually the reason I say this because Tommy I
12:24 actually pulled up the post from Zoe Douglas the one who announced this on LinkedIn as well. Yeah. This this is a winner to me. love this feature. This should have been out a long time ago. The Matrix has needed some love for a long period of time. I’ve had people complaining about why can’t I collapse things on the on the column level? I want to have it. How do we get it in there? So, now that this exists, I’m very pleased with that as But wait, Mike, I wasn’t even done with the list of things I should have. I know. I know I cut you off a little early, but I think this is going to change it for you. Okay.
12:54 Mobile mobile enhancements. Rotate view added to the mobile report for landscape viewing and export data to Excel directly from mobile visuals. Those were not available until August. Now, wait, what? Oh, yeah. view. Wait, is this like on Hold on, I don’t understand. The rotate view button is now available in the PowerBI mobile report footer, letting you switch layouts with a single tap for larger visuals in a wider view.
13:26 Okay., I could just not that big a deal for me. The one I’m mystified here is export to Excel on your phone. Yes. So, it will bring up like the share feature just like if you were sending a copy of something. So you could choose a visual in the mobile report and export that data into CSV maybe in the mobile but I don’t know man do so let me let me ask a follow-up question to this one Tommy on your mobile device
13:56 Mhm. How much time do you en maybe two parts? How much time are you spending in Excel on your phone? And then two, do you even enjoy Excel on your phone? No. Let me answer the second question first. No. Okay. I have used Excel on my phone and there are so many dogone buttons on the dumb app that I can’t figure out how to make anything work around correctly. And so export to So interesting,
14:26 but am I really going to want to like export from a mobile phone app to get to my Excel app? Like does it I have to I have to test this one out, Tommy. This feels like a bit too janky to me. But let’s massage this a little because I may not want to open this in Excel, Mike. I have the data. For example, I can always feed this into Copilot, right? And just share this with Copilot directly. I don’t think the goal here is for users to go, “Oh, good. I have a CSV file. Let me go to Excel on my phone and
14:57 now develop this.” Unless they’re in a crunch on something that they were supposed to report to the boss. I think a lot of the cases are going to be sharing. Who’s asking for this? Who’s asking for Now, now I can say this is give me more tokens. Oh my gez. No way., so one thing where I think this could make sense. Mhm. iPads still act like a mobile device. They’re still pretty good. Yeah. There. And And so I think the real
15:28 estate on a screen for Excel on an iPad is actually pretty decent. And so I do have an iPad. I I I do use Excel with it. My family has an iPad that it’s kind of like a family iPad that the family can use. They also use Excel on an iPad. And to be very clear, Tommy, my iPads come with keyboards. I had the little keyboard attachment thing. You like magnet it on. You get the keyboard. Sure. Yeah, that’s that’s not bad. It’s that’s experience. It’s not it’s okay. A phone I think is pushing it. I think I think a phone form factor is a
16:00 bit too small for what I really want to do and I me personally whenever I’m in Excel on my phone, I just get angry and annoyed and frustrated with it. It doesn’t really help me. Okay. Yeah, the only way I would use Excel on my phone is if I could solely use Anthropics Claude code or co-work with my phone. Like, so inside Excel on my phone, if I could just say, “Hey, agent and I don’t want to use co-pilot cuz I don’t I feel like it’s absolutely useless inside Excel.” And I
16:30 , it it was interesting like a little bit, but like the other agents have gotten so much better than it now. I don’t I don’t even want to touch it. So if I could use Grockbot, if I could use cursor or something like that or if I could use anthropics cloud code on my mobile phone with my Excel and just narrate to it what I wanted to build because I I feel like my opinion here Tommy I don’t know if you’re about this the agents are really good at manipulating Excel now. Sure. Oh yeah. And I think that’s the
17:01 point of the mobile feature. I don’t think for people to download Excel on their computer, but Mike, wait, there’s more in terms of we are the I think the theme of August are things that should have been out, but we’re glad they’re out. That’s what we’re going to call August. Here’s another one. Enhancements to PowerBI embedding SharePoint online. You now have the ability to directly select a workspace you want to embed in SharePoint rather than copying and pasting the full URL. If you ever did it before, you need the URL of the report.
17:32 Along that you can also embed a single visual. Simply select embed a single visual toggle and select a visual. These were so everything that I have noted was not available until now until August of 2026. So Mike the theme of August to me is things we really should have had but we’re happy we have them now. Yeah. Yeah. I agree. All right. So do you still keep the gold for you? Yep. Yep. Okay. So, gold is still gold. Matrix
18:04 improvements are just listen, anything we can do to improve a matrix. Yeah. Anything we made that’s people people run in tables and matrixes and that the more we can make those things better and more like Excel, I think the happier people will be in the long run. Anyways, so love that feature as well. So, one feature here I’m just going to call out here, Tommy. Mhm. This is a sad day on this feature. Modern visual defaults and customization themes formating is now generally
18:34 available now. Customizing themes like it. That’s good. Modern visual defaults. I still I I really dislike the modern visual defaults. I still don’t like them. When you click on a visual and the title it it the I have visuals now. I was I was literally doing a demo the other day and I put the visual on the page and when I click on the visual, two little boxes appeared, the title and the subtitle, they just appeared on the visual and I’m like I don’t want I literally I
19:06 literally just made a I just made a a table and now even though it was a small table, , 30% of my screen width I was gonna say that. I was gonna say 30%. 30% of the screen of that little table just got disappeared because it said, “Hey, do you want to type in a title here? Do you want to type in a subtitle here? You can just type it into the visual.” Like, I understand that’s how you want to edit it. But if I’m just making the visual and dropping some columns into it verbatim, I don’t need that to show up. That just the whole someone needs to rethink that
19:37 whole design experience. I just really dislike the on visual editing. I feel like if you’re if you’re going to on visual edit things, the visual should look the way it does without any extra items on it. And when you double click into the visual, the border turns blue. When you’re when the border turns blue, then show all the things,, go in go into edit mode of the visual and then click on the objects on the visual to adjust them. Don’t just let me edit them verbatim because it it’s adding all kinds of extra padding the visual. It moves around. It’s just weird. I don’t
20:07 like it. It’s just bad. Yeah. Listen, whatever happened I would you rather that or on object interaction being the only option because for me do you use on object still they’re it is gonna I turned it on by default now just to get used to it because I have it’s general it’s G now they say that for four years at some point in time they’re going to just kill the old thing. now that now that this one is generally available they’re going to kill the old thing. It’s not going to stay around forever. They’re going to get rid of it. Also, they’re going to force us to adopt this
20:37 new visual editing experience. It’s going to happen moving forward. Maybe they can refine the experience and make it a little bit better, but I really dislike this new visual experience. Not good. Do not like it. I am so willing on this point to bet you a stake when it comes to whether or not on objects is going to become the default because Mike, if they were going to do it, they would have done it by now. How is it not becoming the default? I think that I think it’s gonna be hard for you to evaluate it. Let me say this
21:07 way, Tommy. It’s gonna be hard for you to poly market this one because I don’t want to poly I don’t like this. Yeah, it’s going to be hard for you to poly market this one because I think the the line like it’s going to be a very gray line moving forward around what on visual editing or modern visual defaults are going to be looking like moving forward. Like I think I think this on visual editing stuff is going to be very gray moving forward and they’re going to continue refining it. It’s going to get better over time. So, I I think you’re going to have a hard time saying it’s on. We’re going to argue more about is
21:38 it actually on or is it actually off because it’s going to be hard to determine what’s really going on there. So, I wouldn’t bet on this one. I will say though, I’m pretty sure that Microsoft’s moving more towards these on visual editing things. What I will say though one of the new default features of your reports are now gray background pages, which I absolutely love. This is amazing. Have you seen that, Tommy? Say that one more time for me. The background. So, okay. So, when you So, with the most
22:10 recent update of the desktop Yeah. Yeah. There are brand new modern visual defaults that are turned on when you use it, right? So, so let’s I’m going to be clear about what I’m difference here, right? Modern visual defaults are what default settings are there when you start a report versus on object interactions, which is a little bit different, right? That’s that’s another feature that’s somewhere else. So the modern visual defaults means every page you used to get before today was white. A white background page.
22:41 Oh, the worst. Using modern visual defaults, the page now becomes gray. So the page background
22:47 now is gray and your visual backgrounds are now white. So polishes a little bit more of that. And I believe also if I look at the the report that Zoe put out here, if I had to look very very closely intent in this visual image, I believe the visual image here also is showing rounded edges on all the visuals. So even though you have like a line chart, there’s a little tiny rounded corner on the the edges of that, which I think it’s a bit it’s it’s a bit softer than what we’ve had in the past,
23:18 right? Microsoft went very hard towards this sharp edge thing and now they’re doing this whole glass morphism design on things now and I think this is just softening the edges a little bit of visual. So I I do like this as well. So I think this one’s also a winner. I I do like the fact that the visual styles panel is now available to us. instead of having to go through the view ribbon and then click stylize. Like there’s like three clicks just to get to stylize the theme file. I’m glad that there’s improvements here,
23:49 but still it’s there’s nothing it’s better, but it still doesn’t give you what you really need to control for visual styling. It’s it’s better, but it’s not it’s I don’t think it’s ever going to really cover what Power Designer can do for you, which is all the styles, all the visuals, everything that you need to do. So, anyways, good idea. too. The one thing I’ll just quickly mention there too that I found I’m surprised you didn’t mention this was normally any things they do from a styling or visual or really just desk updates they’re always showing a screenshot from the desktop.
24:20 Yes. But the screenshot here that’s the web and Tommy I know I know listen Tommy if I die on the web thing that’s fine but on object is not happening. That’s just I’m going to still say on that but here but no but this is interesting though too that they’re showing screenshots are from the web. Yeah from the web. I Tommy I’m really thinking so my opinion here why would they come out with PowerBI desktop bridge then anyways that’s another conversation for another day. Yes, understood. But I I think also
24:52 Tommy like when I look at this, I think the right mode for Microsoft is to land features first and fast with the service. I think I think I think this is the new way honestly think about all the work that they have to do. the binary to build PowerBI desktop is 600 megabytes in size. the amount of machine or computers or computing like you’ve got a team of many many developers building on top of PowerBI desktop and they’re all pushing code and and
25:23 compiling code and making things. I got to be frankly honest Tommy building website design and having many different websites all work together as one uniformed application is way easier to manage than one massive desktop build. The the the the amount of the the git repo they must have for power desktop must be puking at this point because it must be massive because how big this thing is. I can’t even imagine. I’ve built a couple desktop applications. One of them being
25:53 business ops, right? Business ops is our desktop application. I built business ops and even that was like big and heavy and took a lot of time to build. I can’t imagine what it would take to build best desktop. It it just would be astronomically large. So anyways, I think the the speed to market, the speed to build things, the speed with less risk will be Microsoft building more features inside the service first and then trickling those features into
26:23 PowerBI desktop later on. I think that makes sense. Yeah, I see this starting to happen and I think this is going to be the way they go moving forward. All right. Well, hey, some good August updates here, Mike. Let’s run through one more and then we’ll save. We had another article but we’re already well it’s already 8 o’clock so let’s run there’s a lot of feature updates this time a lot of feature updates. So Mike just to continue with our MCP and agent skills. Well if you used the MCP before notebooks are great semantic model
26:54 building’s great now even PowerBI report designing is great but I think Brad would have been a little upset who we had on the podcast last two years ago almost at this point. Because what about love for warehouses and the thing is fabric’s great for agentic SQL development but it’s pro it’s struggled with full warehouse workflows understanding objects running queries diagnosing failures performance well now the MCP and skills help close that gap
27:27 we have secure government access to fabric warehouses fabric specific knowledge for author and consumption operations and the ability to iterate based on live results. So this is really a true warehouse assistant. The MC the fabric data warehouse MCP server. It’s a dedicated MCP server that provides agents with a man secure way to execute TSQL. Agents can connect with the global endpoint authentication using mic Microsoft Entra ID Oath 2. 0 exposes execute query and no local MCPS are
27:58 required. Governance is built in and there’s a few skills for it. authoring consumption and operations so that they work together and I think this is huge because especially something like the warehouse I think it should probably have its own MCP probably you don’t want to just bloat out the fabric remote MCP so this is great to see I I’m going to ask you a question Tommy’ll see if see how good I was just talking with Matias the other day around MCP servers for
28:29 Microsoft okay [snorts] how Many MCP servers do you think exist now for fabric built by Microsoft? Yeah, built m by Microsoft with the Microsoft fabric service. How many do you think? Built by Microsoft right now. I believe there’s the semantic model and the remote fabric. Two eight. Eight. There’s eight of them, Tommy. What? So hold on. There’s there there’s many MCP servers.
29:00 So, no, the what you can install like that they show not I’m not saying back hidden that I can’t see. No, there’s eight different MCP servers. Some of them are remote. Some of them are installable locally. There’s eight of them now. Tommy, eight MCP servers that you got to figure out which one to use, where, when, turn it on, turn it off when you need it inside your agents now. So, Matias was just doing a study on this. He had his his agents go research this and go get information for him. But he says there are now officially eight, I don’t know
29:30 if this is the ninth one that we’re seeing appear, but there are eight official MCP servers that you can go get access to. There’s a fabric one, there’s a fabric remote, there’s a PowerBI one, there’s a PowerBI remote, there’s a data warehouse, there’s a real time analytics, there’s now there is the real time one. Yeah, that’s right. There’s a lot of other MCP servers. There’s a there’s a fabric core MCP server, right? So you can tell this MCP server to go make items and get definitions of items, talk to the APIs from Fabric. There’s so many different MCP servers now. And I think it’s I
30:00 think it makes sense, right? You want to give access to your agents to do specific things on items, but I think there’s an ontology MCP server. Like there’s there’s stuff everywhere. every little feature is getting like a net new MCP server. So I I think while I’m happy that they’re there, I think we’re pres I think we’re presenting a brand new problem, which is well, how where’s all the documentation on this? Microsoft, you you need to build like great
30:31 who at Microsoft is writing up or documenting or giving us a library of here’s all the MPC servers that are existing. Here’s how you connect to each one of them. Are they all the same? Are they all handling? What a pain in the butt to do the remote ones, too. , yes. And the the the PowerBI modeling MCP server is not inside the Microsoft Store. You you you can’t get it. It’s it’s not, right? It’s not labeled the same way. Yeah. Oh, no. It is. It’s a GitHub. You
31:02 have to download and clone. Then you have to run this execution. You can get it from the store. No, no, that’s not what I’m saying. What I’m saying is inside the VS Code extension when you go in there Mhm. when you go into the VS Code extension using VS Code. yes. I use VS Code for everything now. that’s usually where I do all my agent work. So I’m I’m going to assume that’s where people are doing their work. But in inside the VS Code when you go in there and you go search for PowerBI MCP
31:33 server, it shows up as an extension. it doesn’t show up in the MCP extensions, right? So, it’s it’s like mislabeled in the Microsoft Store. And so, if you go over to if Tommy, do you use the there’s a new one, an agent mode? So, if you go to VS Code, you there’s like another window that has agent mode. Yep. Okay. So, the agent window from VS Code lets you go search really cleanly for all the MCP servers that exist. the PowerBI online MCP server is not in that list because it’s mclassified in the
32:03 Microsoft store as a regular extension and not an MCP extension. So therefore, you can’t even add it into the agent extension. You have to go manually add it using JSON, which is annoying to me. Anyway, it’s just this is the problem. This there’s so much going on with all the different MCPs. It’s like you got to like spend some time understanding how to get it going before you can really use it and be effective with it. So like it’s not like a plugin. Yeah. It’s not like an add-on. Yeah.
32:33 Yes. And I will say though the only not push back because everything you said I agree with when you do take the time to set it up and go through the headaches of it. It’s well worth it. Oh, I do agree. Yeah. I I will it’s if you are a solid developer of things, it is definitely where you want to go. All right, love this conversation. I do think we need to get onto our main topic. We are 30 minutes in. We haven’t even hit the main topic yet. I do have another topic I want to talk about, which is disposable
33:05 apps at some point, but maybe that’s a whole another topic someday. Let’s do that different day. But anyways, Tommy, go ahead. What’s the main article we’re talking about today? Let’s transition over to the main main article here from let me get the the title here. It’s critically evaluating large language models. What data visualization can teach us from EVA? Yep. So Eva, this is great. We’ve done a few articles from EVA and the website is nightingale dvs. com and we this is not our first article
33:35 that we’ve done on the podcast. We really love Eva’s work because it’s all focused on data visualization and this one is a unique approach because usually Eva’s always talked about different things that how people conceive or look at ones. I think we got cognitive load from her potentially some really great nuggets. However, what about the building of visuals? Our last episode we did Mike 556 was on PowerBI desktop bridge which is really dedicated for design planning and
34:06 creating your layout and your visuals for reports and AA’s going through two
34:11 core problems that notice and really how do you actually approach visualizing data with an LLM? The fir the two ones are hallucination is inherent. LMS are going to generate incorrect information even when asked to compile from simple public data sets and there’s a lack of transparency because at the end of the day Mike whether you’re using claude copilot gro chat GPT well chat GPT doesn’t have MCP support or whatever or your own local model
34:41 they’re products they’re not necessarily neutral tools so there’s a design that prioritizing appearing helpful sometimes over being accurate And there’s those cheerful tones. There’s a lot of times says, “I did your work.” However, you know, but it doesn’t tell you if things were actually messed up. Nicola talked about this when we went through hallucination of good code, right? But the things might have not been done the right way. So these problem exist and Ava goes through why they exist but there’s really three categories and I I
35:12 want to talk to you about because Ava’s lens on this is very much focused on public data and not necessarily this is not from a PowerBI or fabric point of view. It is collecting or extracting the data from somewhere and then building the visual just using the LLM itself. Ava talks about accuracy, security, and sovereignty are huge things. Is my data true and verified? Is my data safe and
35:42 private? Am I using it the way I want to? Again, Ava’s coming from these three different the experience of using the chatbot without,, any context behind that. But Mike for us is a little different because we have fabric and MCPS. But before we get to that, Mike, just going through this as you look at your own experience using LMS in your workflow, could would you agree with those two core problems that lack
36:13 of transparency and hallucination and what do you do about it? I’m not sure if I understand exact. So I would agree with I think the hallucin hallucination parts, right? I think there’s always this risk and and you know, regardless Every harness that you are building with a large language model now is designed to mitigate or reduce hallucinations from the agent, right? Most harnesses have like memories, they have like skills that you put in, they have
36:44 safeguards put into them.,, I don’t know all the pre-prompts that go into like the Claude code agent when you send,, tokens to the agent or a prompt to the agent, but I had to presume there’s a lot of pre-prompting that happens before I even show up, before my my prompt gets sent to that agent, right? So, there’s a lot of work around these agents to keep them from hallucinating. I think if you don’t know what you’re building and you’re just asking again I
37:15 this is where I’m going to maybe maybe I slightly disagree with some of the opinion here on this hallucination part right if you don’t know what you’re doing you get more hallucinations yes if you do know what you’re doing and you do know how to prompt how to get the answers how to test and verify that the answers are correct you get better results out from it. And what I find when I do development with an agent is this actually dovetales very
37:45 well with what I wanted to talk about earlier. Yeah. I was given a TMDL file from someone just I wasn’t able to see the data. I was only going to see the tindle. File. Okay. Yep. Just a single file timal. I needed to understand what was in that I could have loaded into PowerBI desktop, but instead I prompted my agent and said, “Build me a model view using this timle. Show me all the tables, columns, measures, and show me all the
38:15 relationships.” And within Tommy, I I guarantee you within 30 to 40 minutes, it had read that TIMDLE. It had built what I wanted. It had made the things that I needed, and I was moving report tables around and refining. So the first prompt I gave it, it did a good majority of the work. It it got 70% of the way there and then I spent the next 3 hours refining, tweaking, feature like designing inside of it. So I find
38:45 like a lot of times you can get very fast a large majority of your work and get done maybe call it 60 70% of the way there on that first pass in your prompt. If you can give it a good enough prompt, reason with it a bit, do the planning in front, you can get really close at the very beginning, once you get the answer, and this is I think maybe back to the hallucinations, right? Then it’s a lot of like refining, testing, checking, making sure that it does the right thing. And you want to put the large language models where
39:15 they’re good at, right? I think the large language model is really good at writing a SQL statement and executing it and let me look at the result of the table. Like you still have like there’s no there’s no excuse here Tommy to remove the engineer from the process. I still need someone who understands how to read SQL but you want the agent to go build it for you and return the results. So and if there’s stuff in the SQL you don’t understand you ask the agent why why you wrote it that way. I I don’t think that’s the ar to me I don’t think that’s the argument that AA
39:46 is putting up that hallucination you like in terms of dealing with it. It’s that you are always going to come in contact with this. I’m curious, Mike, with your workflow that you just mentioned, can you tell me, were you using any MCPs? Were you using any skills? Or did you provide any additional context or none? I just I just threw it I just threw it at an agent. I said, “Agent, build this.” And I gave it literally. Was it a custom agent or just a standard out of the box open chat GPT? It wasn’t a chat GPT. It was Luna actually.
40:16 Luna. Okay. which is I think I think is chat GPT but but it was a Luna model and it was really good. It was very solid. Produced a lot of really good results and I didn’t give it much there weren’t many guard rails on it but it built to from what I wanted a lot of what I needed. Now, what I feel like I’m seeing a lot more recently, Tommy, is agents are getting much more proficient at not just reading in a bunch of text and getting answers out from like just absorbing text and then regurgitating
40:46 what it found. What I’m seeing agents do a lot more, and what I think this is an effect of the harness, the harness is doing a lot more of like, oh, I’m an agent. I need to go find this thing. I need to go replace this pattern. I need to go build this thing. And instead of the agent actually going to do it with like absorbing text and code and then regurgitating new code, I see in the tooling, it’s actually making and using a lot more tools. Tool call, made a tool, load a script, then ran the script. So it’s building scripts that
41:19 are extremely deterministic and using those to get the job done for these accomplishments. So I think there’s a lot of things there that it’s doing. I’m gonna I’m gonna challenge you here a bit because I don’t think here and this is my challenge here. I’m going to push back. I don’t think you’ve said anything new here in terms of the general feel for how you’re using AI, but I want to throw this back on the context of from a data visualization point of view, right? Because this has always been the holy grail, not just to
41:49 build a single visual, but to build a report, right? in terms of in a theme with the proper context using a tim file and understanding the semantic model to me is an it’s an inherently different skill set and prompt that you would to say help me build a visual or a set of visuals for a certain data right and I think this is where it’s also too where there’s I don’t think as much context or tools that a lot of LLMs have for
42:20 visualization right it can do some things. But the example that’s given in the article is about pulling a data set, not rendering it correctly, not rendering the data correctly, but still then creating its own test data from it and then providing some visuals from it. And there’s a lot of slop that comes with that. And to me, Mike, everything you said about the tindle, fine. Yeah, sure. But that’s not nothing new under the sun. So I I think there’s a big difference here when we’re talking about
42:51 that reading code and then generating visual off of data. I I’m going to I think fundamentally I’m going to disagree with you a little bit, Tommy, here. Okay. Because I I perceive this being handled slightly different than what you’re describing, right? You’re describing like a a a vanilla model, a vanilla system. I think there are frameworks that the agent can know how to use to build decent things. So one of them is Vega Vega light right
43:23 this is a very clearly known defined specification right I think so there’s different levels of getting agents or AIs to build things right I do think Tommy to your point the acumen of the user of the agent right if you’re not really good at understanding how to wield agents telling it what it needs to know giving it clear instructions on what you what you want also I’m finding myself building a lot more design templates beforehand without real data, getting
43:55 visuals and things built the way I want to see them before I actually start adding real data to them. And I think there’s also a progression here, Tommy, that I like to use where it’s start with a SQL statement or a table of data. I feel like the agents or the AIs are much better at helping me produce a table of data first. Help me get this table data built. Then from there after I vet that data then I have the agent go ahead and build okay now render this as a bar chart use this framework and I think the
44:26 combination of taking the data development into steps is much better what I don’t like doing is throwing an agent at a blank canvas and say build this just just build this whole report. So I I think maybe where I’m coming from in my conversation here, Tommy, is I’m thinking a lot more about the sequential steps of how me personally I would build with this. So I don’t see in this article as much of an issue as this lady is describing as Eva’s describing because I like what you’re Yeah, because I’m I’ve already know I’ve
44:57 already done the data challenges. I’ve already like when I’m building even when I’m building reports myself, I’ll go to a PowerBI report. I’ll drag columns in to get data tables built. Once I understand some of those data tables, then I’ll turn them into visuals. then I’ll start thinking about the story. So to me, there’s a lot of data discovery that I’m doing before I even get to an impactful or a useful visual to a user. Yeah. Now we’re on the same page here. Now we’re on the same page because I love what you’re saying here. I you mentioned something and I’m going to change a single word
45:27 because you said allow there there you’re good. I’ll allow it.
45:32 Well, I you can put one word in my mouth, Tommy. That’s okay. You mentioned that there are frameworks that LLMs and agents can know to achieve their goal. I’m going to argue the word I’m going to change is can. I would say the there are frameworks frame there are frameworks that AI must know I think to achieve its job here. So if I if I’m going to provide the prompt or have an execution I want to provide what framework it’s going to work on. Is it Vega? Is it
46:03 PowerBI desk bridge?, and also provide the context and skill for that rather than just doing on its own. So, I I agree with your statement, but I think the framework that you’re describing is a function of the harness you put the agent in. 100%. Well, yes and no. Because if your harness doesn’t allow So this is where I think this idea of a custom harness or building custom harnesses around specific experiences can really improve
46:34 the value or the the the effort that you put into use talking to an agent, right? So there’s two ways to think of this, right? If I show up to an agent that’s just in cloud code and or command terminal interface, a CLI, and I just start prompting it and say, “Here’s what I want you to go build.” Right? When I do that with like nothing around it, it’s there is a little bit of a harness from whatever the CLI is providing. But if if I start from scratch like that, I’ve got to really shape. Okay, AI, I want you when I say this, you mean you need to go use this. Hey, here’s some boundaries on
47:05 what systems and platforms you need to use. Okay, you need to go to this website and research what Vega Vegaite is doing. So, I have to go give the agent context to understand what those things are doing. You have to do that every time when you’re when you’re talking from like a pure model standpoint from base from ground floor ground floor. When I start moving into like harnesses, the harness can be designed around building reports. The harness can be
47:35 designed around using Vega light. The harness can be designed around using Flint visuals, the new visual spec that Microsoft came up with just for agents. Right. So, right. When you look at these other frameworks, if your harness incorporates those frameworks into the harness, you can program more deterministic results using the harness. And so this is, we’ve talked about this in the podcast a lot, Tommy. I think there’s a concept of companies building custom harnesses for
48:07 what they want to do. And I don’t think I’ve seen enough yet in the marketplace for data visualization where I feel like there’s a really good custom harness that’s out there. I think people are dabbling on the periphery of what this looks like, but we don’t yet have a really good report, custom visual visual building harness that really removes a lot of these extra steps for me. Explain to the AI how does it work? How do we build things? How do we make these things work? I would I would argue with the harness side
48:37 because I think for me the three there’s three really components to the recipe to avoid hallucination and to avoid the slop that comes from especially when you’re trying to do design. Now the three things I’m going to say here Mike I have translated to really not just data visualization but all processes and all work I do. But Mike, if you want to avoid these problems here, where I don’t want the AI to always search the web to find the relevant context, if I’m asking for Vega
49:08 or if I’m asking for certain tooling, like I to me, I’m now at the point where if I’m going to do something worth my money or worth someone else’s money, it must use an MCP. It must use a plethora of skills and it must have proper context. Mike, for me the skill one is the skills that I built and I I think I want to do a user group on this, but I built something called skill vault that can push to cloud desktop to cloud code. I can sync them and even in in cloud
49:40 now my skill creator is updated to will actually when it writes the skill will convert it to notion. it it sinks things around and what I do on larger projects rather than saying use this skill or use this individual skill when it’s certain projects in in cloud or whatever the harness is I say use relevant skills to this project first list out what skills it may use it’s going to list it out and I’ll tell
50:10 which ones to use otherwise I’ll say the individual ones so skills are such a necessary part here where I don’t think you’re going to get around it. You were going to you’re asking for trouble at this point without using skills. The other side of that is use MCP if available. If there is an MCP around there, you really need to I think just lean into that whether it takes a lot of time to set it up. It’s worth the long run. And again, the context. If you’re just doing a two paragraph prompt or a giant prompt without, I think, data
50:42 validation tied into it on what it’s supposed to test, you’re you’re still going to ask for trouble here. And that goes back to that second brain we talked about where I can’t imagine at this point, Mike, in my workflow writing a giant prompt. I want you to test this, but they have metrics that use this. The model does this. I need validation. and I need a schema that it has available to it before it actually starts exploring. So Mike, regardless of the harness, those are my three pillars, my three components that avoid the I
51:14 think a lot of the things that Ava is alluding to here. Is there anything you would add or do you think I’m I’m making a mountain out of a molehill on those three things? Well, I think those are I think those are definitely good principles to kind of guide by. But one of the assumptions I think you’re having here, Tommy, is some So, I think you’re affirming my point around the harness being so important here because all you’re what you’re describing to me is custom skills, MCPs. , these are all things that we plug into the harness. harness.
51:44 And so I think you’re just affirming my point earlier and like reinforcing it more that says, look, the harness and all of these plugin things, things that I can stylize and customize and make specific to my project. That’s really where the value comes from. And all those plugins are doing, it’s just giving additional context to the AI to tell it what you really want to do, what you want to do, right? That way I can I don’t have to describe all those skills to the agent every single time. I don’t have to describe to it all the tools and
52:14 how to talk to an API and talk to the service and and do different things with stuff, right? So everything you’re telling me is you are taking pre-baked knowledge, packaging it up in MCPs and skills and give it to the agent. And I think this just what you’re describing to me is a custom harness that’s used to do a very specific task. And so I agree with you, Tommy. you’re you don’t need those things to get good visuals and good report building out of
52:45 an agent. However, you’re going to spend a lot more tokens and time reinventing things that people have already figured out and you’re going to spend more effort getting a report that actually is right and true with checks because you don’t have these skills in place. And I think that’s one of the reasons why I have found with using the skills for fabric skills for fabric and using that to help author reports huge difference like with and without them huge
53:16 difference. Now, those skills for fabric are quite large. There’s lots of them. And that’s okay, Michael. That’s okay. And and why would you say that, Tommy? Like, I I’m just saying in general, I know you want to say, why are there so many scripts? Why are there so many file? No. Okay, you’re reading ahead. This the skills are like 400 plus lines long. In general guidance, skills should be less than 400 lines long. So, like I think there’s a little bit of extra verboseness. the skills were
53:47 written from a company who has no limit on skills like like no limit on tokens like I yes Microsoft is telling their employees don’t use Anthropic anymore but like they’re not saying like stop using agents they’re not going in they’re not going to one of their employees and saying hey Tommy thanks for working at Microsoft you only get 10, 000 AIC credits to do your work this month just keep it keep it within a budget like no they’re using GitHub copilot like all day long. I mean, they’re they’re probably building
54:17 hundreds of millions of tokens because they’re not buying the software for themselves. They’re using Luna and ChatgPT and all the models they have. Any model that’s free, they can use whatever model they want to use on cheap and free. And they’re paying for the electricity. That’s it. If you if you only had to pay for the electricity, Tommy, to run your models, , this is like a Think of it, Tommy. This is a sweet deal for Microsoft.
54:47 Listen, every customer, you pay us. Hey, every customer, you pay us money per token and then we’re going to take that money from tokens and we’re going to reinvest that and build physical hardware data centers that now all we need to do is pay electricity for our own employees to write code on. Like, how how great of an accelerator is that for your company? You own the hardware that runs the large language models and you’re only paying the electricity bill. What if I’m acting like I’m only paying the electricity bill now? I rely on it
55:18 so much. No, that would be a lot. But I think that goes you’re going into the second area here that I want to massage with you a little bit. I know we’re getting your time here, but one of the other things talked about where these LLMs, whether you are using Microsoft or Anthropics, they’re they’re products. They’re not necessarily neutral tools and they’re going to try to appease you. this parts hallucination but it will try to do something that’s going to please you to make you happy whether or not sometime that’s accurate in the data because again especially with PowerBI Mike we’re
55:49 not just building a single visual off a single CSV you’re dealing with context evaluation filter evaluation semantic models and multiple columns to build to that single visual Mike I I want to quickly read I I sent this to you in the try to send it in the chat, but what LLM can do to make you happy are pretty crazy. This is a meme and I know you love memes, Michael. I love memes, especially on a podcast when people can’t see it.
56:19 Oh, yeah. Like like our 50th. Yeah. And the first half of this is Bilbo Baggins from The Hobbit and he’s clearly on a phone. And then you can see this chat with Chat GPT and it’s Bill Bo Baggins from The Hobbit who says, “After all, why not? Why shouldn’t I keep it?” to which chat GBT requires. You’re absolutely right. You found it. It’s been with you a long while. And it’s only natural to feel fond of something that served you so well, especially when someone like Gandalf suddenly seems to want it for himself. And
56:49 and one of my favorite things in the world where it will confirm things for you that are not necessarily right. And
56:55 you you can fall in this trap so easily with context with filter evaluation. We’ve done this ourselves with human error. And I think this is really important too when we are designing and building the last point or the last execution that you do is not the design but it’s validation. Mike, I cannot stress this enough in my own work with this now and testing and the workflow validation has become
57:25 such an essential point for me on design for data visualization with LLMs because without it will look nice. You’re not going to see,, something’s wrong with this visual. It will generate and make it look nice and aligned perfectly. Seth would be proud. Pixel perfect alignment. But you click on something and something may be screwy and I think this is a huge part of the workflow that we need to think as consultants or just professionals in the space. Have you run
57:55 into this at all and what’s your take on this? I have not run into the same problem that you are. I I I I understand I think I understand your principle the concept you’re talking about but I don’t I I think validation is definitely a part of everything we’re doing with agents. 100% agree with that. not disagreeing with any the validation steps. I’m not finding the agents wandering off so far and not building things that I don’t want. Right? When we use deterministic systems like PowerBI desktop, when we use determin deterministic systems like
58:27 skills and skills for fabric, I’m finding that and and this is also a part of I think a learning curve too for me Tommy, right? The difference between an engineer that uses an agent and gets the results they want and an engineer that uses an agent and does not get the result they want. A lot of people complain about the agent and the models hallucinating and not getting what you want. Yes, true. It happens. But whose fault
58:57 is that? Is it the fault of the agent? So, or is that actually a fault of you not understanding how to build better requirements, how to describe and talk like there’s a way that you need to talk to the large language model. And we’ve said this in the podcast a long time ago. It’s just like when Google search came out. Yeah. I I didn’t know how to Google search. I didn’t know what I was doing. So, for me, I had to learn how to type in phrases into Google better. So, two things were happening. I was getting better at typing in the key phrases I
59:27 needed for Google. Google’s algorithm was getting better and giving me the results that I actually wanted. Yeah. I don’t if I search for Google Tommy and I don’t find the answer on page one, I go search for something else because I know that I haven’t found what I’m looking for. It’s the same thing with agents. If I don’t oneshot it on my first prompt, something’s wrong. My prompt is wrong. It was me. And so I think the difference here, Tommy, is Yeah. In the same way that like I’m the
59:57 issue, I don’t know enough to make the agent because the same model, the same harness is being used by anthropic and they’re getting great results out of it. And the same model and same thing is being used by brand new users and they’re not getting great results out of it. So what’s what is the one variable in this system that’s changing? It’s the people and what they know, not the harness, not what the age is doing. Let me massage this a little. There’s actually, if you scroll down the article, there’s this really great graphic here, and I just want to read
60:28 something from here because I think this is a really great way of a solution to go off of what you just said. Yep. , and it’s the importance of prototyping. So, there’s an really another excellent blog blog post in the article from Nightingale on generation AI. Was prototyping really a bottleneck? And where what if the slow parts about prototyping are actually what makes it worth doing? So here and this is the quote. The key point is that prototyping is where we test the intellectual rigor of an idea. I love this graphic he
61:00 included about the intellectual refinement and technical refinement of ideas. It makes it clear how incorporating AI early on the idea stage can make shady shoddy ideas look sleek while really generating slop and how the better opportunity for keeping AI in the zone of missing missing skills and resources for ideas that already prove themselves in a prototyping stage. This is a great graphic. Yeah, it’s hard. Yeah, it’s hard to refine ideas with an AI tool. You have
61:30 to be in the driver’s seat because the tool is a reflection of what you asked for. Yes. Paired with a repackaged statistically likely out output what others have already said. People tend to assume that the ideas they have in their heads are really good even if they’re really aren’t used to release testing. So this is exactly what you’re saying. It’s exactly what I’m saying. So it’s it’s not it’s not that the agent isn’t capable. It’s not that the harness is bad. It’s that you don’t know how to wield it correctly. your way of describing it. You’re describing slop to the agent. You’re giving it too many
62:01 boundaries. You’re giving it too much things. You’re trying to go from 0% to 99% complete in one prompt. That’s not how it works. And so I think this diagram is really good, right? I really like this.,, how technically refined is the artifact, right? Well, if the item that you’re trying to build is not technically refined, you don’t need a lot of intellectually refined definition to get an to get a good answer, right? You you use low fidelity prototypes. Wireframe for me out
62:32 something. Get me the early stages. This is what I’m finding right now, Tommy. I’m building new apps and what I’m doing is I’m getting a team of agents together and I’m not giving the agents the full reign to go build the full app. What I used to be doing is I’d build the whole app all in one shot. here. I want to build an app that does this. And then what it would do is it would build really rough UI. The graphics wouldn’t look right. I didn’t have any framework to go off of. It was building junk that I didn’t like. And it was making,, every button was in a little bit slightly different position. They looked different.
63:03 It was bad. I had poor framework. So, what I’m doing now is, hey, I have an idea for an app. Refine for me the idea first. Idea gets refined. Great. Send this over to a different agent. And now go build UI designs, right? Here’s frameworks that I like. Go copy or mirror or build derivatives off of this other stuff that I like. Give it requirements. Hey, don’t just use any UI framework. Go use Fluent 2 from Microsoft. That’s what we’re going to design in. Hey, when you build designs, you can only make components.
63:34 Don’t just go build it. So, there’s all these like requirements that I’m giving it and I’m acting more like a program manager, a developer at a better level. And you’re not taking the engineer out of this process. You’re not. What you’re doing is you’re using engineering skills. You’re just not writing the code. And so this chart exactly describes my new process. And I’m moving much more from low fidelity concepts, medium fidelity designs, and then into high fidelity,
64:05 like really working and polishing. And I’m spending a little bit of time in low fidelity, maybe a tad bit more time in medium, but I spend a lot of time in the high fidelity polishing stage of with an agent. So if I had to transpose this graph into something, I’m doing one or two turns in the light blue bottom lefthand corner of this graph. in the far right hand corner where they’re talking about highfidelity prototypes. I’m doing many turns, multiple conversations with the agent to
64:36 really refine the style and design of something that I’m building. And so that’s really where things get really ex nice. And then if you follow this graph, yes, then you’re able to push out something that’s actually production. I love that, Mike. This this was a great article. So, my closing thoughts here because I know we’re getting past time here is it’s funny how this ties into a lot of things we’ve talked about in the past on our podcast and just on our discussions as well. But for me, I am just going to go back to if you want to build
65:07 effectively data visualization and obviously to me right now your most direct way with PowerBI, it’s it’s going to be I’m not going to say only in PowerBI desktop bridge, but it’s really a great great road and in a way that Microsoft has paid for you. But don’t stop there. to me use the MCP use the skills that Microsoft provide or build your own for your organization and proper context if you are not using those stool those legs you’re
65:37 not going to have a very stable table I know I should have a fourth leg because that’s a table but let’s just go with three and I think yeah three points define a plane I think I’m good with that there’s chairs that are made with three legs what an engineer right there stands up so those three pillars are such an essential part to avoid avoid the hallucination. And I think definitely take out take a look at this article. So Mike, what are your closing thoughts? I think my closing thoughts on this one is if you’re not getting the results out of an agent the way you expect,
66:08 you need to focus on your skills and your harness and what what you’re giving to the agent. likely you’re not doing a good enough job with providing good skills to the agent and you need to add better requirements and and more think out and flesh out the design. this technique has been around for many many months. It’s always plan mode first, a lot more tokens on higherend models with plan mode. Period. Yeah. Yeah. Yeah. And then you can go down to cheaper
66:38 models and let cheaper models execute on that plan. So spending ample time Everyone I talk to who’s using large language models, they’re spending more and more time and more of their tokens in the reasoning portion with an agent upfront and that greatly improves the output of this. So that’s one one major takeaway here is right don’t underestimate your harness. If you’re not getting the results that you need, it’s probably you. It’s probably not the agent and it’s probably your part of your harness. So you’ve got to improve
67:08 your harness. You’ve got to improve your prompting skills to the agent so you get better results. What do you think, Tommy? What’s your final thoughts? I love No, I already said my final thoughts. So, you miss it. The three You can’t tell. You can’t ask me for final thoughts and then be like, I already gave my Like, you said your final thoughts without announcing it was your final thought. No, I did say it. Go roll the tape. Roll the tape. I said my closing thoughts. How about that? Okay. My Okay, roll the tape. Go back. Tommy already Tommy already told you what he thinks. So, that being said,
67:38 thank you so much for listening to this episode and the podcast., we we hope you enjoyed this conversation. This was a great article. I think this is helping people understand where and what to do to use agents. And with that being said, , thank you all so much and we appreciate you and we’ll see you next time. Tommy and Mikey, dance to the day, the laughs in the mix. Fabric and AI, I get your fix. Explicit measures. Drop the beat now. Hest kings feel the
68:10 crowd. Explicit measures. Drop it loud.
Thank You
Want to catch us live? Join every Tuesday and Thursday at 7:30 AM Central on YouTube and LinkedIn.
Got a question? Head to powerbi.tips/empodcast and submit your topic ideas.
Listen on Spotify, Apple Podcasts, or wherever you get your podcasts.