Medallion Patterns are Changing – Ep.558
A year ago, Tommy Puglia expected medallion architecture to fade as Microsoft Fabric and AI changed the tools. On this Explicit Measures episode, he and Mike Carlo read Microsoft’s new series on choosing a medallion pattern for the Fabric Data Warehouse and come out with a narrower conclusion: the bronze-to-gold shape is staying, and the warehouse has to justify where it sits inside it.
News & Announcements
-
Chicagoland Power BI user group — September 24 — The user group is back on a monthly schedule, and Meetup already lists the meetings through December. September 24 covers using AI with Microsoft Fabric through a “second brain” approach, whether you keep that system in Notion or Obsidian. Register so they know you are coming.
-
Don’t Let Your Agent Touch Fabric Until It Reports Back — Tommy’s post, and the rule he uses on every Fabric MCP session: the first prompt does not build, change, or run anything. It reports the architecture, which tables hold physical data, where the shortcuts are, and where conflicts and missed opportunities sit. The prompts and code are at the bottom of the article, and the site’s prompts section links each one back to the post it came from.
-
Choosing your medallion pattern in Fabric Data Warehouse — Microsoft opened a five-part series on medallion architecture for the Fabric Data Warehouse, and this episode is a reaction to part one. The article asks how much Spark a team actually needs, then recommends either a warehouse for bronze, silver, and gold or a hybrid that lands raw data in a lakehouse and promotes gold into the warehouse. Four follow-ups are promised: building the layers in a warehouse, medallion best practices, security, and pipeline performance.
Main Discussion
Topic: Where a Fabric warehouse belongs in a medallion, and where it does not
Tommy named the author as Sidney. The piece says it is not a lakehouse-versus-warehouse argument, and it leads with a different question: how much Spark do you need? Model A is a warehouse for bronze, silver, and gold, split by schema or by workspace, for structured data and a SQL-centric team that wants one engine, transactions, and views at every stage. Model B lands raw data in a lakehouse, optionally keeps silver there, and promotes gold into the warehouse, aimed at unstructured data and teams with Python, Scala, and Spark skills. Tommy and Mike’s objection is the pattern that never appears: bronze, silver, and gold entirely in a lakehouse.
-
Medallion is still the working shape. Tommy’s prediction from a year ago — that AI and Fabric would retire the pattern — was an unofficial bet, and he is walking it back. Microsoft is still treating bronze, silver, and gold as a common way to organize data in Fabric. Mike’s read is that the pattern is getting more useful, not less: it is relatively cheap to run, agents can work against the lakehouses, and organizations adopting agents will generate more data that needs inexpensive storage before anyone analyzes it. What is missing, he said, is a real library of medallion and data-engineering skills. Tommy noted that the warehouse skills showing up in MCP are still new and thin.
-
“The team knows SQL” is a weak reason to pick a warehouse. Tommy does not recommend a warehouse for every client engagement. Mike’s clients often came from SQL, and Spark SQL is a real gap, but both hosts said an agent can cover a lot of that gap. Mike would tell the team to learn notebooks and lakehouses before he would freeze the design around current skills. He wants SQL in a notebook — write a cell, see the data, scroll, add another cell — over the old SSMS habit of one script per tab. Tommy’s addition: T-SQL already runs in a notebook, alongside Spark SQL, so a SQL-centric team is not stuck in the warehouse.
-
Gold in a warehouse has to do something a lakehouse gold table cannot. Unless an application is writing into that layer, they do not see the advantage. Reporting gold that can be inserted, updated, deleted, or merged breaks the pipeline the medallion is supposed to protect. If the only case for the warehouse is “we want less Spark,” Tommy called that a staffing choice. Mike also thinks table manipulation in the warehouse is getting more expensive, and he is content to take a V-Ordered lakehouse gold table straight into a semantic model.
-
Direct Lake is the line Mike will not give up. He described it as reading Delta tables straight into the semantic model — close to an instant VertiPaq load, with no import — until row count or model size forces DirectQuery against the SQL analytics endpoint. They did not find Direct Lake in the article. Tommy checked during the show and found you can build a semantic model on the warehouse, including from Power BI Desktop, with relationships, measures, and row-level security, and that the modeling experience looked similar to the lakehouse. In the comments, Nikki asked about Direct Lake, and a Microsoft reply treated “not Direct Lake” as the benefit of gold-tier warehouse storage. That is a deal-breaker for Mike. If the pattern cannot use Direct Lake, he wants a lakehouse-only template. The lakehouse SQL analytics endpoint is read-only; he writes back through Spark, and he does not need warehouse DML on reporting gold.
-
V-Order is still the part he trusts on the lakehouse. Mike has been looking at how VertiPaq works with Delta tables, and he asked Mim on Twitter why Databricks Delta tables are not the same as Power BI V-Order tables. Mim’s thread treated the sort as a black box — which columns are ordered, and how that optimization is chosen — and called it Microsoft’s secret ingredient. Mike’s hunch is that tuning Spark engine settings in Databricks could land files close to a V-Order layout, which is another reason he is comfortable leaving gold in the lakehouse.
-
A warehouse still has a job. It is the application side. Tommy’s floor, if a customer truly needs a warehouse, is bronze and silver in the lakehouse and gold in the warehouse, and he would rather not go further than that. He likes the Spark notebook loop — he called out Spark 2.0 and the pace of new features — and he does not see the warehouse getting faster or cheaper on that same curve. Mike’s framing is to stop making the warehouse a second lakehouse. Use it where applications write back and talk to each other, then push that data into the lakehouse for reporting. State what each engine does, and let the architect choose a warehouse-heavy model or a lake-and-notebook model for the business.
Workspaces, and the four articles still coming
-
Splitting bronze, silver, and gold into three workspaces fights how the work gets done. The article’s governance note says to keep the layers separate, ideally in different workspaces. Mike likes separation, and he treats the workspace as the control plane for who can touch the data. Lakehouse security can now limit columns, rows, and which tables a user sees; the warehouse had richer SQL-style security for longer, and the lakehouse is catching up. Dev, test, and production still matter. Hiring a different engineer per layer does not. In most organizations the same people carry data from bronze through gold, and they need the raw data unless something confidential has to live somewhere else.
-
The later parts are the ones they want to read. Tommy listed four articles still to come: building bronze, silver, and gold inside a warehouse, medallion best practices, security across the layers, and performance optimization. Performance is the one he expects to differ most from lakehouse practice, and security is the one he expects Mike to dig into. His comparison point is the older SQL warehouse, which was staging tables and a final table. Medallion came from lakehouse work and from Databricks, and stretching it across a warehouse so it looks identical feels forced. A stronger warehouse story, for Mike, is the team that wants the hyperscale SQL experience they had on-premises. He is curious whether the rest of the series changes after the feedback on part one.
Their answer to the episode title: the pattern is not changing. Raw data at bronze, through to gold that is ready to share, is still the shape. Some of the tools in the article try to shift the emphasis. The concept holds.
Looking Forward
If you are choosing a pattern this week, start from the question the article skipped: can gold stay as Delta tables in a lakehouse and feed Direct Lake? Put the warehouse where an application needs to write back, then land the reporting data in the lakehouse. Keep bronze, silver, and gold in one workspace unless a real confidentiality boundary says otherwise, and keep dev, test, and production as the split that matters. When an agent is about to build any of it, make the first prompt a read-only survey of tables, shortcuts, and conflicts before anything gets written.
Episode Transcript
0:01 Lighting up the sky. Dance during the day to laugh together. Fabric and AI understand your feelings. Explicit Measures. Let ‘s get the rhythm going now. Pumpkins feel the crowd. Explicit Measures. Good morning and welcome welcome back to the Explicit Measures podcast with Tommy and Mike. Tommy, good morning. Oh, Mike, it’s another great day. Good morning. I ‘m excited to dive into this
0:31 ‘m excited to dive into this with with you today. Tommy, you’ve been having a nice nice streak of great days lately. Every time I talk to you on the podcast, yes. It’s another beautiful day. You’re already up and at work. You probably have a good morning routine, or that Italian coffee maker you have at home at home just works at full speed in the morning. He just has to work. work. I will answer this like this. full return. So, espresso always flows like a river. But no, I was able to ride my
1:01 ride my bike on bike on Monday, Wednesday, and Friday mornings. This helps a lot. So I’m out on my out on my bike at bike at and I get and I get I was going to ask you what time you get up in the morning? And let me let me Yes, Yes, I have to start with this. I’m not a “morning bird.” Tommy, you’re much better than me at getting up this early. What time do you get up, Tommy, so you can go for a go for a bike ride? On Monday, Wednesday and Friday I get up at. I head out the door on my bike at to ride 20 miles. I’m coming back. Yes. My God
1:31 coming back. Yes. My God. 20 miles on a bike. How bike. How long does it take? Is it a couple of hours? Hour? No, 20. So I guess I guess I ride my bike slowly. Slow bike. My speed. I My speed., since mean, since I’ve been sticking to this all summer, I’m going about 17 miles per hour. The fastest I could go on a hill, which my wife didn’t approve of, was about 32. 32. Oh my god, Tommy. Yes. And you need a flashlight for that. So, I’ll be back by. I pick up the kids, we go to morning mass, I take them to
2:01 , I take them to school, and then we’re here, man. So on Tuesday and Thursday I get up a little later. I would happily ride every day, but I also need my sleep. So, you would be happy to ride every day. So, does that exhaust you by eight o’clock in the evening? Are you just like a squeezed lemon? what? I actually sleep poorly on days when I ride my bike. For some reason, I always have a 20-minute nap. Let’s say, after 10 o’clock, I always take a 20-
2:31 take a 20- minute minute break. But on those days you get a boost of energy. energy. That is, at 10 am, you say, on say, on bike ride days you take a little nap, nap, do a do a reboot, and then everything is fine. Good. Ready for the rest of the day. That’s right. So yes. Interesting. I’ve noticed that when I get up this early, I feel energetic, but then I quickly burn out, around 2-3 pm. I’m just so exhausted, man. I can barely keep my eyes open. open. This is where the extra espresso comes into play play.
3:03 Absolutely true. That’s right. Speaking of extra espresso, we’ll give you a good dose of caffeine and effort from us. We’ll talk today, our main topic: “medallion” patterns are changing. So, with the advent of Fabric and all these new tools new tools inside, what happens to medallion templates? Why are they different? What is going on? Do they change at all? We don’t even know. So we’ll try to explore this topic today in the podcast. But before that, we have a few things. We have some news from the streets.
3:34 some news from the streets. A few short messages here. Do you want to start with the news from the streets, Tommy, or the news? news? Why don’t we dive into the news from the streets first? So, a few. , a few. Let’s start with news from the streets. Let’s do that. I like it. Okay, Tommy, what do you have? What have you learned lately? I lately? I think I said news first, but yeah, let’s get to the news from the streets. streets. Oh, sorry. I apologize. As you wish, Tommy. Let’s do this first. So, first a brief announcement, and then we’ll move on to news from the streets. Well done. OK. done. OK. So, first: we are
4:04 So, first: we are getting ready again for our next next user meeting. Again, we keep , we keep talking about this because this is the first time we’re going back to a monthly schedule, and we already have all the user group meetings scheduled for the whole year. If you go to the to the meetup of the year, yes, Chicago Land Power BI on Meetup, you can see September, October, November, December—it’s all set. And our next meeting is on September 24th, where we will talk about using using artificial intelligence to work with Microsoft Fabric using a
4:34 using a “second brain” approach. Regardless of whether you use you use Notion or Osbeden. That seems to be how it’s pronounced. Obsidian. What exactly are you talking about? Well, it’s like Osbeden or Obsidian,. Obsidian. Obsidian. Yes. Yes. And I know this because I play Minecraft a lot. So when Tommy gets his words mixed up a little, I can still understand what he’s talking about. I talking about., we’ve been mean, we’ve been talking for too long, so I understand everything. Oh, don’t worry. It’s just like
5:04 Oh, don’t worry. It’s just like you update or change my questions. You say, ” I’m just rephrased. ” ” Okay, Tommy, let’s get to the news. -huh, okay. So, news from the street. They drove. So, it’s September 24th. Be sure Be sure to register so we know you’re coming. Next. Mike got the article I wrote and I’m just going crazy with my blog, man. This was another one I like. I like. I constantly see something appearing there. This seems very relevant to what we’ve been discussing together in our podcast. But it’s time. It’s like your… it’s
5:34 it’s time. It’s like your… it’s like your second, third, fourth brain, what’s going on here. You talk about it on the podcast and you think, ” Well, Mike talks too much. ” “I’ll ” “I’ll write a blog too to make sure I get my point across.” This is what I want to finish when I can’t finish a finish a sentence. I do this on the blog. blog. Yes. Yes. Well, to be honest, the blog really came about because we discuss a lot of interesting things during the episodes, and I think, “Damn, I want to explore this in more detail.” And that’s exactly what happens with an article: I look at it and think, “Okay, I
6:04 think, “Okay, I wrote that down.”, that’s actually a pretty good idea. Well, I wrote an article about not letting agents touch Fabric until it until it reports back. And I try not to not to just express an opinion in my articles, but to really provide prompts. This is promptingbi. com. And,, one of the things I do there, and yes, I noticed the video too. So one thing to do do later is talk about, actually, when you’re going to going to use Fabric
6:34 use Fabric MCP or bring in an agent to work with Fabric, don’t ask them to build anything right away. Yes. One of my things is that it’s that it’s a mistake, a big mistake. mistake. One of my One of my tried and tested approaches is that every session I session I start with Fabric MCP starts with an exploration prompt that doesn’t build, change, or run anything. Intelligence prompt? I’ve never heard that before. I like this like this wording— wording— exploratory prompt. The reconnaissance prompt just reports
7:04 prompt just reports what’s available, what’s there, the architecture,, is there data, is there pointers, is there actual physical data, what data is that. And the intelligence report simply tells you what is available, what the physical tables are, where there are shortcuts, where there are conflicts, and what opportunities are available and lost. lost. I won’t I won’t list everything I wrote down, but one of the things I do here is validation.
7:34 I actually have prompts here, and what’s really cool right now is that on the blog, if you scroll to the bottom of the article, I have a library. And there you can see see the prompts and code from this post themselves. Which, actually, says something like that. And if you look at the top of my site, there’s now a prompts section that links any prompt, instructions, or agent skills to the article they came from. Steeply. So yes. So, Mike, the main thing here—I don’t know if it’s the
8:04 here—I don’t know if it’s the same process as yours, but I always do this reconnaissance when I’m working on something. Yes, I like that, Tommy. And,, I totally agree with Tommy, as we build our websites, do certain things, and put everything together. There are almost no limits to what you can create on websites these days, right? Oh yes.
8:34 right? Oh yes. That is, you describe what you want, and it simply produces the result. And the fact that you have spent years refining, refining, developing, and creating prompts. This is your secret weapon. I think it’s a great idea to implement this and develop it, integrate it right here, into your site, because it’s great, and now there are a lot of different tips. One point I’ll note— maybe it maybe it applies to me too. I’m not sure I can find them just by browsing your site.
9:04 browsing your site. I understand. So, a search is required on the tips page. . I will work on this. Yes. . Yes. Good. There are many ideas. OK. Why am I suggesting this to you, Tommy? Because yesterday Tommy came up to me and said, “Hey, I was looking for…” - Tommy and I do a podcast all the time. So if you go to PowerBI, I did the same thing. I created a static site with site with PowerBI tips. It runs on GitHub Pages. Very simple. We can produce hundreds of words of content. He takes a full
9:34 He takes a full transcription of our episode. He posts it on the website so we can see everything we discuss, word for word. It’s like our catalog or index. I don’t know what to call it. . We talked about it. Yes. This is our knowledge base of everything we’ve said on the podcast, starting around episode 250. I didn’t redo everything from scratch, but I did a lot. Now that all the episodes are on the site, you can search for any text or things we mentioned and it will show every episode where it was.
10:04 episode where it was. Tom says, “I need a need a search function.” I need better search. I need to exclude or add certain phrases. Something like a phrase search, because I’m looking for, for example, ” soft data” because I wrote an article about it. . Yes, that’s right. You want to refer to yourself, yourself, right? That’s right. I’m like, “What did I say?” I understand. I said this and added, “We need Google search operators.” Mike is like, “Okay,” and six hours later, even faster, he says, ” Yeah, it’s done.” Yes. Bing, done. This is very cool. But this is
10:34 very cool. But this is the speed of thought. Speed of thought is when Tommy comes to me and says, “I need a feature.” I say, “Hey, agent, here’s a new feature we need to add.” “Add her.” And he just chews for a minute or two and says, “It ‘s done.” Perfectly. Let me check. Start development. Run mpmdev. I see it. Boom. Looks great. Okay. Yes. Let’s release. And now it’s on the website. That’s the speed. I think Tommy, my gut feeling is that more and more software is going to have to
11:00 software is going to have to run in this accelerated mode, right? This will be several updates per week. These will be weekly updates., the , the VS Code team now updates its code every week., you keep getting a lot of updates from the cloud code all the time. So I think companies that are able to use AI in a way that allows them to use AI use AI to help with development “on the fly” and
11:30 development “on the fly” and make it an integral part of their process. A lot of people, I think you heard that, Tommy, and you argued with me a little bit too. You said, ” Listen, what about the fact that I don’t really want the AI to work, is that okay?” Can I use use AI-generated things in production? You shouldn’t do this thoughtlessly, right? right? But if you think about what AI can do, it writes code better, faster, more efficiently than you can. And on top of that, you just need to be more
12:00 need to be more architecturally literate about adding your own tests, tests, checkpoints, making sure it gives you diagrams and information. I’m doing a full,, I’m building new websites, Tommy, or web applications. And I go through the whole process with the help of agents. Here’s my idea. Go build me some me some wireframes. Based on the wireframes, I want you to draw up a specification for me. With . With the specification, go to Miro. Draw me all the architectural diagrams. Give me all the all the sequence diagrams. Give me everything. Even
12:31 me everything. Even internally, my teammates teammates say they say they want to make architectural changes to the software. And I say: great, write a brief. Write why we do this. Tell me what we have now, what’s new, what’s the difference. And I said: I need 100% test or code coverage of every function that touches the data system on the backend. How do we replace this? How do we make sure that everything on the backend
13:01 everything on the backend automatically works correctly, as we need it to? That’s pretty close to the speed of thought, right? When you have an idea and idea and you just start planning, I planning,, I was literally mean, I was literally talking to my buddy yesterday and he was talking about it, it wasn’t me, by the way. It wasn’t you. You are one of my friends. This is another friend. It was another friend. another friend. But we were talking about AI-generated content and comparing it to the fact that there’s so much trash around right now, , right? And so many fast things from AI, so we compared it to a
13:32 so we compared it to a lake full of so- called bad fish, but every now and then there’s a really really good fish in there. How do you make sure that your fish is, in some sense, of high quality? Because if someone wants to go fishing, and 90% of the fish there are bad, would you go fishing there? Or if you looked over there and were like, “Oh, look at that dolphin, man. Like, look at that trout. That’s a good catch.” So, this requires requires planning. And I kept telling my friend: you have to understand that there are skills for creating something that
14:02 creating something that creates something else. And this requires experience. You can go to ChatGPT or wherever to to create something, Mike. But I think if you don’t have the right structure, and as I said on the blog, if I work at Fabric, it’s not just updating a notebook, it’s exploration, understanding what we’re working with, and then it starts to plan, for me personally, Mike, I use questioning all the time. So yeah, and that type of planning to
14:32 planning to break down what you’re actually actually trying to achieve to make it AI-ready is really important. I agree 100%, Tommy. This is interesting. Tommy, your kids aren’t old enough to play with AI as an agent, but my son is. So I’m starting to encourage him now, we sit down and work on things together. He is very passionate about video games. And again, he’s still a little kid.
15:03 Nowadays, for kids, every video game is super graphics, ultra 3D, and hyper-realism. And I say, man, there was an era of gaming where you could barely see the pixels on the screen. I think about the games I love. Again, I loved a lot of strategy,? I liked ? I liked Dr. Mario. I liked Civilization. I liked,, the game Dune, the very first version of Dune and maybe Dune 2., there was also Command and Conquer. All those games were
15:34 Conquer. All those games were like, “Oh, man.” Warcraft, I loved it, it was an era, it was the golden age of gaming. How many hours have I spent, Tommy, playing these games. How long did I wait for my,, internet to connect to get an update for the app, and, I turned on the modem so it would download. I waited hours to get this. I would leave it leave it charging charging overnight so I could play a game or several games the next day. Do you remember? Do you remember those times? I remember that.
16:05 I remember that. what game I remember from our early our early life? Do you remember the one where there was only console navigation, like you walk into a room, find a chair, a table, and a wall. And you’re like, “Pull up the chair.” You can’t do this. You’re like, “Okay,, look under the chair.” “You found the key.” You’re like, “Okay.” And, I don’t know what this genre of games is called. . Oh, yes. I know exactly what you’re talking about, Tommy. And this is one of those early games where there were only words. It was like, you read that, right. You read all this. I found them
16:35 I found them incredibly annoying because maybe my mind works non-linearly. I could never get never get them to them to work properly. I will create a version for Fabric. This is what I just wrote down. And I’m going to add that. add that. For example, you go to Lakehouse and see that there is no data there. What are you doing? You’re like: I write code. Code unavailable. Tommy, Tommy, I’m going. Sorry, Tommy is not available right now. I’m going to add some,, yeah, I’m going to add some Easter eggs in there. For example, listen to a podcast. You
17:05 listen to a podcast. You won. won. I like it. Well, you should do it, Tommy. And don’t just do it, do it as a Rayfin application. Then you can deploy it and distribute it for free, and we , and we can give it to our audience to play our play our Rayfin game,,, a library of different things that we’re building and creating together here. So, that would be, that would be a lot of fun. That would be really cool. cool. Okay, I’ll go for it because I feel like it’s a great idea. So, would that mean your agents would be working on this? You this? You won’t actually do anything. Do you think I’m going to write a CS DOS game based on this? And CD too? Hmm,
17:36 ? And CD too? Hmm, great. great. Okay, Mike. So, a little bit from the streets, and I just wanted to demonstrate to our listeners and to you how powerful the whole Fabric world is in PowerBI MCP. Er, when you use it properly. Mike, I’ve been working on a lot of new projects or potential new projects. I have a GitHub repository where all my all my technical tasks, various various work results and all that are stored. And I thought, I was trying to create a
18:06 was trying to create a Power App before, where I would enter information about a project or estimates, then the data would go to a SQL server in Azure, and then I would create a report based on that. But it was a very tedious process, because if I wanted every result of the work, I had to write it down there, and then in my technical task. So I’m like, what, let’s think about it, and I started developing, because I have all this experience with Fabric. I have a project for Fabric guru and another one for my,, ,, consulting
18:36 consulting help. Like, I want you to create a skill that will scan my client client repository, and every time a PDF is updated or changed, it will trigger a series of events. I want to have a database in Fabric with all my all my projects, projects, results, results, statuses; statuses; to review Notion, because they are synchronized, and essentially build me a pipeline, like a record book. And I also want to see not only
19:06 want to see not only what is currently active or available, what the status is, but also what types of deliverables and milestones have worked best in the past. And Mike, what I’ve created now is that whenever I write a new PDF or have one finalized, I just run this skill and it scans all my existing technical tasks. It, in a sense, a sense, extracts the results of the work, the cost, the essence of the result itself,
19:36 the result itself, classifies it and puts it into Microsoft SQL Server, or it seems to be just Lakehouse, and from there builds a semantic model. I then created a created a skillset using Notion to have instructions for the PowerBI desktop bridge using the report builder and planner. And Mike, now I have a direct connection when something is updated in my new projects. If I If I talk to someone today, like, “oh, we want to change something,” it goes through the Fabric update, my
20:06 Fabric update, my reports automatically update, and I can see what stage they’re at. I see what open funnels there are; I can see the volume of backlogs, backlogs, what what percentage of successful deals I have had in the past, and I can filter this based on the type of projects. I have something that we developed, we called it the results economy. Well,, what really works based on hours, where I went over budget on certain things, where hours compared to cost per cost per result. So I see
20:37 result. So I see here,, here,, looking at this, my training is one of the least profitable, because of how many hours are spent compared to the money spent. And Mike, this all comes from PDS from the repository, and we can pull that out and build on that Fabric Lakehouse. Honestly, it took me about an hour to create this. So I think,, going back to our previous point, I just wanted to mention this because it’s
21:07 mention this because it’s really important, Mike. We move at the speed of thought when it comes to what is possible for us. I believe that infrastructure is very important. I think if it wasn’t on GitHub or I didn’t have Fabric, it would be much harder to do. You can create the HTML to do this here. . Yes. But the fact that Microsoft also gave us MCP and this access means that for me we are now at a new stage where I see what I have been able to build. It seems to
21:37 me that everything is accelerating in this direction. We’re moving faster and faster towards this, , less effort on my part. Let me pick up on your comment, Tommy, regarding your complaints about the market situation. I think this is very appropriate. I find it’s better to integrate more more software. And honestly, Tommy, I’m getting to the point now that if the software doesn’t
21:59 if the software doesn’t integrate well with my agents, I don’t want to use it use it. . Yes. Yes. That’s going to sound very bold, Tommy. That would be a very bold statement. Recently I,, you have, so I… let me let me slow down. My brain is brain is thinking too fast this morning. I drank too much coffee. Sorry, Tommy. I assume so, correct me if I’m me if I’m wrong in this assumption. You, you now have now have Excel installed on your computer. Yes. Yes. Good. Why? I do
22:29 Good. Why? I do n’t know, because I haven’t used it in forever. Good. I recently upgraded my computer, which gives you a clean clean system, and I wanted something lightweight like Excel. I adore him. This is a great tool. I tool. I don’t run macros anymore. I’m not doing anything complicated. I don’t use anything like that anymore. Most of the windows on my computer are now occupied by AI chatbots. I have more CLIs open than ever. I have more VS Code open
22:59 more VS Code open than ever. I’m constantly on GitHub. I spend significantly less time on business products like Microsoft. So I thought, I don’t want to keep a big, big, bulky bulky Microsoft product on my computer. You deleted yours… I backed off and switched to Open Office. So Apache, oh yeah, the Apache group, which is the same company that created Spark, has this thing called Open Office. Again, I’m not doing anything too complicated anymore. Most of
23:30 complicated anymore. Most of my work my work is done in Fabric, in the data engineering area, so I’m more of a developer now. So, Tommy, I installed Apache Open Office. It has 390 million million downloads. It is significantly smaller than Excel and other programs. It is light. I can download the Excel file. I can open it. I can view the data in it. I can create lines. I can create a table. Well, just basic stuff. I’m conducting an experiment to see how long I can go without doing anything. And one more
24:01 anything. And one more point, Tommy, he’s pretty good. Open Office works quite decently with AI. So if I say, ” Process this Excel file,” Claude, Grock, or whatever agent you’re using right now, you can just say, “Open this file and do this and that.” Boom, done. I don’t I don’t need Excel anymore. It just opens the binary file, makes changes, and does what I want. I think to myself, oh my god, Tommy, this is
24:31 oh my god, Tommy, this is just incredible. So, I’m conducting this experiment. I don’t know how how successful it will be. I guess sooner or later I’ll give up and have to go back to Excel. But with PowerPoint, with Excel, with HTML documents, Tommy, we can do everything very easily, why do I even need a PowerPoint presentation? Okay, this is a different question than with Excel because
25:01 because it’s the same package. I think that’s right, it’s the same package. I will focus on Excel because I still see the advantage of presentations for conveying certain elements. But regarding Excel, yes, Mike, you could just write on the board how many days Excel is worth. True. True. True., and I still use still use Excel, for example, I create create data catalogs for clients in Excel. It is easy to navigate. navigate. Of course. Various things. But to
25:31 Various things. But to your point, it could be Open Office in terms of verification, as if as if you and I are exceptions, I think. If we were to take a distribution curve of people people using using agent-based solutions, we’d probably be in the 90th, maybe 95th percentile, right? We are at the extreme level of how much this has become part of our workflow. For the average person who also works in an organization where Excel is still pretty much,,
26:01 still pretty much,, constant and just exists, exists, I would say they probably couldn’t do without it. But I think you’re talking about a future where there’s going to be a point, Mike, where applications are going to change radically, because I can even go into Notion, and Notion can create, extract, and read an Excel file. Notion, right? Not this cloud bot, but it can extract information to where I can can use it. Everything, as you say, should be agent- be agent- accessible and ready
26:31 accessible and ready for agent work, right? Where regardless of the format, whether it’s binary or a data structure, everything has to be easily understandable for the agent, otherwise I don’t see the point in it, it’s pointless to me. And this is one of those areas where I still feel some resistance. So,, I’ve also been trying to use Work IQ lately,. lately,. When you’re inside a Microsoft product, it works fine, but as soon as I go into
27:01 as soon as I go into Work IQ, for some reason Work IQ requires me to have a pay-per- use Co-pilot license to access from outside Microsoft. They changed something, and again, I say, “I don’t want to deal with you.” Honestly, if Microsoft is going to be such a closed “wall”, and I understand that it’s for security reasons, well, fine, but I’m trying to trying to do something, right? And if you hold you hold me back and don’t give me an easy way to communicate with various APIs within Microsoft at no additional cost, you’re just making my
27:31 making my job harder. So, this is another area of friction. If you don’t change and adapt, people will leave your system, they will find other tools. Do tools. Do how long it will be before everyone starts creating creating PowerPoint presentations by simply telling AI, ” Hey, I want a slide that looks like this ”? And it creates a completely HTML presentation. It’s like,, a barrier. I would say that a couple of months ago this was a popular topic. Everyone was creating HTML presentation tools.
28:03 It’s common now,, if you , if you have Claude Code. Claude will do it for you even if you didn’t ask. Yes. ask. Yes. You go to Claude and just say: give me a presentation. He just betrays her. He creates an entire infrastructure. There are slides on the left. There are notes at the bottom. The slides are ahead, and you can export it all as HTML. I’m like… And it looks better. And it can even be edited. I was just playing with this. You can click on different blocks and edit the text in them. I’m like, “Oh
28:33 them. I’m like, “Oh my God, this is really something incredible.” In any case, any case, yes, yes, things things are changing a lot. I like the direction of development. We are on the right track. But in the end, you have to be careful. We need to get involved. We need to start studying these things. Speaking of you needing things to be clear to AI, I’ll touch on touch on that quickly to get to our main topic, but it’s not so much news as a slight annoyance or frustration I feel because… Of course. We talked about this, I don’t remember how long ago, but… the
29:07 data preparation function for AI, where you and I imagined that in the semantic model, there is a folder in the pipet file. Yes. And you mentioned there that there’s a data preparation file, or I don’t know if it’s actually a separate file for the instructions. Yes, that’s right. And I want to explore this in more detail, because one of the articles I just wrote was about ” soft” and “hard” data. Again, here’s what we’re talking about: if I give an agent the ability to view a SharePoint site
29:37 to view a SharePoint site and,, have that cross-integration to update instructions. That’s where I see the point of this, because Mike, I’m looking at data preparation for AI and in Power BI, and I don’t see much benefit in just writing instructions myself in this empty empty text box. because 100 %. %. Yes, because this is what we’ve been talking about and emphasizing over the last few months: context is everything, and
30:07 context is everything, and instructions are good. But if it can’t read or provide information about,, what , what Word documents or Excel files are there and what’s important in there, how it relates to the data in this semantic model or model or metadata, this data preparation for AI, then those instructions are not going to have any real value. So I have this cry of the soul, more of a wish, where I really think what we need from Microsoft is, can I zoom in? I’ll
30:37 zoom in? I’ll zoom in another time, we need a folder for AI that can provide a lot of skills and a lot more context that actually works, not just one just one instruction file, which by the way is also limited in in character count, which I don’t like. And he, he doesn’t allow instructions; you can’t say, ” Go to this SharePoint site or this or this GitHub repository and read the content while looking at the semantic model.” This is purely for what’s in this model, which I think makes you
31:07 which I think makes you lose a lot of context. So Mike, yes, yes, you mentioned that you don’t really really like this like this future either. Do you agree with what I’m saying that we need to expand on this? What do we need to do to make it better? This is great, I think it’s a very interesting topic., I don’t know if we know the answer yet. I think, I think our whole world of how we build things is changing a lot., also I think there are a lot of people who were educated and their entire careers were writing code. Yes.
31:37 Yes. And I think they’re a little excited. I read an article by someone on Twitter or X that talked about the biggest problem with using AI: it’s people who have classical training in writing code, building solutions and applications. There’s this resistance to holding on to every function and line of code you’ve written because everything has to be exactly how you how you envision it. But AI
32:07 envision it. But AI can beat can beat you in you in typing speed, right? So your value is not in diving into the code, but in in stepping back from it. Your work should shift towards discussing concepts, patterns, architectural solutions, testing, and determining the truth of data at runtime. I think our thoughts can start to shift
32:37 start to shift from directly working on code to explaining the business rules to the agent about what exactly you want to build. But the question immediately arises: how do we know that this is correct? Because there is always the thought that everyone can hallucinate, make mistakes, and agents can do something
32:59 can do something wrong. It may take two or three tries to get it right, but you as the user should still check and check and confirm it. We are now agent managers, not people managers. I see this in my engineering team. We significantly increase productivity by by using using agents. agents. And I want to point out that it’s not just about productivity. My work for 30 minutes has much more value than
33:29 more value than just writing code or instructions for the same amount of time. My ability to plan tasks is much more valuable than knowing a specific function or 30 minutes of manual manual coding, right? And I think the idea of performance… My ability to drive a car — again, if I didn’t have a car, I would n’t be a fast runner, runner, but the ability to drive a car is much more valuable than being a fast runner, right? This is manual labor. I don’t know if that’s the best
34:00 know if that’s the best analogy, but you get the point . I think so. I think so. That’s the equivalent of me saying, “Tommy, I need you to walk, run, or even even bike to bike to Milwaukee,” right? Yes. Yes. Hey Tommy, ride your bike to bike to Milwaukee. Or, “Tommy, get in the car and drive to Milwaukee,” right? So, each of these three things can lead you to your goal. They will take take different amounts of time for you to cover the required distance. It’s just that with AI you use you use teleport,
34:30 teleport, right? This is even higher than a car. I car., this is the level of mean, this is the level of Star Trek, where you can just end up where you need to be. So this is the world we we live in now., you could could use use code helpers, IDEs, or other tools to write code and build things faster. But suddenly we’ve got this new technology, and now the question is how do we use it around our understanding to
35:01 apply it effectively, which I think is incredibly exciting. exciting. 100% of the time. Well, Mike, I think this is a great transition into architectural planning. Yes, let’s get to it. By the way, I really like our introductions. I know they’re getting longer, but that’s good. This is good for a podcast. I liked these conversations. So, let’s get to what Microsoft wants us to talk about, or doesn’t want us to talk about, which is Microsoft’s new series on choosing a medallion architecture for its Fabric data storage. And, Mike, this is the same series they’re going to release
35:31 going to release. It looks like they did a six-part series on architecture called ” Medallion.” This is for the Fabric data warehouse, not just in general, but for the data warehouse, and there are a few points here that I want to start the conversation with, because I may have been wrong a few few months or maybe a year ago. Does this mean you owe me a steak? It wasn’t a bet on steak. steak. Too bad, I lost the opportunity to have a delicious steak dinner
36:01 the opportunity to have a delicious steak dinner. . Yes. So it was an unofficial bet. Simply put. So we . So we can’t print it on a t-shirt. A year ago, I said that I foresaw the abandonment of the medallion architecture, because I thought that AI and Fabric would fundamentally change it. We, again, the context was that we we used used Medallion because it was the best tools and tools and approach at the time, but technology changes quickly. So,
36:32 changes quickly. So, Microsoft is proposing the Medallion architecture because it is and, I think, will remain one of the most common patterns for organizing data in Fabric for the foreseeable future. And Microsoft states: ” Keep in mind, there are certain ideas and questions you should ask yourself: should you use use lakehouse or warehouse?” Because I think, Mike, that’s the main thing: most people are more comfortable
37:02 people are more comfortable using using lakehouse, simply because it seems less technical. It seems like you’re not going to,, break anything. I used lakehouse to get the data. I can I can use use lakehouse for reports, especially if you don’t have a technical or DBA background. Mhm. But Microsoft asks this question differently: instead of wondering whether you need a lakehouse or a warehouse, they want us to ask: how much Spark do you actually need? And I think
37:34 need? And I think this leads to a larger idea when we look at the options and approach to data retrieval in Fabric. Let me ask you right now whether the warehouse in Fabric really changes the approach to the architecture we use when use when loading data— lakehouse or warehouse—because I don’t think it think it matters that much. I
38:04 matters that much. I matters that much., I think that’s mean, I think that’s why I don’t think it’s happening happening on time, I think, a little bit here too, Tommy. So,, last ,, last weekend or the last few days, I was like, Oh my God. How do I even approach this topic? I think it’s related, but maybe not. OK. T So, let me make a caveat here. Talk about deviation. Not yet. OK. OK. Tangentially related, right? Er, Er, sorry. What? What did you say? say? Tangentially
38:34 Tangentially related. That is, something parallel. We are tangential. We are close, but maybe it’s not the same. So, same. So, you shouldn’t use use difficult words with me, but I’m but I’m sorry. Sorry., it’s not tangerine., this is tangentially… … tangelin related. So, So, tangelini is related. This is not… this is not pasta. This is not linguine., so I was trying to understand a little more about the VertiPaq engine and how it works with Delta tables and Microsoft Fabric. So, one of
39:04 Fabric. So, one of the advantages of Microsoft Fabric is the ability to build “ “ medallion” architectures: bronze, silver, gold levels. Tommy, you said it doesn’t really matter matter —lakehouse or warehouse, it all works. I works. I agree with you. I understand what you mean., I don’t have any animosity, but it seems to me that warehouse is starting to get a little more expensive. Yes. In financial terms, when manipulating tables. And if I can just go from lakehouse to gold table and straight into
39:34 table and straight into semantic model using using V-Order sorting, I’m perfectly happy. I bring up this topic because I don’t think medallion architecture will disappear. I think there’s definitely something to it. I think it’s pretty cheap to run a lot of things on the Medallion architecture. This gets a little more complicated as a system design. But I think it’s possible to bring in agents to simplify some of these tasks. Now, we already have the relevant skills.
40:04 relevant skills. Yes. And right now there are no really good medallion architecture skills, data engineering skills. Data engineering is a library of skills that really helps you navigate from raw data through the entire process. You still need to keep a close eye on the data and clearly explain what exactly you are trying to do to do to transform it. So, I think it’s still relevant. relevant. Yes. But the reason I’m bringing all this up is because I was I was talking to Mim. Do you talking to Mim. Do Mim?, yes. know Mim?, yes. Good. So, Mim is on
40:34 Good. So, Mim is on Twitter, and I wrote to him, “Hey Mim, I don’t know the answer to this, but I have tables in Databricks, and I know that Databricks Delta tables—they’re not the same as same as Power BI V-Order tables. What’s the matter?” And Mim listed about 10 or 15 really good tips, like, “Hey, when you’re building models and semantic models, I know that the models should do this, this, this, and this.” It was wonderful. So I’ll try to find that Twitter thread
41:04 find that Twitter thread where I talked to Mim about this, but he provided some incredible information and insights into how V-Order sorting works. It’s a bit of a “black box”—which columns are sorted and where, and how they optimize the V-Order sort, because that’s actually actually Microsoft’s secret ingredient. But my intuition told me that if I could get Databricks to adjust the Spark engine settings, it should produce files that were pretty close to what we use to use to organize VORDER.
41:35 organize VORDER. So, I was just playing around with all these things, just enjoying the process. But to your point, Tommy, yes, the Medallion architecture, in my opinion, is only getting stronger. She is becoming increasingly important to all of this. It is easy for agents to interact with it. We have a lot of tools to help help build these lakehouses, and the volume of my data. Well, just from creating applications, right? I create applications. Every application generates a lot of additional data. So
42:06 additional data. So as your as your organization grows and you decide, in turn, to bring in more of these agents and so on, I think there’s going to be a lot more data generated, because now agents are creating, agents are building, agents are developing software, which again has to come to a come to a common place to find out whether it’s effective or not, and other things. So you’ll need some need some cheap and cheap and inexpensive storage to simply accumulate large amounts of data and then extract analytics as needed. So yeah,
42:36 , I don’t know. I’m not sure if sure if that answers your question, Tommy. I’m just trying to figure this out. We will work with this. We will work with this. I think the most important thing— I want to at least mention that Microsoft has created some skills for the Medallion approach. I think one is called ET… to U, but as you say, there’s no comprehensive skill set, I think it’s very limited, like one or two. And again, the data warehouse skills in MCP are very fresh, it’s very new. I think the
43:07 new. I think the main thing for me is that I have some reservations about why I wouldn’t just use the use the lakehouse approach from a diagram perspective, as you say, because I would also be concerned about the cost of cost of using storage. And I want to mention here that they have two recommended patterns. patterns. Yes. And let’s get get straight to the article. Yes. Let’s define what exactly you’re seeing here, Tommy. Yes. Continue. Continue. And I have a few complaints about it. But yes. But yes. They say that there are two underlying patterns that they recommend when you again again use the
43:37 use the repository approach. And they’re not saying this about whether you should use use storage or lakehouse, but about how much Spark you need. They start with this very question. Model A is a universal data storage. Use a Use a data warehouse for bronze, silver, and gold levels, separated by schemas or separate workspaces. And they say it’s say it’s best when the data is structured or can be can be structured at the time of loading.
43:58 time of loading. The team The team prefers prefers SQL-centric development. So it’s one driver, one set of skills, transactions, and representations at every stage. Model B is a hybrid of a storage and a lakehouse. Upload Upload raw data to lakehouse. Yes. Optionally process the silver level in lakehouse. Then grant the gold level in the repository. And again, it’s best when you don’t have a… Yes. … Yes. And I would say, I like it. I would also say that you can leave lakehouse for bronze and silver and then load gold into the
44:29 load gold into the repository, or any other combination, but there are certain advantages when it comes to SQL data warehouse. Yes. And that’s what I want to discuss here, because Model B with lakehouse requires strong Python, Scala, and Spark skills, and again, the advantage is that there is no data copying, one supports both. So, their rule of thumb is: if you’re an SQL- centric team, use a use a universal universal data warehouse. If you have unstructured data
44:59 unstructured data and need an engineering approach, use a use a hybrid. And I have a few ideas about this. The first one: they seem to be pushing more teams to choose data warehousing as the preferred approach, because I don’t know about you, Mike, but how often do you recommend recommend this to clients? No, we will make a complete repository for everything. One thing I didn’t understand from this article article is what
45:30 is what, , they talk a lot about a SQL-centric analytics team in this article. OK. Well, I would say that many teams are already SQL-oriented. I completely agree, of course. I have many clients that I work with. They came from SQL. All of their data engineering is based on SQL. My argument here is Spark SQL. Yes. A little different, but any gap in knowledge about how to do things in Spark SQL compared to what you need to know. Yes. Yes. Talk to the agent. The agents will handle
46:00 The agents will handle it. He will insure you. No problem. So, from that perspective, I think: I don’t see the argument of ” just just use the use the data warehouse because that’s what your team your team knows.” Yes, that’s a fair fair argument, but I wouldn’t make any serious decisions based on based on that alone. You might have to tell the team, “Okay, guys, you need to learn to work with more with more laptops and lakehouses.” And honestly, Tommy, development is cool, I love writing in
46:30 love writing in laptops. Moreover, now they have SQL laptops, it’s just great. But I can’t stand the old old writing style, like in SMS, where there’s one script per tab, and to do something, you have to open a new tab, a new script, and you can’t see the previous data anymore. The fact that you can write SQL right in your laptop, see the data right there, scroll down, create a new cell, and add more information—it’s just great. I like this like this scrolling format,
47:00 scrolling format, that’s how my brain thinks, especially when I’m doing I’m doing data research, yeah. yeah. This is a great interface they have designed. So,, not from Microsoft. This has been around in the industry for a long time, and Microsoft just adopted it. But But when I come back to it, I think: what’s the advantage here? I don’t see her. Apart from personal preference, I don’t see any strong arguments in this article that would convince me that we absolutely need to use only use only Lakehouse. By the way, I want to make one more
47:31 I want to make one more point, Tommy. Where is the ” lakehouse only” option? Well, actually. They offer option A — — data storage only. They offer Option B —lakehouse for the bronze level, followed by silver, gold, and vault. OK. And then they give option B—bronze, silver, and gold levels in the vault. So I do n’t understand something. I can totally implement implement the bronze, silver, and gold levels inside the lakehouse. True. True. This option does n’t even exist here. And I
48:01 n’t even exist here. And I think this is a miscalculation. And especially when the question arises again: what is the advantage? What does a data warehouse, say, with a gold tier, do better than a gold tier in a lakehouse? Unless you’re performing transactional tasks, such as writing data from an application that accesses accesses your gold layer, which you wouldn’t be doing anyway. I think for most
48:31 most organizations, if you’re already doing silver and bronze tiers in Lakehouse and using the using the gold tier for reporting—well, you already have Delta Lake, the tables are already in Delta Parquet format. Okay, speaking of the gold table in Lakehouse, where exactly does Lakehouse get better? If the only argument, Mike, is that we don’t want to be too intensive on Spark or PySpark, then it’s a question of
49:02 or PySpark, then it’s a question of resources, not architecture. Let me ask you one more question, Tommy, about something else. When you use Fabric Direct Lake, yes, can you use Direct Lake with SQL data warehouse? what? I’m not entirely sure, because I’m not sure either. I actually tried to quickly Google this and thought, maybe this is where
49:32 thought, maybe this is where I’ll get I’ll get my agent involved and say, “Hey, look at the Microsoft Microsoft Learn documentation, MCP. Is this doable or not?” So, I’m not sure, Tommy, whether it’s possible to connect directly from lake to the model. Now, what I will say about Direct Lake, there are two aspects here. Direct Lake allows you to go from a semantic model, read delta tables or partitions, and immediately load them directly into memory within the model. Right? It’s essentially like instant loading of Vertipac. There are
50:02 loading of Vertipac. There are some other hidden things going on there, certain optimization processes. But, in fact, it happens very quickly. It takes delta tables and pulls them directly into the semantic model. No No import regime. It just works. However, there are thresholds for this. So there are limitations that prevent you from doing that. And if you use use too many too many rows, if your model gets too big, the model says, “Hey, you need to do a Direct Query
50:32 do a Direct Query to the SQL analytics endpoint, which is just a SQL analytics endpoint, not a data store.” So I didn’t see any mention in this article about Direct Lake. Yes. I just checked. So, for Direct Lake, you can create a new semantic model based on the repository, choose which tables to open, create the model, define relationships, measures, and even RLS, and you can also do this using Power BI Desktop
51:02 using Power BI Desktop with the repository. I just also checked what the difference is between the Lakehouse Direct and Lakehouse semantic models. There doesn’t seem to be much difference in SQL endpoint permissions for the repository. Modeling in Power BI is the same. Again . Again, the most important thing they say is: if you want you want to use to use the repository for SQL-centric teams, you already have validated validated data marts, tightly managed BI layers, you can also can also use use views, but beyond that
51:32 beyond that… … So in the comments, in the comments under this post, you can learn a little bit more about what’s going on here. Someone immediately asks,, Nikki Dev asks,, what happens if I want to use Direct? And then the Microsoft employee replies to Nikki and says, ”, the main benefit of using using Gold tier storage is that it’s not Direct Lake,” which, again, is a , is a big no for me. That’s right. . If I use If I use storage and can’t
52:03 storage and can’t use Direct Lake, then it’s not… I don’t need such storage, this architecture went in the wrong direction from the very beginning. There should be an architectural template that that only uses lakehouses. And then he goes on to say, well, you say, well,, know, unlike,, the SQL analytic endpoint is is read-only on top of,, the read-only data endpoint. Okay, fine. But fine. But Fabric storage supports insert, insert, update, delete, and merge operations. I don’t need this
52:33 I don’t need this at my gold level. level. I was just about to say that. Why would you do that? Why do I need a vault at all? Why… why should I manipulate my gold level without first passing the data through the other levels, or you could have less, could be bronze, gold, well… your reasoning makes no sense to me. I don’t like this article, this part of the article about… you’re missing something. There’s a huge, glaring hole here, and to your point, Tommy, this article is starting to sound very much like, ” Hey, use Hey, use data warehouses; we do n’t have enough
53:03 n’t have enough users, users, buy them.” Well,, Well,, here are some patterns that you might you might want to think about, and I think it’s a complete failure in my understanding. You completely misunderstand the essence of why why data warehouses exist. I think they could be here. That’s the thing. I think they could have positioned the storage differently. It’s so funny what you say about Microsoft trying to promote something. The last two clients I had were trying to run Fabric and we had to deal with quota. Both were Both were rejected. I told them to send a to send a support request to say say wait, your Fabric isn’t starting. starting. Wait. Wait.
53:34 Wait. Wait. So both customers were starting for the first time, and oh, you need to request a quota for your region and it was denied. So I said, “Send a support request,” and then Microsoft said, “Hey, do you want to do this in region X?” I’m telling you, this has happened twice in the last two months,, with two different clients, and they say, “Yeah, we’d rather you do this in region Y, because there’s actually more resources available there.” And both times they asked me, “What should I do?” I I do?” I replied, “Just say it’s very important.”
54:04 important.” And that’s all they did, they just said it was necessary, they said we really want it, and they were told—okay, you’ll get it. So I think there’s a lot of pushing here, but before you comment on that, because I want to point it out, go ahead. go ahead. Storage has its purpose, but it does n’t have to be exactly the same as what a what a lakehouse provides. If you need the ability to record back, this is a huge advantage. If you But you But you can still write back to lakehouse. I, I know that through the T- interface, the SQL Analytics endpoint, you can’t write back. But I don’t, but I don’t write
54:35 But I don’t, but I don’t write back back using it. using it. This is a lake house. This, I use Spark. I don’t understand. Before you say that, I think a repository could be great for great for applications to interact with each other, rather than as part of the same workflow for building a semantic model. Honestly, I, I’m trying to wrap my head around it. . Data warehouse is great. This is incredible for a multitude of data operations, especially if you manage data in
54:58 you manage data in many other applications. But let’s say lakehouse lakehouse connects to the repository, honestly. So here . So here is your is your ultimate ultimate reporting solution. Regarding your comment, I do not want my gold level, which is based on my reporting, to be changed or updated. This breaks my pipeline, well, the medallion medallion architecture they’re talking about. Why would you do that? Especially when you just talked about medallion medallion architecture. A repository serves many purposes and can serve better than a lakehouse, but
55:28 better than a lakehouse, but it is not necessary for creating semantic models. And I, I think that’s the most important thing I see here: build your repository. If you have, if you want programs to write back and interact with it, you have a central location. This is great. But when you’re actually going to move on to reporting, then read then read directly from lakehouse, because that’s the best option right now. There is nothing nothing better than lakehouse right now to service your golden data. Me, I would agree with that. And I think maybe the maybe the author of this article,
55:58 author of this article, now I’m,, let me quickly see if I can remember his name. The author of the article, Sidney, in the very last very last comment, seems to be receiving some criticism regarding this material. I think the article is a failure because he should have clearly stated, “Look, this article will focus on a specific design where you can integrate storage.” And at the end he adds this,, this,, comment: “this is not a dispute between the storage and the lakehouse.”
56:28 I understand, but it would have been worth at least acknowledging the real state of affairs and saying: “Listen, if you’re looking for…” Although, in essence, that’s what it is, whatever one may say. Yes., , people are trying to make sense of this information and understand what’s going on. So unless you make a compelling case for where and why storage is part of lakehouse or this architecture, then again, the bare minimum I would advise customers is: if you really need storage, then bronze and silver are lakehouse, and gold is storage. This is
57:01 storage. This is probably the most I’m willing to go to, because I really I really like the… side of Spark and notebooks. I like this like this development pattern. I like the cycle of working with him. I think it works very well. Now we have Spark 2. 0. Lots of cool new features. This field seems to be constantly evolving, evolving, creating new useful tools that significantly speed up my work and allow me to process more data. data. I don’t see storage improving in this aspect… or becoming more economical any
57:31 economical any faster than it is happening in Spark. This is a race, right? Who is doing better? It seems to me that Spark is evolving faster than repositories are changing. . But, Mike, to your point: point: you can still run T-SQL in Notepad, right? So if you’re a team that works primarily in SQL, you can… run the notebook in T-SQL, or in PySpark SQL, or in other formats—you have the choice. So I don’t think that’s an argument worth using. using. I would actually call this the wrong question. Should
58:01 question. Should I use a I use a lakehouse or a storage unit? If you’re thinking, “should I use I use storage,” you’re asking in asking in comparison to what, no matter what, and I think that’s a very important aspect, to be honest, you’re , you’re going to be thinking, should I use I use storage… or,, there’s always an “or” in that statement. And I, I think it’s not too late to do this cycle of what a data warehouse does that a data lake doesn’t do, and talk about application development, like running your applications in the warehouse,, pushing that data into the
58:31 pushing that data into the lake for reporting, instead of just saying, “We’re going to talk about a medallion medallion architecture in the warehouse,” because I, you warehouse,” because I,, I get nervous know, I get nervous when we try to cram everything into one hole, or,, when we try to cram everything into one path, even though there are things that the warehouse does better. Let’s not make it exactly one-to- one like a data lake, because that’s talking about them, talking about them exactly. Show me, show me what it is. I me what it is., and let mean, and let me decide as a system developer, as an architect for my organization.
59:01 my organization. Let me Let me decide which way we go. Will we move forward with this data this data warehouse-heavy model, or will we stay with a more data lake and notebook-oriented model? What makes sense for your business? And I think you’re right, Tommy. Present me with the facts, let me make an informed decision, and in the end, the products will show their worth. If worth. If the product is good and people people use it, it will have more demand. If a product lacks the right features or isn’t or isn’t different enough from the
59:31 different enough from the rest of the product, yes, it won’t get as much attention. He won’t get that much investment. This is normal. The last thing they’re talking about here, and I think this will be my final word, because I have to I have to mow the lawn before a big event my my wife is hosting., that’s what they’re talking about, the split- split- everything model. I want to hear your thoughts on this. They, they recommend at the very end, management recommendations. keep bronze, silver and gold levels separate, ideally in different work areas. And
60:01 work areas. And because they said, “Hey,, yeah, I, I thought that too.” So So you’re going to create a lot of a lot of workspaces if you ‘re going to separate the bronze, silver, and gold levels into different ones. These are the three data stores that you’re building. Yes, me, that was exactly my thought too. in three different work areas. There’s a lot here that I wouldn’t necessarily call the best approach when it comes to a management perspective, because Mike
60:33 you owe it to your team. Yes, I like I like separation. Yes. Yes. But I think my understanding of separation is more about models like workspaces, right? Workspaces are the control plane. Mhm. Mhm. Workspaces determine who has access to what data, they are part of this surface. surface. Now we Now we have more have more control planes, like data lake security. So now . So now you can configure security by columns and rows within the
61:03 and rows within the Lakehouse table or choose which tables the user sees. Right? So Right? So now we have security inside the workspace. Now we have security inside Lakehouse too. At Warehouse it was a little longer. You had more control over security in the data warehouse because it is more like SQL and handles its own security differently than at the Lakehouse level. So there’s probably an advantage in the data warehouse, and again, we’re starting to see some of that advantage emerge for Lakehouse as well. But I look at it as the
61:34 look at it as the surface of the workspace, which is your boundary. boundary. Who can work with this data? Of course, I think having development, testing, and production environments is important, but do I need to break everything down into bronze, silver, and gold levels? I think it’s becoming obsessive obsessive and a bit too much. And if you give your engineers the ability to work with silver data, why can’t they see the raw data? Unless there’s something very confidential there that they need to
62:04 that they need to have access to somewhere else. else. In most cases, they will work with full volume. I would say that most of your your data engineers will work at all three levels. The data will progress from bronze to bronze to gold. You don’t want to hire a separate separate data engineer for each layer. You will want want data engineers working on all three layers. That’s how I’ve traditionally seen it in many in many organizations. So the recommendation to break down the bronze, silver, and gold layers by work areas seems intrusive and disconnected from the reality of development.
62:36 Yes. So, Mike, they actually mentioned at the end that there would be four more articles. This is a five-part series. The second article that will come out is how to build bronze, silver, and gold tiers in a data warehouse, which, again, will be interesting to compare to Lakehouse. The third part will be dedicated to best practices for medallion architecture. The fourth part is security and managing your layers. Part Five: Optimizing the Performance of Your Medallion Pipeline. Of all of this,
63:06 . Of all of this, I’m most I’m most interested in performance optimization, as I as I think it will be the most different from what you do at you do at Lakehouse. So, I’m sure we’ll discuss this in more detail, but I think Mike will also find the security issue quite interesting. I think the last two points, four and five, five, yes, it would be interesting to see how they look at the Medallion architecture and how they apply it to Lakehouse. I guess if I were to compare this to SQL data warehouses, most of them have a staging stage and a final table, right? There’s actually
63:37 right? There’s actually no no Medallion architecture there. The Medallion architecture in the context of a data warehouse is still a Lakehouse feature. This is a concept that mostly comes from Databricks. In fact, I don’t think the Lakehouse architecture meshes very well with the data warehouse architecture in this case. So it seems a little out of place. If you are using staging staging and then final tables, then a data warehouse might be a more appropriate plan, right? In fact, I think the think the system design is potentially even different,
64:07 even different, because you’re used to using SQL on on- premises servers and want the same hyperscale SQL storage, wanting to get that experience in Fabric. I think that’s a more compelling story than trying to trying to use use this element of Medallion architecture here. So to me, it seems a bit far-fetched. . I’m curious to see if to see if they listen to the feedback and if the articles will change. So, So, answering our questions, let’s get back to the topic of this video. Do
64:37 this video. Do Medallion templates change? No, they don’t change. Some of the tools in this article are trying to change the emphasis a bit, aren’t they? The tools The tools we use in we use in these templates may vary slightly, but the concept from raw data or Bronze level to the point where I am I am ready to share this data with other users at users at Gold level remains remains the same. It’s just the way things work. And I really like this article. I think it’s good. But she is really good. In any case, I want to
65:08 any case, I want to thank you for listening to this article and discussing it with us today. We really appreciate all of you. Tommy, where else can I find this podcast? You can find us on Apple, Spotify, or wherever you listen to podcasts. Don’t forget to subscribe and rate. This helps us a lot. Do you have a question, idea, or topic for future episodes? Go to the PowerBI tipsodcast. Leave your name and a great question. And finally, finally, join us join us every Tuesday and Thursday at AM Central Time on PowerBI tips social media. Thank you all very much, see you
65:38 see you next time. Explicit measures. Make it louder. Let it be higher. Tommy and Mike light up the sky. Dance all day, laugh with us. Fabric and artificial intelligence—feel it. Explicit measures. Now add a beat. The crowd just explodes. Explicit measures. Explicit measures.
Thank You
Want to catch us live? Join every Tuesday and Thursday at 7:30 AM Central on YouTube and LinkedIn.
Got a question? Head to powerbi.tips/empodcast and submit your topic ideas.
Listen on Spotify, Apple Podcasts, or wherever you get your podcasts.