Skip to main content
BI4ALL BI4ALL
  • Expertise
    • Artificial Intelligence
    • Data Strategy & Governance
    • Data Visualisation
    • Low Code & Automation
    • Modern BI & Big Data
    • R&D Software Engineering
    • PMO, BA & UX/ UI Design
  • Knowledge Centre
    • Blog
    • Industry
    • Customer Success
    • Tech Talks
  • About Us
    • Board
    • History
    • Sustainability
    • Awards
    • Media Centre
    • Partners
  • Careers
  • Contacts
English
Português
Last Page:
    Knowledge Center
  • SWAT Kimball automating KPI to star schema design with GenAI agents

SWAT Kimball automating KPI to star schema design with GenAI agents

Página Anterior: Blog
  • Knowledge Center
  • Blog
  • Fabric: nova plataforma de análise de dados
1 Junho 2023

Fabric: nova plataforma de análise de dados

Placeholder Image Alt
  • Knowledge Centre
  • SWAT Kimball automating KPI to star schema design with GenAI agents
30 September 2026

SWAT Kimball automating KPI to star schema design with GenAI agents

SWAT Kimball automating KPI to star schema design with GenAI agents

Turning a list of business KPIs into a robust star schema and production‑ready Spark code is one of the most critical and error‑prone parts of any modern BI initiative. SWAT Kimball is a group of GenAI agents designed to automate large parts of this journey—from initial KPI descriptions all the way down to dimensional models, notebooks and Power BI artefacts—while keeping analysts and engineers firmly in the loop.

Business inputs and technical context

SWAT Kimball assumes two main inputs: a business document describing KPIs and a representation of the available data sources. In practice, the KPI document is often a list of KPIs with names, descriptions and mathematical formulas for metrics such as “Total Revenue”, “Gross Profit Margin” or “Sales Growth Rate”. The data input is typically a “print‑out” of one or more databases, containing table names, column names and data types, sometimes derived from a data catalogue.

The agents are optimised to read Markdown, which is compact, token‑efficient and less error‑prone for LLMs than formats like Excel. PDFs and Word documents can still be used, but it is recommended to convert them to Markdown first using an external LLM such as Claude, making the subsequent agent processing more reliable. JSON is also preferred over Excel for reference lists—like catalogues of known KPIs—because it is more structured and LLM‑friendly.

Although early experiments used sample databases such as AdventureWorks and Worldwide Importers, the agents behave similarly when working against real‑world datasets or a merge of multiple databases, which is more realistic for enterprise KPIs that span several systems.

KPI Wizard: mapping business metrics to data

The entry point in SWAT Kimball is KPI Wizard, which consumes the KPI document and the database print‑out. For each KPI, it identifies candidate source fields in the underlying tables and calculates a confidence score expressing how likely each mapping is to be correct. It produces a structured Markdown report showing, for example, that “Total Revenue” should use the ExtendedPrice column in the SalesInvoiceLines table with a 100% confidence.

When confidence is lower, KPI Wizard explicitly flags this and often includes commentary on additional filters or conditions that might be necessary, for example suggesting filters by transaction type to isolate actual sales. KPIs whose mappings fall below a defined threshold (for instance 95%) are grouped into a section for stakeholder validation.

Beyond mapping requested KPIs, KPI Wizard also behaves like an experienced business analyst by proposing new KPIs that would be valuable given the data. Using a JSON‑based list of standard KPIs as a knowledge base, it might suggest metrics like Sales Growth Rate or Gross Profit Margin, with mathematical formulas and detailed business definitions explaining why they are useful.

Confidence Judge: verifying the maths

LLMs are not perfectly deterministic and can sometimes miscalculate or mis‑apply their own reasoning steps. To mitigate this, the Confidence Judge acts purely as a validator of the KPI Wizard’s numerical reasoning, focusing only on the confidence calculations and related logic. It re‑checks each computation, corrects any mistakes directly in the document, and produces a summary of changes for auditability.

This design pattern — “LLM as judge”— is used multiple times across SWAT Kimball to cross‑check and refine outputs from other agents. It reduces the risk of subtle computational errors propagating downstream into modelling and implementation.

Model Storyteller: generating the logical star schema

Once stakeholders validate the KPI mappings, Model Storyteller takes the baton. It reads the KPI Wizard output and generates a comprehensive logical dimensional model that serves as the blueprint for implementation. This includes:

  • The list of tables to use in the model, aligned with the earlier mappings.
  • A base matrix of required dimensions and fact tables.
  • A proposed star schema, with clear separation between dimensions and facts.
  • For each dimension: the SCD strategy (Type 0, Type 1, Type 2, or hybrids), surrogate key strategy, natural key handling, and full list of attributes.
  • For each fact table: the grain, the set of measures, degenerate dimensions, role‑playing dimensions, and classification of measures as additive, semi‑additive or non‑additive.

The output is a rich narrative document that not only specifies structures but also explains design choices, such as why certain dimensions should be SCD Type 2 or which KPIs each fact table is intended to support.

KPI Sentinel: quality gate for the model

Because the logical model is so critical, SWAT Kimball adds another layer of validation via KPI Sentinel. This agent checks whether all requested KPIs are properly represented in the proposed model and whether there are obvious gaps or inconsistencies. It returns a pass/fail assessment with warnings, for instance, where the information seems insufficient to build a particular dimension or where assumptions may need clarification with stakeholders.

At this stage, the human BA, architect or lead data modeller reviews the report and confirms with business stakeholders that the logical model is acceptable before proceeding to implementation. This human‑in‑the‑loop checkpoint is central to the philosophy of treating agents as accelerators rather than autonomous designers.

DIMS Builder and Facts Builder: from design to PySpark code

Once the logical model is approved, SWAT Kimball moves into the implementation phase with DIMS Builder and Facts Builder. These agents read the star schema design and generate PySpark notebooks for each dimension and fact table, following internal frameworks and best practices. There are typically separate agents or modes optimised for Databricks and for Fabric, though in practice any Spark‑based platform that supports PySpark can work.

Because the team already has established patterns and templates for dimensions and facts, the agents can produce highly consistent code aligned with those templates. This consistency significantly reduces the variation in implementations and makes downstream maintenance and reviews easier.

Notebook Inspector and Functional Sentinel: closing the loop

The generated notebooks do not go straight into production. Notebook Inspector aggregates code from all the notebooks and produces an executive summary describing what each one does. This high‑level explanation helps reviewers quickly understand the overall design and focus on the most critical parts.

Functional Sentinel then performs two checks:

  • It validates whether the code follows the agreed coding templates and best practices.
  • It checks whether the code implements what was requested in the logical model and KPI definitions.

Again, using the “LLM as judge” pattern, it produces a report stating whether the implementation is approved or not, along with guidance on what needs to be improved. Any issues can be iteratively fixed either manually by engineers or by re‑invoking specific agents with refined inputs, depending on the complexity of the change.

Power BI agents: semantic models and reports

Beyond Spark code and star schemas, SWAT Kimball also includes agents for Power BI artefacts. Starting from the Model Storyteller’s output, these agents:

  • Derive the subset of the logical model relevant for the Power BI semantic model.
  • Create and save a Power BI semantic model into a project folder.
  • Use a semantic model builder/inspector agent to assess whether the model design fits the business requirements and to generate a validation report.
  • Build an initial report layout using a trio of agents: a background builder (for visual background images), a visuals builder (for charts and tables) and a theme builder (for branding, colours and logos).

The visuals builder can already produce a useful first draft of the report, although positioning and layout still require manual refinement. Even if only 30% of the final report is “right first time”, that still removes a significant amount of manual work for report developers.

Practical considerations and ecosystem choices

SWAT Kimball has been implemented and tested using GitHub Copilot as the main LLM integration layer, leveraging Claude Sonnet 4.6 where it provides better performance and lower token usage than more expensive alternatives. The team is also actively experimenting with Databricks Genie Code, which offers tight integration with Databricks notebooks and system tables, and supports a skills‑based approach rather than huge instruction‑only prompts.

A key design decision is when to encode behaviour as long, robust instructions versus smaller reusable skills:

  • Complex, relatively stable flows (like the full logical model generation) work well with instruction‑heavy agents.
  • Highly interactive, iterative tasks (such as “create just this one new dimension”) may be better served by smaller skills that can be called ad‑hoc.

As LLM providers refine pricing (for example, GitHub Copilot moving towards token‑based pricing tiers) and as tools like Genie Code mature, the CoE continues to benchmark cost, performance and control.

Conclusion: structured acceleration, not magic

SWAT Kimball demonstrates that GenAI agents can take on a large share of the heavy lifting from KPI definition to star schema and Spark code, without removing the need for expert human judgement. By structuring work into specialised agents, adding LLM‑as‑judge validators, and keeping humans in the loop at key checkpoints, organisations can speed up BI projects, reduce inconsistencies and improve documentation—while retaining full control over their data architectures.

Author

Martim Dornelas

Martim Dornelas

Associate Specialist

Share

Suggested Content

GenAI Agents: the new accelerator for Modern BI and Big Data projects
Blog AI & Data Science

GenAI Agents: the new accelerator for Modern BI and Big Data projects

Within Modern BI and Big Data programmes, organisations still lose enormous time on repetitive, technically heavy work: gathering requirements, mapping KPIs, designing dimensional models, generating Spark code, building Power BI models, and migrating legacy workloads. To address this, the CoE has created A Team – a family of GenAI‑powered agents designed as accelerators, not replacements, across the full lifecycle of data projects.

The Challenge of Ensuring High-Quality Power BI Semantic Models
Blog Data Visualisation

The Challenge of Ensuring High-Quality Power BI Semantic Models

As Power BI adoption grows, so does the number of semantic models, and with it, the challenge of maintaining quality and consistency. Teams working on multiple projects often adopt different approaches to naming, design, DAX, and performance, leading to models that are harder to maintain and operate efficiently.

AI Governance – From Compliance Obligation to Competitive Advantage
Blog AI & Data Science

AI Governance – From Compliance Obligation to Competitive Advantage

Most organizations know they need to govern their AI. Fewer have turned that intention into a system that regulators, auditors, and business leaders can all trust.

From Locked Data to Governed Access: GxP-Aligned Data Access for Clinical & R&D Webinar
Tech Talks Data Strategy & Data Governance

From Locked Data to Governed Access: GxP-Aligned Data Access for Clinical & R&D Webinar

In this webinar, BI4ALL and Immuta show how to scale AI in one of the world's most regulated industries without drowning in tickets, manual approvals, and access-control chaos.

The Report is Correct, the Question is Wrong
Blog Data Visualisation

The Report is Correct, the Question is Wrong

It is easy to conflate three different discussions when talking about dashboards that do not deliver: productivity driven by AI, information design, and governance over who decides what is left out. These are distinct problems with distinct solutions. Solving only one of them is not enough, which is why this article addresses them in the right order.

Inside Saint-Gobain’s Data Maturity Framework Journey with BI4ALL
Tech Talks Data Strategy & Data Governance

Inside Saint-Gobain’s Data Maturity Framework Journey with BI4ALL

In this DAMA Portugal (Lisbon Chapter) webinar, Orquídea Fonseca (Data Strategy Lead, Saint-Gobain) and Sandro Scordo (Head of Data & AI Strategy and Governance, BI4ALL) share the story behind building an enterprise data maturity assessment across a decentralised, 80-country organisation.

video title

Lets Start

Got a question? Want to start a new project?
Contact us

Menu

  • Expertise
  • Knowledge Centre
  • About Us
  • Careers
  • Contacts

Newsletter

Keep up to date and drive success with innovation
Newsletter
PRR - Plano de Recuperação e Resiliência. Financiado pela União Europeia - NextGenerationEU

2026 All rights reserved

Privacy and Data Protection Policy Information Security Policy
URS - ISO 27001
URS - ISO 27701
Cookies Settings

BI4ALL may use cookies to memorise your login data, collect statistics to optimise the functionality of the website and to carry out marketing actions based on your interests.
You can customise the cookies used in .

Cookies options

These cookies are essential to provide services available on our website and to enable you to use certain features on our website. Without these cookies, we cannot provide certain services on our website.

These cookies are used to provide a more personalised experience on our website and to remember the choices you make when using our website.

These cookies are used to recognise visitors when they return to our website. This enables us to personalise the content of the website for you, greet you by name and remember your preferences (for example, your choice of language or region).

These cookies are used to protect the security of our website and your data. This includes cookies that are used to enable you to log into secure areas of our website.

These cookies are used to collect information to analyse traffic on our website and understand how visitors are using our website. For example, these cookies can measure factors such as time spent on the website or pages visited, which will allow us to understand how we can improve our website for users. The information collected through these measurement and performance cookies does not identify any individual visitor.

These cookies are used to deliver advertisements that are more relevant to you and your interests. They are also used to limit the number of times you see an advertisement and to help measure the effectiveness of an advertising campaign. They may be placed by us or by third parties with our permission. They remember that you have visited a website and this information is shared with other organisations, such as advertisers.

Política de Privacidade