Deployment¶
This guide covers the supported retail-setup deploy workflow for provisioning,
publishing, updating, and recovering a Microsoft Fabric Retail Demo workspace.
Complete Getting started first if the repository has not
been configured and rendered.
Audience: operators, entry-level developers, and support engineers. Business analysts can use the deployment outcome and troubleshooting sections without following the advanced authentication details.
See the plain-language glossary for terms such as Terraform,
preflight, terminal success, Direct Lake, and DEGRADED.
Exact command order and failure behavior are owned by the deployment framework specification.
Confirm the target
--recreate destroys the configured workspace and every item in it. The
deploy polls Fabric until the workspace name is absent and fails before
apply if deletion times out or pagination is incomplete. Use recreate only
for a disposable demo workspace after confirming the tenant, environment,
workspace name, and Terraform state.
Deployment outcome¶
Every deployment resolves deployment.profile from
contracts/retail-demo.json. The manifest selects stable logical asset IDs,
existing notebook groups, existing pipeline folders, and existing KQL scripts.
It does not redefine physical items or schemas: deploy/config/deploy.yml,
deploy/scripts/build_artifacts.py, and the source folders remain authoritative
for those definitions.
The current deploy plan:
- resolves the exact profile inventory;
- runs local and live tenant/capacity preflight before Terraform destroy/apply, publication, or KQL;
- generates profile-aware Terraform and
fabric-cicdinputs; - initializes and applies Terraform unless skipped;
- captures the selected Terraform-owned identifiers;
- stages and publishes infrastructure without Reporting;
- executes only the selected ordered KQL scripts and validates infrastructure;
- for gated profiles, waits for setup and required ML validation to complete;
- stages and publishes Reporting only after required ML terminal success;
- runs selected optional/experimental ML after Reporting;
- tells the Lakehouse SQL endpoint about newly created tables so Power BI and readiness queries can find them;
- for
full-demo, runs the deployed ontology notebook to create or revalidate the target ontology and its derived graph; - publishes both Data Agents and the complete 48-item task flow, then reads the graph back and compares every task, item binding, and edge;
- runs read-only readiness verification over the complete selected profile.
Each live run atomically updates
deploy/.generated/<env>/deploy-run.json. Inspect it for target names,
the manifest version/hash, profile support and expected counts/folders,
core/optional/preview/manual boundaries, exact asset/group/pipeline/KQL
inventory, required/optional classification,
failed step, exit code, and final SUCCEEDED, DEGRADED, or FAILED status.
Artifact build steps link phase-specific artifact-inventory-<phase>.json
evidence. Credential-like error text is redacted and raw subprocess output is
not stored.
The final journal status is:
SUCCEEDED: every selected required and optional check passed.DEGRADED: required capabilities passed, but at least one optional capability has failed, stale, or missing evidence. The required demo path is usable; inspect the linked report before presenting the affected optional feature.FAILED: a required capability failed or could not provide evidence. Do not present the workspace as ready.
Profiles¶
core is the conservative default for direct retail-setup commands,
standard is the supported opt-in Reporting path, and full-demo contains
live-checked preview/manual surfaces. The guided setup.ps1 and setup.sh
wrappers default to full-demo so that a first-time guided run offers the
complete demo. The
workspace and profile inventory is the single
canonical guide for exact asset IDs, groups, pipelines, KQL scripts, staged
counts, folders, and support status.
All profiles exclude the destructive reset group. Pipeline selection is
explicit; a notebook reference does not implicitly select a pipeline. No
profile enables a schedule or starts the long-running stream.
Runtime, capacity, cost, preview, and manual boundaries¶
The dry-run prints the manifest-owned runtime, capacity, cost, preview, and
manual boundaries for the resolved profile. A live full-demo run reads
Fabric tenant settings and capacity metadata before mutation. Reporting
remains absent unless the current required ML pipeline run, including the
runtime contract validator, finishes successfully.
Prerequisites¶
Local tools¶
- Git
- Python 3.11 or later
- Terraform
>= 1.8, < 2.0 retail-setupazure-identityazure-kusto-datafabric-cicdpyodbc- the locked deploy dependency set, which includes the bundled
mssql-pythonSQL driver on Windows and Linux; an installed Microsoft ODBC Driver 17 or 18 is used when available and remains required on macOS - Azure CLI for the default auth mode, or Azure PowerShell for the lower-level alternative
Install the Python dependencies manually when you are not using the guided bootstrap:
python -m pip install --require-hashes -r .\utility\requirements-deploy.txt
python -m pip install --no-deps -e .\utility
Fabric target¶
The operator must be able to:
- use the selected tenant and active Fabric capacity;
- create or update the target workspace and Fabric items;
- create or publish every item type selected by the chosen profile;
- apply KQL schema when the profile includes Eventhouse;
- perform the profile's documented manual steps.
Use a dedicated workspace. Review workspace and item roles before sharing the demo, because synthetic customer-like fields still require governance.
Authentication¶
Azure CLI¶
Azure CLI is the default and the only mode the guided bootstrap installs and checks automatically.
Set:
retail-setup deploy rejects a configured tenant that differs from the active
Azure CLI tenant.
Azure PowerShell¶
Most users do not need this section
Use Azure CLI unless your organization specifically requires Azure PowerShell or an automated service identity. Azure PowerShell does not provide credentials to the Fabric Terraform provider, so this path needs additional identity configuration.
For a manually prepared workstation:
Set:
The deploy REST and KQL helpers use AzurePowerShellCredential with the
configured tenant. The Fabric Terraform provider cannot consume an Azure
PowerShell session. Therefore this mode supports either:
- Python-only publication with
--skip-terraformand previously captured, validated outputs; or - normal Terraform with exactly one provider-supported service-principal
secret/certificate, OpenID Connect (OIDC) federated identity, or
managed-identity credential configured through
the provider's
FABRIC_*orARM_*environment variables.
Without one of those boundaries, apply and destroy fail before any command
runs. Terraform always receives the configured tenant_id; Azure CLI use is
explicitly disabled in Azure PowerShell mode, so an unrelated CLI login is
never used silently. A conflicting FABRIC_TENANT_ID or ARM_TENANT_ID also
fails before mutation. The guided scripts/setup.* prerequisite check still
expects Azure CLI, so invoke retail-setup directly for this advanced flow.
Configure an environment¶
The workspace name defines the environment. configure normalizes the name to
lowercase hyphenated text and omits a leading retail-demo- prefix. For
example, workspace retail-demo-alice uses environment alice. Each workspace
gets its own ignored target overlay, generated inputs, Terraform data
directory, state, outputs, and run journal.
Run:
retail-setup configure --workspace-name retail-demo-alice --profile core
retail-setup render --env alice
configure prints the derived environment key. Configure another workspace to
create another local environment without replacing the first.
Valid profiles are core, standard, and full-demo. Omitting --profile
prompts with core as the default. The selected value is stored in the
workspace environment overlay as deployment.profile.
Important configuration boundaries:
| Path | Purpose | Tracked |
|---|---|---|
deploy/config/deploy.yml |
Shared deployment defaults | Yes |
deploy/config/environments/<env>.yml |
Workspace-specific target overlay | No |
utility/config.yaml |
Local generation scale and seed | No |
utility/out/ |
Rendered workspace-specific notebooks | No |
Keep the Eventhouse and KQL database display names aligned. The Eventhouse creates one default KQL database with the Eventhouse display name; provisioning a separately named database is not a supported topology.
Starter pool or custom pool¶
core and standard use the workspace starter pool. Only full-demo selects
the custom pool. Live preflight requires an active F64-or-larger capacity and
checks configured node vCores against the capacity's base vCores. Legacy
--use-custom-spark-pool and
spark.use_custom_pool overrides fail clearly; select the profile instead.
Review the source-defined pool configuration against the target capacity.
Preview the plan¶
The preview prints the resolved inventory, profile blockers, ordered commands, and confirmation gates. It validates the existing environment configuration and authentication boundary but does not execute subprocesses or:
- contact identity endpoints or validate live permissions;
- contact Fabric;
- validate capacity state;
- prove live generated IDs unless
--skip-terraformis selected; - run Terraform plan as a separate command.
With --skip-terraform, dry-run validates the existing captured outputs just
like a live run. Malformed or invalid existing YAML always fails; the default
core fallback applies only when a legacy environment overlay is genuinely
absent. The CLI asks for confirmation before invoking Terraform apply. Terraform then
prints its change preview and proceeds with -auto-approve; there is no
separate plan review before the confirmation.
Preflight¶
Preflight means “check before making changes.” On a live run, local profile
preflight checks manifest/source pointers,
rendered notebooks, selected pipeline references, KQL sources, selected Power
BI/agent/task-flow sources, custom-pool configuration, prior local Terraform
state and captured profile outputs, and blockers. A second read-only live
preflight checks the selected capacity's identity, active state, capacity
size/tier, region, and Spark worker fit. For full-demo, it also reads the
Fabric admin tenant-settings
API and verifies the required Ontology and Data Agent switches.
After Terraform (or immediately when --skip-terraform is selected), target
access validation requires captured workspace IDs to match managed resource
IDs in state and a live workspace with the configured name that the operator
can read.
Correct environments proceed without prompts. If a required tenant switch is disabled, interactive deploy prints the exact Admin Portal setting and offers to recheck after an administrator changes it. Fabric documents tenant-setting inspection but no write API, so setup never silently changes tenant policy. Noninteractive deploy and invalid/inaccessible capacity fail with exact remediation before mutation.
Whenever environment-local Terraform state exists, missing or stale captured
outputs and missing state profile evidence fail closed, including for state
with no managed instances. A profile downgrade that would delete an Eventhouse
or custom Spark pool is rejected unless the operator explicitly selects
--recreate.
Run a normal deployment¶
Interactive:
Pre-confirm the Terraform apply gate:
Without --yes, an existing workspace detected by display name produces a
choice:
- keep it and update in place; or
- reset it and follow the recreate path.
With --yes, the CLI does not perform that interactive existing-workspace
check. It still runs mandatory setup and required ML gates for profiles that
publish Reporting.
Understand generated files¶
Deployment writes or refreshes:
| Path | Content | Git status |
|---|---|---|
deploy/.generated/<env>/terraform.tfvars |
Terraform input generated from merged YAML | Ignored |
deploy/.generated/<env>/terraform.tfstate |
Environment-local Terraform state | Ignored |
deploy/.generated/<env>/.terraform/ |
Environment-local Terraform data directory | Ignored |
deploy/.generated/<env>/fabric-cicd/config.yml |
fabric-cicd environment and item scope |
Ignored |
deploy/.generated/<env>/fabric-cicd/parameter.yml |
Workspace, item, OneLake, KQL, and agent rewrites | Ignored |
deploy/.generated/<env>/terraform-output.json |
Captured live Fabric item identifiers | Ignored |
deploy/.generated/<env>/database.kql |
Combined ordered KQL script | Ignored |
deploy/workspace/ |
Staged Fabric item folders | Ignored except .gitkeep |
Terraform operations for different environments use different backend and data
paths. Full publication still shares deploy/workspace/ staging; use separate
checkouts for concurrent full deploys. Never commit local target overlays,
credentials, tokens, or generated bindings.
If an older checkout left state at deploy/terraform/terraform.tfstate, deploy
stops before Terraform. Verify which workspace owns that state, then move it to
deploy/.generated/<env>/terraform.tfstate. Never copy one legacy state into
multiple environments.
Existing workspaces¶
For a workspace that Terraform should resolve rather than create, set
workspace.existing_id in deployment configuration and run the normal
Terraform path. This resolves the workspace only; it does not automatically
discover or import pre-existing Lakehouse, Eventhouse, role-assignment, or
Spark resources into Terraform state. Avoid name collisions and import or
reconcile existing child resources deliberately before apply.
Use --skip-terraform only when all required resources already exist and
deploy/.generated/<env>/terraform-output.json contains correct identifiers
from an earlier deployment:
Before publication, the command verifies the captured environment, workspace
and resource names, required Fabric IDs, and absence of placeholder IDs. A
full-demo capture must also include the custom Spark pool ID. A first-time or
wrong-workspace --skip-terraform run fails before mutation.
--skip-terraform cannot be combined with --recreate.
Recreate a disposable workspace¶
Preview:
Execute:
The current sequence is profile preflight, configuration generation, Terraform initialization, Terraform destroy, bounded polling until the workspace name is absent, Terraform apply, and normal publication. Preflight therefore cannot destroy the workspace on failure. Preserve run evidence before retrying if Fabric does not release the workspace within the documented timeout.
Do not use 99-reset-lakehouse as part of normal deployment. It is a manual,
destructive data reset asset.
Post-deploy work¶
Understand local validation¶
deploy.scripts.validate_deployment checks generated files, YAML, staging, and
placeholder rewrites. It does not query the live workspace. A successful deploy
still requires live item, binding, KQL, run, and data checks.
Generate data¶
For core, run setup notebooks 01 through 04 in order. core intentionally
has no automatic pipeline run.
For standard and full-demo, the deploy command already waits for
setup-pipeline and ml-required, publishes Reporting only after both
succeed, and waits for selected post-Reporting ML pipelines. --yes does not
skip these gates. No profile starts the long-running stream or enables an automatic schedule;
the committed daily-maintenance schedule is disabled.
To retry one pipeline deliberately, run:
python -m deploy.scripts.run_pipeline `
--environment alice `
--pipeline setup-pipeline `
--auth-mode azure_cli `
--tenant-id <tenant-id> `
--wait
--wait polls the exact submitted run to a bounded terminal state. Omit it
only for an intentional asynchronous manual trigger.
If a deployment stopped but every selected setup/ML pipeline was later recovered manually, adopt the exact successful Fabric job IDs instead of rerunning hours of Spark work:
- Open the Fabric Monitoring hub.
- Open each pipeline's run history.
- Copy the job ID from the successful run you want to adopt.
- Supply every pipeline selected by the stored profile.
| Profile | Required run IDs |
|---|---|
core |
None; setup notebooks are run manually. |
standard |
setup-pipeline and ml-required |
full-demo |
setup-pipeline, ml-required, ml-optional, and ml-experimental |
The example below is for full-demo. For standard, omit the
ml-optional and ml-experimental lines.
python -m deploy.scripts.adopt_pipeline_runs `
--environment alice `
--run setup-pipeline=<job-id> `
--run ml-required=<job-id> `
--run ml-optional=<job-id> `
--run ml-experimental=<job-id>
Supply every pipeline selected by the profile. Recovery validates the captured
workspace against the live configured target, confirms that setup finished
before required ML and required ML finished before optional ML, and accepts
only explicitly named, completed runs from the last seven days. It then writes
a current deploy-run.json, runs readiness, and finalizes the journal only
after that verification result is known.
Synchronize the Lakehouse SQL endpoint¶
Reporting profiles do this automatically after setup and machine learning. The step tells Fabric's SQL endpoint that newly created Lakehouse tables exist, which allows Power BI and readiness queries to discover them. The step is optional because required-table availability is checked again by fail-closed readiness.
Run it manually when tables exist in the Lakehouse but Power BI or SQL queries cannot see them:
The command waits for Fabric's metadata operation to finish and reports any table that did not synchronize.
The same artifact build also sets the deployed report's saved date filters to
the month containing the configured history end_date. To change that default,
update the history configuration, rerun retail-setup render, and deploy again.
Recover the ontology and task flow¶
Normal full-demo deployment completes this phase automatically. It runs the
deployed 30-create-ontology notebook to create or revalidate
RetailOntology_AutoGen, waits for the item to appear, publishes both Data
Agents, and deploys the
source-controlled task flow. The task-flow write is read back and compared
with all 11 tasks, 48 selected item references, and 11 edges before the command
reports success.
Use this idempotent recovery command only when that final phase was interrupted or an operator removed one of its items:
The command reuses the same completion path. A Data Agent is a conversational Fabric experience that answers questions from an approved data source; the task flow is the workspace navigation map. Missing selected references, ambiguous task-flow records, or a graph that does not persist exactly all fail closed.
Validate live readiness¶
For standard, deployment runs the live verifier after the normal exact-run
setup and ML gates. full-demo runs the complete verifier after ontology,
Data Agent, and task-flow publication. Verification is read-only and never
adds another pipeline run. A required failure or unknown
result fails deployment; an optional failure or unknown result records
DEGRADED in deploy-run.json. The journal links:
For core, run setup notebooks 01-04 manually, then run:
You can rerun the read-only command for any profile. Add --run-pipeline only
to explicitly start and wait for the profile-required post-publish pipeline:
The flag is invalid for core. It does not start streaming or other pipelines.
See the operations readiness checklist
for prerequisites, exit codes, taxonomy, freshness windows, and report
contents. Offline validate_deployment.py behavior is unchanged and does not
substitute for this live check.
If the optional stream has never been started, three streaming checks are
expected to be UNKNOWN: Silver watermarks, Eventhouse ingestion, and
checkpoint identity. Required capabilities can still pass, so the overall
result is DEGRADED, not FAILED.
Update an existing deployment¶
For normal source or configuration changes:
- pull or check out the intended revision;
- rerun
retail-setup configurewhen target or generation values changed; - rerun
retail-setup render; - inspect
retail-setup deploy --env <env> --dry-run; - run the normal deploy without recreate;
- rerun only the data or optional workloads affected by the change;
- validate live bindings and freshness.
Environment-local Terraform paths make switching between workspaces safe.
Preserve each environment's ignored .generated/<env>/ directory when its
state and output evidence must survive cleanup.
Troubleshooting¶
| Symptom | Action |
|---|---|
| Azure CLI tenant mismatch | Run az login --tenant <configured-tenant> and confirm az account show. |
| Azure PowerShell Terraform boundary | Use validated --skip-terraform outputs, switch to Azure CLI, or configure exactly one provider-supported service-principal, OIDC, or managed-identity credential. |
| Capacity not found or inactive | Confirm the capacity display name, state, tenant, and operator access. |
| Custom Spark pool fails | Select core or standard, or use an active F64-or-larger capacity that passes live vCore sizing. Legacy pool overrides are unsupported. |
| Required ML gate fails or is not terminal-successful | Inspect the exact ml-required pipeline/notebook run and 15-validate-required-ml-contract; Reporting is intentionally unpublished. |
| Full-demo tenant preflight fails | Sign in as a Fabric administrator, enable the exact named Ontology/Data Agent tenant setting in the Admin Portal, then recheck. |
| Terraform executable missing | Install Terraform or use --skip-terraform only with valid prior outputs. |
| Workspace already exists | Update in place, configure workspace.existing_id, or explicitly recreate a disposable target. |
| Captured workspace is unreadable or mismatched | Run Terraform refresh/apply for the environment, recapture terraform output -json, and confirm the operator has a workspace role. |
| Fabric item publish fails | Inspect the failing item type and generated .generated/<env>/fabric-cicd/parameter.yml; do not treat later steps as completed. |
| KQL application fails | Inspect deploy/.generated/<env>/database.kql, target IDs, database name, and operator permissions; rerun the ordered script as one database script. |
| Local validation passes but workspace is unusable | Perform the live checks in the operations guide; local validation is offline only. |
Readiness verification is UNKNOWN |
Install the locked deploy dependency set, confirm identity permissions, and restore the matching Terraform output and deployment journal evidence. mssql-python provides a bundled SQL driver on Windows and Linux; macOS requires system ODBC Driver 17 or 18. |
Readiness verification is DEGRADED |
Read the optional failed/unknown check rows; do not present that optional capability until its evidence passes. |
| SQL endpoint metadata is missing or stale | Run python -m deploy.scripts.refresh_sql_endpoint --environment <env>, wait for all required tables, then reload Power BI and rerun readiness. |
| Power BI opens on the wrong month | Confirm the configured history end date, rerun retail-setup render --env <env>, and deploy again. The staged report uses the latest generated month, not a hard-coded calendar month. |
| Setup pipeline fails | Inspect the exact run and retry with deploy.scripts.run_pipeline --wait; required ML and Reporting remain unpublished. |
| Ontology, Data Agent, or task-flow item is absent | Run retail-setup post-ontology --env <env>. It idempotently creates the ontology when needed, republishes both agents, writes the full graph, and verifies persistence. |
| Live rows are absent | Check notebook sink parameters, Query URI resolution, KQL permissions, connector errors, and ingestion timestamps. |
--skip-terraform rejects outputs |
Use outputs captured for the same environment and workspace; run the normal Terraform path to refresh missing or stale identities. |
Current limitations¶
IMP-* identifiers are named improvement items in the technical backlogs.
Open a linked item to see its exact remaining acceptance criteria.
- Live alternate-authentication and renamed-target smoke verification: IMP-001
- Separate live
coreandstandardprofile proof keeps IMP-012 open. - Recent optional streaming freshness evidence: IMP-013
These limitations are part of the current contract. Do not describe local staging success, a task-flow node, a pipeline trigger, or a checked-in optional asset as proof of a usable live deployment.