Section 04
Data Architecture
The engine room of the Vision Statement. Every campaign audience, revenue metric and market-entry decision the vision depends on is meant to be produced here — a medallion lakehouse on AWS, serverless-first and deliberately staged so that spend scales with proven value, with consent and data-residency controls built in from day one. The committed MVP stack is costed in the reconciliation below.
Which stack is actually committed — the reconciliation
MVP: S3 · Glue Data Catalog · Athena · dbt Core · QuickSight
At MVP scale the committed bill of materials is S3, Glue Data Catalog free tier, Athena, dbt Core and QuickSight (Metabase OSS self-hosted as the A$0 alternative, at the vendor’s own published price of zero — its pricing page publishes “Free unlimited users” for the self-hosted Open Source edition) at a calculated A$46.43/mo run cost and 8.0 setup days, 7.0 of them committable before the first-party data is disclosed. The run cost is calculated line by line: S3 storage 5 GB at US$0.025/GB-month = US$0.125, Athena 10 GB scanned at US$5.00/TB = US$0.05, and QuickSight 1 author at US$24/mo plus 3 readers at US$3/mo = US$33.00, with Glue Data Catalog on its published free tier and dbt Core, open source, at A$0 — US$33.175/mo, ÷ 0.7145 (RBA rate, 21 August 2026) = A$46.43/mo. The four rates are the vendors’ own published Sydney prices, applied to assumed volumes of 5 GB stored, 10 GB scanned per month, and 1 author plus 3 board readers. That is the funded day-1 system, and it supersedes the design on this page. The full lakehouse shown below (Kinesis streaming, Airbyte, Redshift Serverless, SageMaker, the five marts) is post-G2, trigger-gated growth design: a 10-layer stack is a category error for this entity at MVP scale, and each heavier layer may be re-proposed only against a measured trigger — storage > 100 GB sustained, scans > 1 TB/month, dashboard readers > 8, a Spark-only transform, or the first AU customer record. The Stage 1–3 teams below are likewise growth design: the day-1 roster prices no data-engineering hires — one fractional analyst, not “2–3 data engineers”.
Source: The MVP sizing statement, bill of materials and run cost; the growth layers and their re-proposal triggers; and the day-1 roster.
Interactive · End-to-End Architecture
From Source Systems to Marketing Activation
Every box below is a working part of the design, and every one of them opens: what it is in one plain sentence, what happens to the data there, the sample file or dashboard on the prototype page that stands behind it, how that layer is charged, and what can go wrong there. Use Follow one ticket to watch a single sample transaction travel the whole path.
Hover any node for a plain-language explainer. Click to pin it — pin two to compare them side by side.
Source Systems
Ingestion
Governed Lakehouse
S3 + Iceberg · Glue + dbt · Athena / Redshift Serverless
Data Marts
Activation
These four controls are not a stage — they run underneath every node above, and each one is a condition the stage has to satisfy before data moves on.
First-Party Platform DB: The ticketing app’s own database — the single place where a sale actually happens and the only system that knows a real buyer bought a real seat. Payments & Settlement: The payment provider’s record of what money moved — authorisations, refunds and the payouts that eventually land in the bank. Marketing Platforms: The advertising and messaging tools — what was sent, to whom, what it cost and which ticket sale the platform claims it caused. Public & Government Data: Official statistics — the census and survey releases that supply the denominators no company can produce for itself. Partner & Promoter Feeds: What the people who actually own the shows send in — venue capacities, performance dates, allocations and settlement terms. Batch ETL: The scheduled collection round — once a night (or once an hour) it fetches whole files from each source and brings them in together. Event Streaming: The live wire — transactions and app events arrive the second they happen instead of waiting for the nightly run. Landing Zone: A locked room where every incoming file is kept exactly as it arrived — encrypted, timestamped and untouched. Bronze — Raw: The permanent raw record: every field the source sent, kept in full, so the business can change its mind later and recompute from scratch. Silver — Validated: Where messy raw data becomes trustworthy data — same facts, made consistent, de-duplicated and legally usable. Gold — Business-Ready: The one version of each number the business argues from — revenue, refund rate, contribution per ticket — defined once and computed the same way every time. Finance & Unit Economics: The money view: what was sold, what was refunded, what tax is owed and what actually settled. Customer & Consent: Who the buyers are, what they have agreed to, and what the business is therefore allowed to send them. Events, Venues & Inventory: What is on sale, where, how much of it is left, and how much of the room actually filled. Marketing & Growth: What it costs to win a buyer, and whether the campaign actually caused the sale. Markets, Partners & Risk: The board’s view: which market to enter next, who the counterparty is, and what is currently going wrong. Campaign Activation: Turning a governed segment into an actual audience the marketing tools can send to. Dashboards & BI: The screens Leadership actually looks at — every tile carrying the file it came from and the label of what kind of data it is. APIs & ML Models: Machine-to-machine delivery — governed data feeds for partners, and models that price or predict. Consent & Privacy Enforcement: The control that decides, for every row, whether the business is legally allowed to use it for this purpose. Data-Quality Gates: Automated tests that stop a bad number reaching a dashboard, instead of someone spotting it afterwards. Lineage & Traceability: The paper trail: for any number on any screen, which file it came from and when that file was captured. Residency & Transfer Controls: The rules about where data is allowed to physically sit, and what has to be signed before it crosses a border.
Per the reconciliation above, this is the post-trigger target design, not the committed MVP build. Each node names the basis on which that layer is charged and points back to the one costed stack — the committed MVP bill of materials at the top of this page.
Consumption
Five Data Marts
Finance & Unit Economics
Reconciled financial truth for revenue recognition, settlement tracking and unit-economics analysis
Key metrics: GTV, Recognised Revenue, Take Rate, Refund Rate, Settlement Variance, Contribution per Ticket
Access: CFO, Finance team, External auditors (read-only)
Refresh: Daily reconciliation; real-time for settlement exceptions
Sensitive: Payment tokens, tax fields (restricted)
Customer & Consent
Understand buyer behaviour under strict consent and privacy controls
Key metrics: Active Buyers, Repeat Rate, Consented Reachable Audience, Deletion SLA Compliance
Access: Customer team, Marketing (consented attributes only), Privacy officer
Refresh: Event-driven for consent; daily for behavioural aggregates
Sensitive: Name, email, phone, location (tokenised)
Events, Venues & Inventory
Track event supply, venue utilisation and operational performance
Key metrics: Sell-Through Rate, Venue Capacity Utilisation, Lead Time to Purchase, Scan Rate
Access: Operations team, Event managers, Promoter portal (filtered)
Refresh: Real-time inventory; daily performance metrics
Sensitive: Operational access tokens (restricted)
Marketing & Growth
Measure acquisition efficiency and campaign performance
Key metrics: CAC, ROAS, Conversion Rate, Incrementality
Access: Marketing team, Growth team, CMO
Refresh: Daily campaign metrics; weekly attribution models
Sensitive: Ad identifiers and audience membership (no raw PII)
Markets, Partners & Risk
Support market-entry decisions, partner due diligence and risk monitoring
Key metrics: Market Opportunity Score, Partner Pipeline Value, Control Status, Incident Count
Access: CEO, Strategy team, Legal/compliance, Board (summary)
Refresh: Monthly market indicators; event-driven incidents
Sensitive: B2B contacts and due-diligence records (restricted)
Growth Path
Scalability Roadmap — Posture, Infrastructure and Team Shape
Infrastructure and team shape for each stage stand as design intent. The A$46.43/mo in the reconciliation above prices the day-1 MVP stack and nothing beyond it: it buys none of the Stage 1–3 infrastructure listed here, and it is not a floor or a per-stage cost. Each stage is entered on its stated trigger.
Stage 1
Serverless, scheduled processing, one BI environment
Infrastructure: Single AWS account; S3, Glue, Redshift Serverless, QuickSight
Team: 2–3 data engineers (can be contractors); fractional architect
Trigger to advance: Prove reconciled metrics and demonstrate user value
Stage 2
Autoscaling ingestion, separated dev/test/prod, stronger observability
Infrastructure: Multi-environment AWS; dedicated networking; monitoring and alerting
Team: 4–6 data/platform engineers; dedicated BI developer; data analyst
Trigger to advance: Add availability guarantees, private networking and formal on-call support
Stage 3
Multi-account, queue buffering, load testing, regional recovery design
Infrastructure: Multi-account AWS; cross-region recovery; CDN; advanced security
Team: 8–12+ platform, data and security engineers; dedicated governance
Trigger to advance: Dedicated platform/security team and negotiated cloud commitments
Procurement
Technology Options by Layer
Each row below gives the basis on which that layer is actually charged, rather than a rate: each names three alternative products for one layer, and the workload that would select between them is settled at gate G2. The five components that are committed at MVP scale — S3, Glue Data Catalog, Athena, dbt Core and QuickSight — do have both, and their published rates, assumed volumes and arithmetic are set out in full in the reconciliation above, which is where the A$46.43/mo run cost comes from. The recommended, alternative and premium-alternative products are this proposal’s own technology choices.
| Layer | Recommended | Alternative | Premium Alternative | Pricing Basis |
|---|---|---|---|---|
| Source Contracts | Open JSON schemas | AWS Glue Schema Registry | Confluent Schema Registry | Internal labour / managed usage |
| Batch Ingestion | AWS-native jobs | Airbyte (open-source) | Fivetran (MAR-based) | Compute / infra / MAR |
| Streaming | Amazon Kinesis | Amazon MSK (Kafka) | Confluent Cloud | Throughput / storage |
| Landing/Raw Storage | Amazon S3 | Azure Data Lake Storage Gen2 | Google Cloud Storage | Storage + requests |
| Processing | AWS Glue/Spark | Databricks | Snowflake Processing | Job compute / DBU / credits |
| Warehouse/Query | Redshift Serverless + Athena | Snowflake | BigQuery | RPU / credits / bytes |
| Orchestration | AWS Step Functions + dbt Core | Prefect | dbt Cloud | Usage / workspace / seats |
| Governance | Lake Formation + Glue Catalogue | Databricks Unity Catalog | Alation/Collibra | Cloud use / platform / quote |
| BI | Amazon QuickSight | Power BI | Looker | User / session / capacity |
| AI/API | SageMaker + API Gateway | Databricks ML | BigQuery ML + API | Compute / inference |
Governance
Technology Approval Gates — No Gate, No Spend
These are engineering approval gates, numbered TG-0…TG-4 so they cannot be confused with the financial decision schedule, which runs G0–G2 only: G0 due diligence & terms, G1 discovery — the primary demand study, not data feasibility — and G2 MVP build. There is no G3 on that schedule. Mapping: TG-0 sits inside G0’s scope; TG-1 is a technical check consumed within G1’s discovery scope, not G1 itself; TG-2 is G2’s exit test; TG-3 and TG-4 lie beyond the financial schedule and would each require a new decision paper.
TG-0: Entity and Rights
Certified entity, ownership, IP, data and contract rights
If failed: Stop all work
TG-1: Data Feasibility
Representative extracts, data dictionaries, control totals and consent samples
If failed: Do not build
TG-2: MVP Value
Three certified dashboards, reconciliation above agreed threshold, named users
If failed: Fix or stop
TG-3: Production
Security review, recovery test, retention policies, vendor DPAs, runbook
If failed: No production PII
TG-4: Scale
Demonstrated unit economics, reliability and procurement case
If failed: Avoid committed licences
Compliance by Design
Retention & Cross-Border Transfers
Retention Policy
| Data Type | Retention | Rationale |
|---|---|---|
| Transaction records | 7 years | Tax and audit obligations |
| Customer consent records | Relationship + 3 years | Regulatory evidence |
| Marketing campaign data | 3 years | Performance analysis and attribution |
| Raw Bronze data | 3 years minimum | Reprocessing capability |
| Operational logs | 2 years | Security and incident investigation |
| Deletion request records | Permanent (metadata only) | Proof of compliance |
Transfer Mechanisms
| Route | Legal Mechanism |
|---|---|
| Australia to EU | Standard Contractual Clauses (SCCs) or adequacy decision |
| Australia to UK | UK International Data Transfer Agreement or UK SCCs |
| Australia to USA | EU-US Data Privacy Framework (participating vendors); SCCs otherwise |
| Australia to Canada | APP 8 (Privacy Act 1988 (Cth)) reasonable-steps assessment before disclosure |
| Any transfer | Transfer Impact Assessment required before production data flows |
On the Australia-to-Canada route, the EU adequacy decision for Canada’s PIPEDA is not the legal basis: it is a GDPR instrument governing EU-to-Canada transfers only, and it has no legal force over an Australian-origin disclosure. That disclosure is governed instead by Australian Privacy Principle 8 of the Privacy Act 1988 (Cth).
Committed MVP Run Cost
A$46.43/mo
Calculated — the committed MVP bill of materials, monthly-cancellable. The working: S3 storage 5 GB at US$0.025/GB-month = US$0.125, Athena 10 GB scanned at US$5.00/TB = US$0.05, and QuickSight 1 author at US$24/mo plus 3 readers at US$3/mo = US$33.00, with Glue Data Catalog on its published free tier and dbt Core, open source, at A$0 — US$33.175/mo, ÷ 0.7145 (RBA rate, 21 August 2026) = A$46.43/mo. The four rates are the vendors’ own published Sydney prices, applied to assumed volumes of 5 GB stored, 10 GB scanned per month, and 1 author plus 3 board readers. Exceed a volume and the line is recalculated.
