·VegaLoop Team

DynamoDB Data Modeling for Multi-Domain Health Data

How a table-per-bounded-context design handles nutrition, activity, and goals without turning into a mess.

architecturenutritiontraininggoals

Health data is messy. You’ve got meals logged at odd hours, runs tracked by GPS, sleep scores arriving from a watch at 3am, and weekly goals that tie all of it together. A relational database wants you to normalize that into neat little tables with foreign keys. DynamoDB wants you to think differently.

The challenge with DynamoDB data modeling for a health platform isn’t the technology itself. It’s the mindset shift. You design around how your application reads data, not how the data relates logically. That inversion trips up experienced engineers and makes the whole thing feel backwards until it clicks.

Access patterns first, schema second

In a relational world, you model your entities and then write queries. With DynamoDB, you flip that. You list every question your application will ask, then design your keys to answer those questions efficiently.

For a wellness platform, the core access patterns look something like this: show me everything a user did today, show me their nutrition for the past week, show me their current training block, show me progress toward active goals. Each of those is a query, and each query should be answerable with a single request.

Most DynamoDB writing at this point says “single-table design” and means one giant table for the entire application. We don’t do that, and it’s worth being precise about why, because the distinction shapes everything downstream.

The mental exercise of listing access patterns up front is genuinely useful even outside DynamoDB. It forces product conversations early. “Do we need to show a user their weekly protein average?” becomes a design decision, not an afterthought during a sprint review. You end up with a tighter contract between the frontend and the data layer because both sides agreed on what questions matter before anyone wrote code.

One table per bounded context, not one table for everything

Our backend is split into bounded contexts. Each context owns its own DynamoDB table. Nutrition items never live in the activity table. Goal progress never lives in the nutrition table.

Within each table, we apply the classic single-table techniques. The nutrition table holds meals, daily summaries, scheduled meals, and the recipe catalog side by side, distinguished by key prefixes. But the boundary of a table is the boundary of a context, not the boundary of the platform.

The reasons are mostly organizational, and that’s not a weakness. A context’s functions get IAM permissions to their own table and nothing else, so a bug in meal logging physically cannot corrupt workout history. Each context evolves its schema, indexes, and capacity settings independently. Monitoring and alarms map cleanly to domains. And when you’re debugging at 2am, “which table is this in” has an obvious answer.

What you give up is the ability to fetch a user’s meals and workouts in one query. We’ll come back to how we handle that, because the answer isn’t “two queries at read time.”

Partition keys as a shared contract

Here’s where the tables stop being islands. Every context table uses the same partition key convention: USER#<user-id>. It’s a deliberate cross-service contract, not a coincidence, and it buys two things.

First, within any context, all of a user’s data lives in one partition, so every query starts with “for this user, show me…” and lands on a single partition key. Nobody generates enough health data to stress DynamoDB’s partition limits, and writes distribute across users naturally, which sidesteps the classic time-series hot-partition problem without any date-sharding tricks.

Second, deletion becomes mechanical. When a user deletes their account, an event fans out to every context, and each one drains the USER#<user-id> partition in its own table. There is no scavenger hunt across secondary indexes and orphaned rows. The partition key convention makes “remove everything about this person” a well-defined operation in every domain, which matters enormously for deletion requests.

Sort keys do the heavy lifting

Sort keys are where DynamoDB’s power hides. A well-designed sort key lets you answer multiple access patterns from the same partition without scanning.

In the nutrition table, a logged meal lands with a sort key like MEAL#2026-09-02#<entry-id> and the day’s rollup lives at DAILY_SUMMARY#2026-09-02. In the activity table, a run is WORKOUT#2026-09-02#<workout-id>. In the goals table, a goal is GOAL#<goal-id> and each progress entry is PROGRESS#<goal-id>#<timestamp>. With sort key begins-with queries, you can pull all meals for a date range, all summaries for a week, or all progress entries for a specific goal. One table per context, many views into its data.

The delimiter choice matters more than you’d think. We use # as a separator because it sorts predictably in UTF-8 and doesn’t appear in our entity identifiers. Some teams use | or :: but whatever you pick, be consistent. A mixed delimiter strategy is a debugging nightmare six months later.

Sort key design also determines how you handle “get me the latest” queries. Because the date is embedded in WORKOUT#2026-09-02#..., a reverse-order query with a limit of 1 gives you the most recent workout instantly. No scanning, no filtering. That pattern powers the “last workout” widget on the home screen.

Aggregates are computed on write

For anything that appears on a dashboard, we pre-compute. When someone opens the app to check how their nutrition is affecting recovery, they shouldn’t wait for a query that fetches a week of meal records and sums them in application code.

The daily nutrition summary is the clearest example. Every meal write synchronously recomputes the day’s totals and saves the DAILY_SUMMARY#<date> item in the same partition. The summary is always consistent with the meals because the same code path maintains both. Reads of “how am I doing today” are a single item fetch.

The write amplification is manageable because the aggregates are small and the writes are user-paced. Nobody logs a thousand meals a second. Paying a little on every write to make every dashboard read a point lookup is the right trade for an app people check dozens of times a day.

Events between contexts, transactions within them

Health data doesn’t live in silos. Finishing a workout should advance goals, inform nutrition targets, and show up on the dashboard. With one table per context, you can’t wrap all of that in a single transaction, and we think that’s a feature.

Within a context, we use TransactWriteItems where atomicity genuinely matters. When a workout syncs in from a wearable, the workout record and an idempotency pointer keyed by the external ID commit together, so a retried webhook can’t create a duplicate. In the identity context, a profile write and its email-uniqueness reservation commit as one transaction. In every case the transacted items share the same user partition in the same table.

Between contexts, we publish events. Each context owns a topic; interested contexts subscribe through queues and consume at their own pace, with dead-letter queues catching failures. When a workout completes, the goals context’s consumer picks up the event and writes a projection row (ACT_WORKOUT#<date>#<workout-id>) into the goals table. Goal progress is then computed entirely from rows the goals context owns. It never reaches into the activity table, at write time or read time.

Some of these flows run both directions. Nutrition publishes daily summary updates that the activity context materializes as caloric intake context for training. Activity publishes workout completions that nutrition materializes as training load for meal recommendations. Each context keeps a local, query-optimized copy of exactly the foreign data it needs.

The cost is eventual consistency between contexts. A logged meal advances goal progress a moment later, not in the same millisecond. For health data, that’s the right call. What you must not have is a corrupted invariant within a domain (a duplicated workout, a calorie total that doesn’t match its meals), and the design makes those impossible where they matter while letting cross-domain effects settle asynchronously.

The dashboard is a read model

Remember the query we gave up: “show me everything a user did today” across domains. The answer is a dedicated dashboard context whose table is nothing but projections.

The dashboard subscribes to events from activity, nutrition, identity, goals, and integrations. Its consumer materializes DAILY#<date> summaries, recent-item rows like RECENT_WORKOUT#<date>#<id> and RECENT_MEAL#<date>#<id>, and ACTIVE_GOAL#<goal-id> rows. Opening the app queries one partition in one table and gets the whole picture, already shaped for the screen.

This is the pattern that replaces the “one giant table so cross-domain reads are one query” argument. You don’t need every domain in one table to get single-query dashboards. You need a read model that’s fed by events and owned by the thing doing the reading.

Global secondary indexes for cross-cutting queries

Not every access pattern starts with a user ID, and that’s what GSIs are for. Ours are boring on purpose, and each earns its keep with a real, frequent query.

The identity table has an email-lookup index so login and uniqueness checks resolve an email to a user without a scan. The goals table indexes goals by status, so “show me this user’s active goals” is a direct query rather than a fetch-and-filter. The insights table maintains an active-day index: users who logged anything on a given date get a marker row, and the nightly job that computes readiness features queries that index instead of scanning the table for who was active. The nutrition table indexes the curated recipe and meal plan catalogs so browsing them never scans past user data.

GSIs cost storage and write capacity, since every write to indexed attributes replicates. A discipline we follow: every GSI proposal requires a written justification that includes the access pattern, expected query volume, and projected cost impact. This prevents GSI sprawl, which is the DynamoDB equivalent of index bloat in PostgreSQL.

Sparse indexes come free

DynamoDB GSIs only index items that carry the indexed attribute. Leave the attribute off and the item simply doesn’t exist in the index. This gives you sparse indexes for free, and they’re one of the most underrated tools in the box.

The meal plan catalog index works this way: only curated meal plan rows write the index key, so meals, summaries, and everything else in the nutrition table are absent from it. The index contains exactly the catalog and nothing more. The active-day index is the same idea: only the small marker rows carry the key, so the index stays tiny relative to the table, and the nightly job’s query touches only what it needs.

TTL is a janitor, not a data policy

Every one of our tables enables TTL, but not for user health data. Your meals and workouts live until you delete them; the account-deletion path handles that explicitly via the partition drain. TTL cleans up the operational debris around the event-driven design.

Event consumers write dedup marker items so a redelivered message is processed once; those markers expire after their useful window. OAuth state rows for connecting a wearable expire in minutes. Weather and air-quality cache entries expire in about an hour. None of this data has any value past its moment, and TTL deletes it at no cost without a cleanup job to build, monitor, or fix.

The distinction matters because TTL deletions are lazy — an item can persist up to 48 hours past its expiry. For a dedup marker, who cares. For a user’s data-deletion request, that latitude is unacceptable, which is exactly why deletion is an explicit drain and not a TTL.

The tradeoff you accept

A relational schema is immediately readable. Our design is not: nine context tables, each with its own key map, stitched together by topics and queues, with projection rows that are deliberate copies of another context’s data. New engineers look at any one table and see chaos until someone walks them through the access pattern map, and they look at the event wiring and ask why a workout row exists in the goals table.

You accept that for predictable performance at any scale, hard isolation between domains, zero operational overhead for connection pooling or query optimization, and per-request pricing that aligns with actual usage. For a health platform where usage is spiky (everyone logs breakfast around the same time, nobody logs much at 3am), that elasticity matters.

The onboarding cost is real, though. Each context maintains a document mapping every access pattern to its key structure, expected volumes, and the feature it supports. New team members read it before touching the data layer. It takes about an hour per context to internalize, and after that the design feels natural rather than chaotic.

What this means for your data

Every time you log a meal, finish a workout, or check your weekly progress, the data needs to land somewhere fast and come back faster. DynamoDB’s model forces you to think about that retrieval path before you write a single line of code, and the bounded-context split forces you to decide who owns each fact. Both constraints produce better outcomes for the people using the platform.

The data model is invisible to you as a user. But it’s the reason your dashboard loads quickly, your goal progress can’t silently diverge from your workouts, and your historical data doesn’t slow down as months of tracking accumulate. Good infrastructure disappears. You just see your progress.

Note: This article is for general information only and isn't medical advice. Everyone responds to training and nutrition differently. Talk to a doctor or qualified professional before making changes to your training or nutrition.