Everyone gives the same answer to this one. "Embed if it's one-to-few, reference if it's one-to-many." Then the interviewer asks about a specific case, comments on a blog post, and the one-liner stops being enough.
I designed the schema for my own CMS on MongoDB, so these are the calls I actually made and why.
The question behind the question
It's usually concrete:
"You've got blog posts and comments. Would you embed the comments inside the post document, or put them in their own collection and reference them?"
They're not testing whether you know the words "embed" and "reference." They're testing whether you can predict how the data behaves over time, and whether you know the constraint that makes the wrong choice actually break: a single MongoDB document maxes out at 16MB. Embed something that grows forever and you will hit that wall.
The three questions that decide it
1. Does it grow without bound?
If the child collection has no natural ceiling, comments, activity log entries, messages, it goes in its own collection. A post with 4,000 comments embedded is a 4,000-element array you load in full every time someone opens the post, and one day it stops fitting.
If the count is small and capped, embed it.
2. Does anything need to query it on its own?
"Show me the 20 most recent comments across the whole site" is almost impossible if comments are buried in post documents. You'd be scanning every post and unwinding arrays. If the child has its own list views, its own pagination, its own filters, it wants to be a collection.
3. Is it shared, or owned by the parent?
An author isn't part of a blog post -- the same user writes 50 posts -- embed the user object in all 50 and now a name change is a 50-document update, and they can drift out of sync. Shared data gets referenced.

What I embedded, and what I referenced
In the actual CMS:
- Post images: embedded. Each post has an array of image subdocuments (
url,order, a generated id). They're capped at 10, they're always rendered with the post, and nothing else ever asks "give me all images." Textbook embed. - Tags: embedded. A plain string array on the document. Small, bounded, and a multikey index on it handles "posts tagged
javascript" fine. No reason for a collection. - Author: referenced. The post stores an
authorId. Users are their own collection with their own auth data, and one user has many posts. Classic reference. - Topic: neither, on purpose. The post stores the topic as a plain string that happens to match a name in a managed
BlogTopicscollection. Not anObjectIdreference. That means renaming or deleting a topic doesn't cascade through every post or leave dangling ids, the coupling is deliberately loose. Worth mentioning in an interview because it shows you know reference isn't the only non-embed option. - Comments: referenced. Their own collection, each with a
postId. Unbounded, paginated on their own, and I need "recent comments" views. This is the exact case from the question, and I built it the way the "right" answer says to.
The Mongoose part they follow up on
If the interview is Mongoose-specific, expect: "how does that look in code?"
Embedded is a subdocument schema nested in the parent schema, it's just there when you load the parent. Referenced is type: ObjectId, ref: 'User', and you pull the related doc with .populate('author').
The thing to say out loud: populate is not a join. Mongo isn't relational. Under the hood it's a second query (or a $lookup aggregation stage), run per level you populate. Reference the wrong things and you turn one read into five. That trade, storage and consistency vs read-time work, is the whole embed-vs-reference decision in miniature.
And the anti-pattern to name: the unbounded embedded array. It's the single most common MongoDB schema mistake, an events: [] or comments: [] that looked fine with test data and falls over in production.
Quick recap
- The hard constraint is the 16MB document limit, embed something that grows forever and it eventually breaks
- Embed when the child is bounded, loaded with the parent, and never queried alone
- Reference when it grows without bound, has its own queries, or is shared across parents
populateis an extra query per level, not a free join, so referencing has a real read cost- The classic mistake is the embedded array with no ceiling; if you can't name a maximum, don't embed it
The reason "embed comments or reference them" is such a common question is that the honest answer is "reference them, and here's the one number that tells me so." If you can say 16MB and unbounded growth in the same breath, you're past it.




Comments
No comments yet — be the first to share your thoughts.