Overview
Embodied intelligence is shaped not only by the scene in front of an agent, but also by accumulated experience. Familiar places carry histories of use: objects change state, actions become procedures, and repeated episodes reveal regularities that no single frame contains. Egocentric video records this experience from the actor's point of view.
While egocentric imitation and cross-embodiment methods use first-person human video to learn how behavior can be executed, MEMORA studies a complementary question: what should experience become for a future planner? We formulate Embodied Action Memory as the capability to form, maintain, and use embodied experience as a persistent memory state.
MEMORA organizes egocentric experience as editable and consolidated embodied memory, rather than retaining it only as episodes to retrieve or demonstrations to imitate. The resulting semantic–procedural state preserves what happened, what changed, and what recurs so that later goals can be grounded in participant-specific experience.
MEMORA-Bench evaluates this lifecycle on 45 hours of egocentric video across 18 participants. Across four open-weight language models, full MEMORA performs best overall among the evaluated memory interfaces. It gains up to 20.5 points on the experience-dependent, memory-grounded EAM-QA subset and up to 16.6% relative Robot-Grounded Plan score on Generalize. A qualitative two-task robot deployment illustrates the path from remembered human experience to physical execution.

