Learning Human-centric Motion Representation for Action Analysis
Michigan State University
Abstract
We introduce H-MoRe, an innovative pipeline designed to learn precise, human-centered motion representations.
Our approach dynamically retains essential human motion features while filtering out background noise. Unlike traditional methods that rely on fully supervised learning with synthetic data, H-MoRe employs a self-supervised learning paradigm directly from real-world scenarios, incorporating both human pose and body shape information.
Drawing inspiration from kinematics, H-MoRe encodes absolute and relative movements of body points into a matrix representation, termed world-local flows, to capture subtle motion details. This method provides a detailed understanding of human motion, making it highly adaptable to various action-based applications.
Dense Human-centric Flow
H-MoRe represents motion as human-centric world-local flows, locking onto the moving body where generic optical flow is distracted by the background.
Approach
Qualitative Results
Against State-of-the-Art Flow
Flow visualizations generated by our H-MoRe and seven optical flow estimation algorithms.
Frame-by-Frame
Drag the slider to inspect the H-MoRe inference result for each frame.

BibTeX
@inproceedings{huang2025hmore,
title = {Learning Human-centric Motion Representation for Action Analysis},
author = {Huang, Zhanbo and Liu, Xiaoming and Kong, Yu},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2025},
}