Human motion prediction system and method in dynamic scene
Through hierarchical interactive feature representation and Transformer prediction backbone network, the accuracy and efficiency problems of human motion prediction in dynamic scenarios are solved, and real-time and accurate motion prediction in dynamic environments are achieved.
Patent Information
- Application Number
- CN202510446941.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art human motion prediction methods in dynamic scenarios fail to fully model the dynamic interaction between individuals and the environment, resulting in insufficient prediction accuracy and applicability, and high calculation costs, making it difficult to meet the needs of real-time applications.
The hierarchical interaction feature representation and Transformer prediction backbone network are adopted, and through the coarse to fine interactive reasoning module and the DCT frequency regulation mechanism, people-human interaction and people-scene interaction information are integrated to improve prediction accuracy and efficiency.
It significantly improves the accuracy and computing efficiency of human motion prediction, and can perform real-time and accurate motion prediction in dynamic scenarios.
Smart Images

Figure CN120412082A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a system and method for human motion prediction in dynamic scenes. Background Art
[0002] In the field of human behavior prediction, existing methods usually only focus on the motion trajectories of individuals, or simply introduce environmental features into the prediction framework. However, the real-world dynamic environment is full of complex interpersonal interactions and scene factors. The social interaction between individuals and the scene interaction between individuals and the environment play crucial roles in motion prediction. Existing methods are often limited to single factors, either only considering the relationship between people (social interaction) or only focusing on the impact of the scene on individual motion. There are few motion predictions that can comprehensively model in dynamic scenes. Although there are very few recent methods that attempt to generate human actions in dynamic scenes, their essence belongs to long-term motion generation, which focuses more on the diversity and aesthetics of the generation effect rather than the accuracy of motion prediction (motion prediction). More importantly, it has many limitations: First, it relies on the instance segmentation information of the scene, resulting in limited applicability; Second, it is based on diffusion-based models, with high computational costs and slow inference speeds (the prediction time for a single sample exceeds 1 second), making it difficult to meet the requirements of real-time applications; Finally, it only encodes environmental information in isolation and fails to explicitly model the dynamic interaction between individuals and the environment, thus affecting the accuracy and generalization ability of the prediction. Summary of the Invention
[0003] The objective of the present invention is to provide a human motion prediction solution that can fully integrate social relationships and environmental factors and is applicable to dynamic scenes, so as to improve prediction accuracy, computational efficiency, and applicability.
[0004] To achieve the above objective, the technical solution of the present invention discloses a human motion prediction system in a dynamic scene, characterized in that the overall architecture includes:
[0005] Hierarchical Interaction Feature Representation:
[0006] Model the human-human interaction and human-scene interaction respectively in a hierarchical manner to comprehensively model the motion factors in the dynamic environment;
[0007] Based on the prediction backbone network of Transformer, a coarse-to-fine interaction reasoning module is proposed, which uses self-attention mechanism and cross-attention mechanism to integrate interaction information at different levels and improve prediction accuracy, where:
[0008] The coarse-to-fine strategy adopted by the coarse-to-fine interactive reasoning module is divided into:
[0009] Spatial dimension: First inject global information, and then gradually introduce local details to achieve coarse-to-fine feature fusion;
[0010] Frequency domain dimension: Adopt the DCT frequency regulation mechanism to dynamically adjust the weights of different frequency components during the prediction process to enhance the prediction effect.
[0011] Preferably, in the interactive modeling of human-human interaction and human-scene interaction, high-level semantic features and low-level geometric features are fused, taking into account both global semantics and local geometric relationships, and combining explicit interactive distance calculation and implicit feature learning to improve the modeling ability of interactive relationships.
[0012] Another technical solution of the present invention is to provide a method for predicting human motion in a dynamic scene, which is characterized by including the following steps:
[0013] Step 1: Extract features from the input data, where the input data includes:
[0014] The historical motion sequence of the target individual, providing individual motion pattern information;
[0015] The historical motion information of the interactive individual, used for modeling human-human interaction;
[0016] 3D scene point cloud data, used for modeling human-scene interaction;
[0017] Step 2: Hierarchical interactive feature modeling, specifically including:
[0018] Human-human interaction modeling: Calculate the interactive distance between individuals, and construct interactive features by combining self-encoding and human relationship encoding;
[0019] Human-environment interaction modeling: Use PointNet++ for multi-level scene information extraction to ensure efficient encoding of environmental information;
[0020] Step 3: Backbone network reasoning
[0021] Parse the human motion pattern through the self-attention mechanism to improve the time series modeling ability, combine the cross-attention mechanism to fully integrate human-human interaction and human-environment interaction information, and adopt a coarse-to-fine interactive feature injection strategy, including:
[0022] First use high-level features for global semantic understanding, and then gradually introduce detailed features to improve local accuracy;
[0023] Adopt the DCT frequency suppression mechanism: Dynamically adjust the weights of different frequency components during the prediction process. Among them, the low-frequency components have a greater impact in the initial stage, mainly focusing on the global motion trend, and the high-frequency components gradually recover in the later stage to enhance the detail prediction ability.
[0024] Step 4, Prediction Output
[0025] Through the GCN decoder and the inverse DCT, decode and restore the frequency-domain representation to the time domain, and finally output the motion trajectory of the target individual at future time steps.
[0026] Preferably, in step 1, when performing feature extraction, use the DCT + GCN encoder to extract motion features and construct a hierarchical interaction feature representation by combining interaction information.
[0027] The present invention proposes an innovative human motion prediction method, which significantly improves the accuracy and computational efficiency of human motion prediction through hierarchical interaction modeling and coarse-to-fine interaction reasoning. This technology can be widely applied to fields such as autonomous driving, robot interaction, intelligent monitoring, VR / games, etc., and is of great significance for enhancing the decision-making and planning capabilities of intelligent systems. Brief Description of the Drawings
[0028] Figure 1 It is the schematic diagram of the present invention;
[0029] Figure 2 Illustrates the specific process of the present invention. Detailed Embodiments
[0030] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0031] The present invention significantly improves the accuracy of human motion prediction by introducing multi-level interaction modeling and an efficient backbone network.
[0032] One aspect of the embodiments of the present invention is to disclose a human motion prediction system in a dynamic scene, and the overall architecture includes the following two parts:
[0033] First) Hierarchical Interaction Feature Representation
[0034] Model human-human interaction and human-scene interaction separately in a hierarchical manner, and comprehensively model the motion factors in the dynamic environment. Integrate high-level semantic features and low-level geometric features in the interaction modeling, taking into account both global semantics and local geometric relationships. Combine explicit interaction distance calculation (Explicit Interaction Distance Modeling) with implicit feature learning (Implicit Feature Learning) to effectively improve the modeling ability of interaction relationships.
[0035] Second), based on the Prediction Backbone with Transformer network, propose the Coarse-to-Fine Interaction Reasoning Module: adopt the self-attention mechanism and cross-attention mechanism to integrate interaction information at different levels and improve the prediction accuracy.
[0036] The coarse-to-fine strategy proposed in the embodiments of the present invention is divided into two aspects: (1) Spatial dimension: first inject global information, and then gradually introduce local details to achieve coarse-to-fine feature fusion; (2) Frequency domain dimension: adopt the DCT frequency regulation mechanism to dynamically adjust the weights of different frequency components during the prediction process to enhance the prediction effect.
[0037] Another aspect of the embodiments of the present invention is to disclose a human motion prediction method in a dynamic scene, which specifically includes the following steps:
[0038] Step 1, Data input and feature extraction
[0039] The input data of the present invention includes: 1) The historical motion sequence (motion) of the target individual, which provides individual motion pattern information; 2) The historical motion information of the interacting individuals (Multi-Person Interactions), which is used to model human-human interaction; 3) 3D scene point cloud data (3D Point Cloud), which is used to model human-scene interaction.
[0040] Preprocessing and encoding: Extract motion features through the DCT+GCN encoder, and construct a hierarchical interaction feature representation in combination with interaction information.
[0041] Step 2, Hierarchical interaction feature modeling, specifically including:
[0042] 1) Human-Human Interaction (HHI) modeling: Calculate the interaction distance between individuals, and construct interaction features by combining self-attention and relation encoding between people.
[0043] 2) Human-Scene Interaction (HSI): Use PointNet++ to extract multi-level scene information to ensure efficient encoding of environmental information.
[0044] Step 3, backbone network inference
[0045] The present invention proposes Transformer-based Interaction Reasoning: Analyze human motion patterns through the self-attention mechanism to improve the time series modeling ability. Combine the cross-attention mechanism to fully integrate human-human interaction and human-environment interaction information to enhance the prediction robustness. Coarse-to-Fine Feature Injection: 1) First, use high-level features for global semantic understanding, and then gradually introduce detailed features to improve local accuracy. 2) DCT Frequency Rescaling: The low-frequency components have a greater impact in the initial stage, mainly focusing on the global motion trend. The high-frequency components gradually recover in the later stage to enhance the detailed prediction ability.
[0046] Step 4, prediction output
[0047] Through the GCN decoder and Inverse DCT, decoding is achieved and the frequency domain representation is restored to the time domain. Finally, the motion trajectory of the target individual at future time steps is output.
[0048] Example 1
[0049] When a person attempts to pick up an object in a dynamic environment, the system disclosed by the present invention can real-time predict its most likely subsequent actions, such as reaching, grasping, lifting, etc. The system first inputs the historical motion sequence of the target individual and combines the information of the static object positions and other dynamic elements in the scene, such as whether other people are approaching the object. Through human-human interaction modeling, the system disclosed by the present invention can identify the behaviors of other people and infer their potential impacts, such as whether there will be a competitive taking. At the same time, through human-environment interaction modeling, the system calculates the distance between the target person and the object and predicts a reasonable grasping angle and gesture based on the affordance features. Then, based on the Transformer interaction reasoning mechanism of the present invention, the system uses self-attention and cross-attention to integrate various interaction information, combines the coarse-to-fine feature injection strategy, analyzes the overall trend first, and then optimizes the local details, and finally generates the future motion trajectory of the target person, such as the hand moving towards the object, etc. This method can reasonably adjust the action generation according to the behavior patterns of static objects and other dynamic individuals in the environment to ensure compliance with physical rules and common sense.
[0050] Example 2
[0051] When modeling VR / game scenarios, the system disclosed by the present invention can automatically predict the movement behavior of characters in complex dynamic environments, making the performance of virtual characters more natural.
[0052] Example 3
[0053] In the context of autonomous driving, predicting the movement trajectories of pedestrians is crucial for safe driving. The present invention can be used in autonomous driving systems to accurately predict whether a pedestrian will cross the road and adjust driving decisions accordingly. For example, when an autonomous vehicle approaches a crosswalk, sensors collect the historical movement trajectories of pedestrians and analyze whether the pedestrians have the intention to cross in combination with the surrounding traffic environment, such as the status of traffic lights and vehicle flow. The system disclosed by the present invention first analyzes the relationships between multiple pedestrians through human-human interaction modeling, such as whether there is a collective crossing trend, and at the same time uses human-environment interaction modeling to calculate the relationship between pedestrians and the road by combining information on sidewalks, zebra crossings, and road boundaries. Then, based on the Transformer interaction reasoning mechanism, it first processes global information, such as the overall pedestrian flow trend, and then refines the prediction of individual behaviors. This can reduce sudden traffic accidents and improve the safety and robustness of autonomous driving.
[0054] Example 4
[0055] The present invention can be used in intelligent monitoring systems to predict abnormal behaviors and provide safety warnings. If the system predicts a potentially dangerous human action, it will trigger a safety alarm in advance, notify security personnel to intervene, and control the flow direction of people in case of emergency to guide safe evacuation. Compared with traditional intelligent monitoring, the present invention can actively predict risks instead of just detecting existing dangers, improving the intelligence level and response speed of the monitoring system.
Claims
1. A human motion prediction system in a dynamic scenario, characterized in that, The overall architecture includes: Hierarchical interaction feature representation: Model human-human interaction and human-scene interaction separately in a hierarchical manner, and comprehensively model the motion factors in the dynamic environment; Based on the Transformer prediction backbone network, a coarse-to-fine interaction reasoning module is proposed. The self-attention mechanism and cross-attention mechanism are used to integrate interaction information at different levels and improve the prediction accuracy. Among them, the coarse-to-fine strategy adopted by the coarse-to-fine interaction reasoning module is divided into: Spatial dimension: First inject global information, and then gradually introduce local details to achieve coarse-to-fine feature fusion; Frequency domain dimension: Adopt the DCT frequency regulation mechanism to dynamically adjust the weights of different frequency components during the prediction process to enhance the prediction effect.
2. The human motion prediction system in a dynamic scenario according to claim 1, characterized in that, In the interaction modeling of human-human interaction and human-scene interaction, high-level semantic features and low-level geometric features are fused, taking into account both global semantics and local geometric relationships, and combining explicit interaction distance calculation and implicit feature learning to improve the modeling ability of interaction relationships.
3. A method for predicting human motion in a dynamic scene, characterized in that, It includes the following steps: Step 1: Extract features from the input data. Among them, the input data includes: The historical motion sequence of the target individual, providing individual motion pattern information; The historical motion information of the interacting individuals, used to model human-human interaction; 3D scene point cloud data, used to model human-scene interaction; Step 2: Hierarchical interaction feature modeling, specifically including: Human-human interaction modeling: Calculate the interaction distance between individuals, and construct interaction features by combining self-encoding and human relationship encoding; Human-environment interaction modeling: Use PointNet++ for multi-level scene information extraction to ensure efficient encoding of environmental information; Step 3: Backbone network reasoning Analyze the human motion pattern through the self-attention mechanism to improve the time series modeling ability. Combine the cross-attention mechanism to fully integrate human-human interaction and human-environment interaction information, and adopt a coarse-to-fine interaction feature injection strategy, including: First use high-level features for global semantic understanding, and then gradually introduce detailed features to improve local accuracy; Adopt the DCT frequency suppression mechanism: Dynamically adjust the weights of different frequency components during the prediction process. Among them, the low-frequency components have a greater impact in the initial stage, mainly focusing on the global motion trend, and the high-frequency components gradually recover in the later stage to enhance the detailed prediction ability; Step 4: Prediction output Through the GCN decoder and the inverse DCT, decoding is realized and the frequency domain representation is restored to the time domain, and finally the motion trajectory of the target individual at future time steps is output.
4. The human motion prediction method in a dynamic scenario according to claim 3, wherein In Step 1, when extracting features, motion feature extraction is performed through the DCT+GCN encoder, and hierarchical interaction feature representation is constructed by combining interaction information.