A Human Motion Prediction Method, Device and Medium for Human-Computer Interaction Scenarios
By aligning and fusing the movement characteristics of the human body and the robot, and using an adaptive weighted aggregation module, the accuracy of human body motion prediction in human-computer interactive scenarios is solved, achieving higher precision prediction effects.
Patent Information
- Application Number
- CN202510595691.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The prior art has low accuracy in human body motion prediction in human-computer interaction scenarios, mainly due to the reduction in prediction accuracy caused by the differences in robots and real human bodies.
By obtaining the historical motion sequences of the human body and the robot, extracting and aligning the motion characteristics, combining spatial and temporal interaction characteristics, the adaptive weighted aggregation module is used to fusion interaction characteristics to predict the trajectory of each link node of the human body.
The accuracy of human body motion prediction in human-computer interaction scenarios is improved, and the interaction between human-body and robot can be predicted more accurately, especially in complex interaction scenarios.
Smart Images

Figure CN120095838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and in particular, to a human motion prediction method, device, and medium for human-computer interaction scenarios. Background Art
[0002] Robots play an important role in the industrial production field. Human-computer interaction is one of the important functions of robots. Therefore, understanding and analyzing human actions is of great significance to robots. Human motion prediction refers to predicting future human motion based on the historical motion skeleton sequence of the human body. The existing technology directly applies multi-person motion prediction to human motion prediction in human-computer interaction scenarios. That is, the existing technology regards the robot as one of the people in multi-person motion to predict the motion of the human body for interaction. Since there are differences in the structures of robots and real human bodies, directly regarding the robot as a human body and using the multi-person motion prediction method to predict the motion of people in human-computer interaction scenarios will reduce the prediction accuracy.
[0003] In summary, the human motion predicted by the existing technology in human-computer interaction scenarios has low accuracy.
[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a human motion prediction method, device, and medium for human-computer interaction scenarios, which solves the problem that the human motion predicted by the existing technology in human-computer interaction scenarios has low accuracy.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a human motion prediction method for human-computer interaction scenarios, which includes:
[0008] Obtain the human historical motion sequence and the robot historical motion sequence in the human-computer interaction scenario, and obtain human motion features based on the human historical motion sequence, and obtain robot motion features based on the robot historical motion sequence;
[0009] Perform alignment processing on the human motion features and the robot motion features to obtain human aligned motion features and robot aligned motion features with the same dimension;
[0010] Extract human interaction features between humans and robots based on the human aligned motion features and the robot aligned motion features;
[0011] Obtain the predicted trajectories of each joint point of the human body based on the human interaction features.
[0012] In one implementation, based on the human historical motion sequence, human motion features are obtained, including:
[0013] Apply a motion encoder to the human historical motion sequence to obtain human motion features, where the motion encoder consists of a discrete cosine transform module and a multi-layer perceptron.
[0014] In one implementation, based on the human alignment motion features and the robot alignment motion features, human - robot interaction features are extracted, including:
[0015] Extract spatial interaction features and temporal interaction features between the human and the robot based on the human alignment motion features and the robot alignment motion features;
[0016] Map the spatial interaction features to each joint of the human body to obtain spatially - interacted reconstruction features;
[0017] Map the temporal interaction features to each joint of the human body to obtain temporally - interacted reconstruction features, and use the temporally - interacted reconstruction features and the spatially - interacted reconstruction features as human - robot interaction features.
[0018] In one implementation, based on the human - robot interaction features, predicted trajectories of each joint point of the human body are obtained, including:
[0019] Perform weighted calculation on the human motion features, the temporally - interacted reconstruction features, and the spatially - interacted reconstruction features to obtain weighted human motion features;
[0020] Based on the weighted human motion features, obtain the global trajectory of the human body over the entire prediction duration;
[0021] Based on the global trajectory, predict the trajectory of the human body as a point to obtain a pseudo - trajectory;
[0022] Based on the global trajectory and the pseudo - trajectory, obtain the predicted trajectories of each joint point of the human body.
[0023] In one implementation, performing weighted calculation on the human motion features, the temporally - interacted reconstruction features, and the spatially - interacted reconstruction features to obtain weighted human motion features includes:
[0024] Set a first weight for the spatially - interacted reconstruction features; set a second weight for the temporally - interacted reconstruction features;
[0025] Based on the first weight and the second weight, obtain a third weight, and use the third weight as the weight of the human motion features;
[0026] According to the first weight, the second weight, and the third weight, perform weighted calculation on the human body movement feature, the time interaction reconstruction feature, and the space interaction reconstruction feature to obtain a human body weighted movement feature.
[0027] In one implementation, the first weight and the second weight are associated with the human body movement speed and the human body skeleton structure.
[0028] In one implementation, the first network obtains the human body movement feature based on the human body historical movement sequence; the second network obtains the robot movement feature based on the robot historical movement sequence; and then the first network obtains the predicted trajectories of each joint point of the human body based on the human body movement feature and the robot movement feature; wherein, when training the first network, the second network is used as an auxiliary training task to train the first network.
[0029] In one implementation, the total loss function used when training the first network is composed of the loss function corresponding to the first network and the loss function corresponding to the second network.
[0030] In a second aspect, an embodiment of the present invention further provides a terminal device, wherein the terminal device includes a memory, a processor, and a human body movement prediction program for a human-computer interaction scenario stored in the memory and executable on the processor. When the processor executes the human body movement prediction program for the human-computer interaction scenario, the steps of the above-mentioned human body movement prediction method for the human-computer interaction scenario are implemented.
[0031] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a human body movement prediction program for a human-computer interaction scenario is stored. When the human body movement prediction program for the human-computer interaction scenario is executed by a processor, the steps of the above-mentioned human body movement prediction method for the human-computer interaction scenario are implemented.
[0032] Beneficial effects: The present invention extracts the human body movement feature and the robot movement feature from the historical movement sequences of the human body and the robot, performs alignment processing on the human body movement feature and the robot movement feature, then extracts the features generated by the human body during the interaction process based on the aligned movement features of the two, that is, the human body interaction feature, and finally predicts the trajectories of each joint point of the human body according to the human body interaction feature. From the above analysis, it can be seen that the present invention not only considers the influence of the robot on the human body movement trajectory in the human-computer interaction scenario, but also aligns the movement features of the two to alleviate the inherent differences between them, and finally improves the prediction accuracy of the human body movement trajectory. Description of the Drawings
[0033] Figure 1It is the overall flowchart of the present invention;
[0034] Figure 2 It is the double-branch network structure diagram in the embodiment of the present invention;
[0035] Figure 3 It is the schematic diagram of the time-space interaction module in the embodiment of the present invention;
[0036] Figure 4 It is the schematic diagram of the adaptive weighted aggregation module in the embodiment of the present invention;
[0037] Figure 5 It is the schematic diagram for comparing prediction results in the scenario of human and robotic arm;
[0038] Figure 6 It is the schematic diagram for comparing prediction results in the scenario of human and robot dog;
[0039] Figure 7 It is the internal structure principle block diagram of the terminal device provided in the embodiment of the present invention. Specific embodiments
[0040] The following combines the embodiments and the accompanying drawings of the specification to clearly and completely describe the technical solutions in the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0041] Through research, it is found that robots play an important role in the industrial production field, and human-robot interaction is one of the important functions of robots. Therefore, understanding and analyzing human movements is of great significance to robots. Human motion prediction refers to predicting future human movements based on the historical motion skeleton sequences of human-robot interaction. The prior art directly applies multi-person motion prediction to human motion prediction in the human-robot interaction scenario, that is, the prior art regards the robot as one of the people in multi-person motion to predict the motion of the human for interaction. Since there are differences in the structures of robots and real humans, directly regarding the robot as a human to adopt the multi-person motion prediction method to predict the motion of people in the human-robot interaction scenario will reduce the prediction accuracy.
[0042] To solve the above technical problems, the present invention provides a human motion prediction method, device and medium for the human-robot interaction scenario, which solves the problem that the accuracy of the human motion predicted in the human-robot interaction scenario in the prior art is relatively low.
[0043] The human motion prediction method for the human-robot interaction scenario in this embodiment can be applied to a terminal device, and the terminal device can be a terminal product with a data playback function, such as a computer, etc. In this embodiment, as Figure 1As shown in the figure, the human motion prediction method for the human-computer interaction scenario specifically includes the following steps:
[0044] S100, obtain the human historical motion sequence and the robot historical motion sequence in the human-computer interaction scenario, and based on the human historical motion sequence, obtain the human motion characteristics, and based on the robot historical motion sequence, obtain the robot motion characteristics;
[0045] S200, perform alignment processing on the human motion characteristics and the robot motion characteristics to obtain the human aligned motion characteristics and the robot aligned motion characteristics with the same dimension;
[0046] S300, extract the human interaction characteristics between humans and robots based on the human aligned motion characteristics and the robot aligned motion characteristics;
[0047] S400, obtain the predicted trajectories of each joint point of the human based on the human interaction characteristics.
[0048] In the first embodiment, steps S100 to S400 are completed on the first network, and the first network is Figure 2 the upper half of, which is used to predict the trajectories of each joint point of the human, as Figure 2 shown, the first network includes a motion encoder, a spatio-temporal interaction module, an adaptive weighted aggregation module, a trajectory analysis module, a motion analysis module, and a motion decoder.
[0049] Figure 2 The lower half of is the second network for predicting the trajectories of each joint point of the robot. The second network has the same structure as the first network, and the first network and the second network constitute a dual-branch network.
[0050] The motion encoder of the first network consists of a discrete cosine transform module and a multi-layer perceptron (the multi-layer perceptron is MLP, and MLP is Multilayer Perceptron).
[0051] During prediction, the human historical motion sequence in the time domain form is input into the discrete cosine transform of the first network, and the result output by the discrete cosine transform is then input into the multi-layer perceptron to obtain the human motion characteristics in the frequency domain form. The robot historical motion sequence in the time domain form is input into the motion encoder of the second network to obtain the robot motion characteristics. The spatio-temporal interaction module of the first network obtains the human interaction characteristics based on the human motion characteristics and the robot motion characteristics. The human interaction characteristics are then sequentially processed by the adaptive weighted aggregation module, the trajectory analysis module, the motion analysis module, and the motion decoder of the first network to obtain the predicted trajectories of each joint point (this trajectory is a time domain trajectory).
[0052] The human historical motion sequence includes the human point cloud skeletons of each frame, and each frame of the human point cloud skeleton represents the position information of each joint point of the human body at each moment during the interaction with the robot (the position information of each joint is represented by the coordinates of each joint). The robot historical motion sequence also includes the robot point cloud skeletons of each frame, and each frame of the robot point cloud skeleton represents the position information of each joint point of the robot at each moment during the interaction with the human body.
[0053] Train the first network so that the trained first network can be used to predict the trajectories of each joint point of the human body. When training the first network, the second network is used as an auxiliary task, that is, training the Figure 2 dual-branch network in it. During training, collect the human body sample motion sequence and the robot sample motion sequence generated during the interaction between the human body and the robot as samples, and input the human body sample motion sequence and the robot sample motion sequence used as the training data set into the motion encoder of the first network and the motion encoder of the second network respectively. The first network predicts the coordinates of each joint point of the human body at each subsequent moment based on the human body sample motion sequence, that is, the first network outputs the predicted values of the coordinates of each joint point , according to the predicted values and the true values of the coordinates of each joint point corresponding to the human body at each subsequent moment , calculate the motion error of the first network ; The second network predicts the coordinates of each joint point of the robot at each subsequent moment based on the robot sample motion sequence, that is, the second network outputs the predicted values of the coordinates of each joint point , according to the predicted values and the true values of the coordinates of each joint point corresponding to the robot at each subsequent moment , calculate the motion error of the second network .
[0054] ;
[0055] ;
[0056] represents the predicted value of the -th joint point on the human skeleton point cloud of the t-th frame, represents the true value of the -th joint point on the human skeleton point cloud of the t-th frame, where the t-th frame corresponds to the t-th moment in each subsequent moment, L represents the number of joint points of the human body used as a sample, M is the number of frames included in the subsequent moments, represents the human body; represents the predicted value of the -th joint point on the robot skeleton point cloud of the t-th frame, represents the true value of the The true value of each joint point, represents the robot, K represents the number of joint points of the robot, represents the Euclidean distance.
[0057] Add the motion error of the first network and the motion error of the second network to obtain the total loss function of the dual-branch network :
[0058] ;
[0059] According to the value of the total loss function adjust the parameters of the first network and the second network for iterative training until the number of iterations reaches the set number or the total loss function is less than the set value, terminate the training, and obtain the trained first network. The above adjustment of the parameters of the two networks is to adjust Figure 2 the parameters of the motion encoders of the two networks, the parameters of the spatio-temporal interaction module, the parameters of the adaptive weighted aggregation module, the parameters of the trajectory analysis module, the parameters of the motion analysis module, and the parameters of the motion decoder in
[0060] Embodiment 2, based on Embodiment 1, in this embodiment, the specific steps of extracting human motion features in step S100 include: applying a motion encoder to the human historical motion sequence to obtain human motion features, and the motion encoder is composed of a discrete cosine transform and a multi-layer perceptron module.
[0061] That is, first apply a discrete cosine transform to each frame of human body image in the human historical motion sequence, and then input the result of the discrete cosine transform into a multi-layer perceptron (the multi-layer perceptron is MLP) to initially obtain the features of the human historical motion sequence, that is, human motion features .
[0062] Perform the same processing on the robot historical motion sequence in this embodiment to obtain robot motion features .
[0063] Among them, the human historical motion sequence is represented by the robot historical motion sequence is represented by where represents the number of frames in the historical motion sequence, represents the set, J represents the number of joint points of the human body to be predicted, and K represents the number of joint points of the robot.
[0064] Embodiment 3, based on Embodiment 1 or Embodiment 2, in this embodiment, the alignment processing in step S200 is implemented through Figure 2 the spatio-temporal interaction module in, and the function of the spatio-temporal interaction module is asFigure 3 As shown, it includes three functions: alignment, interaction, and reconstruction. Among them, alignment means aligning human motion features and robot motion features into a feature space of the same dimension through an MLP (MLP stands for multi-layer perceptron), so as to solve the inherent differences between humans and robots, such as skeleton differences and motion behavior differences. That is, unifying the number of joint points corresponding to the features contained in the human motion features and the number of joints corresponding to the features contained in the robot motion features through deep learning, so that the number of joint points of the two is the same.
[0065] In this embodiment, the interaction function of the spatio-temporal interaction module of the first network is used to implement step S300, including the following specific steps S301, S302, and S303:
[0066] S301, according to the human aligned motion features and the robot aligned motion features, extract the spatial interaction features and temporal interaction features between humans and robots.
[0067] S302, map the spatial interaction features to each joint of the human body to obtain spatial interaction reconstruction features .
[0068] S303, map the temporal interaction features to each joint of the human body to obtain temporal interaction reconstruction features , and use the temporal interaction reconstruction features and the spatial interaction reconstruction features as human interaction features.
[0069] The above interaction function includes spatial interaction and temporal interaction. Based on the attention mechanism and the residual connection layer, the spatio-temporal interaction relationship between the human body and the robot is captured, and attention interaction is performed in the temporal dimension and the spatial dimension respectively to capture a more comprehensive interaction relationship between the two.
[0070] Among them, spatial interaction is to identify the close relationship between humans and robots through the coordinate similarity of these two features, namely the human aligned motion features and the robot aligned motion features, so as to capture the interaction behavior characterized by proximity. Among them, proximity is the distance when humans and robots are in contact, and the human features corresponding to when humans and robots are in contact are the spatial interaction features, including touch features and kicking features.
[0071] Temporal interaction is to obtain temporal interaction features by detecting the interaction trend across multiple frames and the interaction behaviors occurring at a long distance between humans and robots (long distance means the distance when humans and robots are not in contact).
[0072] In this embodiment, as Figure 2 shown, the spatio-temporal interaction module of the second network processes the robot motion features in the same way as the spatio-temporal interaction module of the first network to obtain the robot contact reconstruction features when the robot comes into contact with the human body (the robot contact reconstruction features are the spatial reconstruction features of the robot), and the robot non-contact reconstruction features when the robot is not in contact with the human body (the robot non-contact reconstruction features are the temporal reconstruction features of the robot).
[0073] This embodiment enables the model to comprehensively understand the context dynamic information in the human-robot system through temporal interaction and spatial interaction, so that it can proficiently handle diverse interaction relationships. Reconstruction is to restore the motion features after interaction to the original skeleton structures of the human and the robot.
[0074] Embodiment 4, based on Embodiment 1 or Embodiment 2 or Embodiment 3, in this embodiment, step S400 includes the following specific steps S401 to S406:
[0075] S401, set the first weight of the spatial interaction reconstruction features ; set the second weight of the temporal interaction reconstruction features ; ; ;
[0076] S402, according to the first weight and the second weight , obtain the third weight (the third weight is ), and use the third weight as the weight of the human motion features ;
[0077] S403, according to the first weight, the second weight, and the third weight, perform weighted calculation on the human motion features, the temporal interaction reconstruction features, and the spatial interaction reconstruction features to obtain the weighted human motion features :
[0078] ;
[0079] Steps S401, S402, and S403 are implemented on the adaptive weighted aggregation module as shown in Figure 4 , and and are associated with the human skeleton and the human movement speed.
[0080] S404, based on the weighted human motion features , obtain the global trajectory G of the human body within all prediction time durations.
[0081] As shown in Figure 2 , the trajectory analysis module passes the weighted human motion features through an MLP-based spatial encoder Encode the spatial dimension, and then use an MLP-based temporal encoder to encode the weighted motion features of the human body in the time dimension to obtain the global trajectory G of the human body over the entire prediction duration. For example, if the entire prediction duration includes the duration corresponding to two frames, then the global trajectory G is the trajectory of each joint point of the human body within the two-frame duration.
[0082] S405. According to the global trajectory G, predict the trajectory of the human body as a point to obtain the pseudo-trajectory Z.
[0083] Generate a center point coordinate for each time frame as the trajectory information of that frame, that is, the pseudo-trajectory Z. Use an MLP (MLP is a multi-layer perceptron) on the global trajectory G to generate the pseudo-trajectory Z.
[0084] S406. According to the global trajectory G and the pseudo-trajectory Z, obtain the predicted trajectories of each joint point of the human body.
[0085] S404 and S405 are implemented on the trajectory analysis module as shown in Figure 2 . After the trajectory analysis module generates the pseudo-trajectory Z, it then outputs the trajectory based on the pseudo-trajectory Z and the global trajectory G :
[0086] ;
[0087] As shown in Figure 2 , the motion analysis module then predicts and infers the trajectories of each joint point based on the trajectory , that is, maps the trajectory into the trajectories of each joint point, thereby obtaining the predicted trajectories of each joint point , , where N is the number of frames to be predicted.
[0088] This embodiment uses the pseudo-trajectory as the global trajectory instead of the trajectory of the human hip joint as the global trajectory. This is because the assumption that the hip joint trajectory approximates global motion is not always applicable, especially when the motion is mainly concentrated in the upper body. In addition, the trajectory analysis module indirectly assists in the conversion from global coordinates to relative coordinates, allowing the subsequent motion analysis module set to predict local postures. The motion analysis module stacks multiple MLP modules to analyze and predict detailed future motions from both temporal and spatial perspectives.
[0089] Apply the prediction method of the present invention and the prediction method of the prior art to the human-robot arm interaction scenario respectively to predict human motion and robot arm motion; apply the prediction method of the present invention and the prediction method of the prior art to the human-machine dog interaction scenario respectively to predict human motion and machine dog motion. Figure 5 Corresponding to the human-robot arm interaction scenario,Figure 6 Corresponding to the scenario of human-robot dog interaction. Such as Figure 5 and Figure 6 In, siMLPe-H, siMLPe-HR, GCNext-H, and GCNext-HR are four baseline methods in the prior art (i.e., four existing prediction methods), Ours-HR is the prediction method of the present invention, and GT is the ground truth. From Figure 5 It can be seen that the method of the present invention predicts more reasonable human postures, especially in terms of the postures of the left leg and right hand. From Figure 6 It can be seen that the present invention provides more accurate predictions for the interaction position and timing of the human hand touching the robot, showing significant improvements compared to the baseline models.
[0090] In summary, when predicting human motion, the present invention takes into account the influence of the robot on the human, rather than analyzing human motion in isolation from the external environment, has a more comprehensive understanding of human motion, and thus improves the prediction accuracy. In addition, considering that the interaction between the human body and the robot is dynamic, the present invention proposes a new adaptive weighted aggregation module for adaptively fusing interactive features and non-interactive features. The context-based weights are organized in a per-frame format and can flexibly consider the influence of each time step, which enables the present invention to adaptively adapt to complex interaction scenarios.
[0091] Based on the above embodiments, the present invention also provides a terminal device, and its principle block diagram can be as Figure 7 shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a human motion prediction method for a human-machine interaction scenario. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.
[0092] Those skilled in the art can understand that Figure 7 the principle block diagram shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0093] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a human motion prediction program for a human-computer interaction scenario stored in the memory and executable on the processor. When the processor executes the human motion prediction program for the human-computer interaction scenario, the following operation instructions are implemented:
[0094] Obtain the human historical motion sequence and the robot historical motion sequence in the human-computer interaction scenario,
[0095] and obtain human motion features based on the human historical motion sequence, and obtain robot motion features based on the robot historical motion sequence;
[0096] Perform alignment processing on the human motion features and the robot motion features to obtain human aligned motion features and robot aligned motion features with the same dimension;
[0097] Extract human interaction features between humans and robots based on the human aligned motion features and the robot aligned motion features;
[0098] Obtain the predicted trajectories of each joint point of the human body based on the human interaction features.
[0099] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A human motion prediction method for human-computer interaction scenarios, characterized in that, Including: Obtain the human historical motion sequence and the robot historical motion sequence in the human-computer interaction scenario, and based on the human historical motion sequence, obtain human motion features, and based on the robot historical motion sequence, obtain robot motion features; Perform alignment processing on the human motion features and the robot motion features to obtain human-aligned motion features and robot-aligned motion features with the same dimension; Extract human interaction features between humans and robots based on the human-aligned motion features and the robot-aligned motion features; Based on the human interaction features, obtain the predicted trajectories of each joint point of the human body.
2. The human motion prediction method for the human-computer interaction scenario according to claim 1, wherein Obtain human motion features based on the human historical motion sequence, including: Apply a motion encoder to the human historical motion sequence to obtain human motion features, and the motion encoder is composed of a discrete cosine transform module and a multi-layer perceptron.
3. The human motion prediction method for the human-computer interaction scenario according to claim 1, characterized in that Extract human interaction features between humans and robots based on the human-aligned motion features and the robot-aligned motion features, including: Extract spatial interaction features and temporal interaction features between humans and robots based on the human-aligned motion features and the robot-aligned motion features; Map the spatial interaction features to each joint of the human body to obtain spatial interaction reconstruction features; Map the temporal interaction features to each joint of the human body to obtain temporal interaction reconstruction features, and use the temporal interaction reconstruction features and the spatial interaction reconstruction features as human interaction features.
4. The human motion prediction method for the human-computer interaction scenario according to claim 3, wherein Obtain the predicted trajectories of each joint point of the human body based on the human interaction features, including: Perform weighted calculation on the human motion features, the temporal interaction reconstruction features, and the spatial interaction reconstruction features to obtain weighted human motion features; Based on the weighted human motion features, obtain the global trajectory of the human body within all prediction time durations; Based on the global trajectory, predict the trajectory of the human body as a point to obtain a pseudo-trajectory; Based on the global trajectory and the pseudo-trajectory, obtain the predicted trajectories of each joint point of the human body.
5. The human motion prediction method for the human-computer interaction scenario according to claim 4, wherein Perform weighted calculation on the human motion features, the temporal interaction reconstruction features, and the spatial interaction reconstruction features to obtain weighted human motion features, including: Set the first weight of the spatial interaction reconstruction features; set the second weight of the temporal interaction reconstruction features; Based on the first weight and the second weight, obtain a third weight, and use the third weight as the weight of the human motion features; Based on the first weight, the second weight, and the third weight, perform weighted calculation on the human motion features, the temporal interaction reconstruction features, and the spatial interaction reconstruction features to obtain weighted human motion features.
6. The human motion prediction method for human-computer interaction scenarios according to claim 5, wherein The first weight and the second weight are associated with the human motion speed and the human skeleton structure.
7. The human motion prediction method for human-computer interaction scenarios according to claim 1, characterized in that Obtain the human body movement characteristics based on the human body historical movement sequence through the first network; obtain the robot movement characteristics based on the robot historical movement sequence through the second network; and then obtain the predicted trajectories of each joint point of the human body based on the human body movement characteristics and the robot movement characteristics through the first network; wherein, when training the first network, use the second network as an auxiliary training task to train the first network.
8. The method for predicting human motion for a human-computer interaction scenario according to claim 7, wherein, The total loss function used when training the first network is composed of the loss function corresponding to the first network and the loss function corresponding to the second network.
9. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a human body movement prediction program for a human-computer interaction scenario stored in the memory and executable on the processor. When the processor executes the human body movement prediction program for a human-computer interaction scenario, the steps of the human body movement prediction method for a human-computer interaction scenario according to any one of claims 1-8 are implemented.
10. A computer-readable storage medium, characterized in that, A human body movement prediction program for a human-computer interaction scenario is stored on the computer-readable storage medium. When the human body movement prediction program for a human-computer interaction scenario is executed by a processor, the steps of the human body movement prediction method for a human-computer interaction scenario according to any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Lightweight human skeleton interaction behavior reasoning network structure based on multilayer perceptron
CN115862152A
Pedestrian trajectory prediction model, pedestrian trajectory prediction method and related device
CN118247307A