Human body motion prediction method and device for human-computer interaction scene, and medium

By aligning the motion characteristics of the human body and the robot in the human-computer interaction scenario and extracting the human-computer interaction features, the problem of low prediction accuracy of human-body motion in the prior art is solved, and higher prediction accuracy and more comprehensive human-computer interaction understanding are achieved.

CN120095838AActive Publication Date: 2025-06-06PEKING UNIV SHENZHEN GRADUATE SCHOOL

Patent Information

Application Number
CN202510595691.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The accuracy of human movement predicted by the prior art in human-computer interaction scenarios is low, mainly due to the differences in the structure of the robot and the real human body, resulting in a decrease in prediction accuracy when directly applying the multi-person motion prediction method.

Method used

By obtaining the historical motion sequence of the human body and the robot in the human-computer interaction scenario, extracting the motion characteristics of the human body and the robot, and performing alignment processing, the human body alignment motion characteristics and the robot alignment motion characteristics with the same dimensions are obtained. Then, based on these aligned motion characteristics, the human interaction characteristics between humans and machines are extracted, and finally the trajectory of each joint node of the human body is predicted based on these characteristics.

Benefits of technology

The prediction accuracy of human body movement trajectory in human-computer interaction scenarios is improved, the impact of robot on human body movement trajectory is considered, and the structural differences between human body and robot are alleviated through alignment processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120095838A_ABST
    Figure CN120095838A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human-computer interaction, in particular to a human-computer interaction scene-oriented human motion prediction method and device and a medium. Human body motion features and robot motion features are extracted from historical motion sequences of a human body and a robot, the human body motion features and the robot motion features are aligned, features, namely human body interaction features, generated in the interaction process of the human body are extracted on the basis of the aligned motion features of the human body and the robot, and the human body interaction features are extracted. And finally, according to the human body interaction characteristics, predicting the track of each joint point of the human body. According to the analysis, the influence of the robot on the human body movement track is considered in the man-machine interaction scene, the movement characteristics of the robot and the human body movement track are aligned, the inherent difference between the robot and the human body movement track is relieved, and finally the prediction precision of the human body movement track is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a method, device and medium for predicting human motion in a human-computer interaction scenario. Background Art

[0002] Robots play an important role in the field of industrial production. Human-machine interaction is one of the important functions of robots. Therefore, understanding and analyzing human movements is of great significance to robots. Human motion prediction refers to predicting future human motion based on the historical motion skeleton sequence of the human body. The existing technology directly applies multi-person motion prediction to human motion prediction in human-machine interaction scenarios. That is, the existing technology uses the robot as one of the people in the multi-person motion to predict the motion of the interacting human body. Due to the difference in structure between the robot and the real human body, directly using the robot as a human body to predict the motion of the person in the human-machine interaction scenario using the multi-person motion prediction method will reduce the prediction accuracy.

[0003] In summary, the existing technology has low accuracy in predicting human motion in human-computer interaction scenarios.

[0004] Therefore, the prior art still needs to be improved and enhanced. Summary of the invention

[0005] To solve the above technical problems, the present invention provides a method, device and medium for predicting human motion in human-computer interaction scenarios, which solves the problem of low accuracy of human motion prediction in human-computer interaction scenarios in the prior art.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for predicting human motion in a human-computer interaction scenario, comprising: Acquire a human body historical motion sequence and a robot historical motion sequence in a human-machine interaction scene, and obtain human body motion features based on the human body historical motion sequence, and obtain robot motion features based on the robot historical motion sequence; Performing alignment processing on the human motion feature and the robot motion feature to obtain human alignment motion features and robot alignment motion features having the same dimension; Extracting human interaction features between a human and a robot based on the human alignment motion features and the robot alignment motion features; According to the human body interaction characteristics, the predicted trajectory of each joint point of the human body is obtained.

[0007] In one implementation, obtaining a human motion feature according to the human historical motion sequence includes: A motion encoder is applied to the human body historical motion sequence to obtain human body motion features, wherein the motion encoder is composed of a discrete cosine transform module and a multi-layer perceptron.

[0008] In one implementation, extracting the human interaction feature between the human and the robot based on the human alignment motion feature and the robot alignment motion feature includes: Extracting spatial interaction features and temporal interaction features between humans and robots based on the human body alignment motion features and the robot alignment motion features; Mapping the spatial interaction features to various joints of the human body to obtain spatial interaction reconstruction features; The time interaction feature is mapped to each joint of the human body to obtain a time interaction reconstruction feature, and the time interaction reconstruction feature and the space interaction reconstruction feature are used as a human body interaction feature.

[0009] In one implementation, obtaining predicted trajectories of various joints of a human body based on the human body interaction features includes: Performing weighted calculation on the human body motion feature, the time interaction reconstruction feature, and the space interaction reconstruction feature to obtain a human body weighted motion feature; Based on the weighted motion characteristics of the human body, a global trajectory of the human body within the entire prediction duration is obtained; According to the global trajectory, predict the trajectory of the human body as a point to obtain a pseudo trajectory; According to the global trajectory and the pseudo trajectory, the predicted trajectory of each joint point of the human body is obtained.

[0010] In one implementation, weighted calculation is performed on the human motion feature, the temporal interactive reconstruction feature, and the spatial interactive reconstruction feature to obtain a human weighted motion feature, including: Setting a first weight of the spatial interactive reconstruction feature; Setting a second weight of the temporal interactive reconstruction feature; Obtaining a third weight according to the first weight and the second weight, and using the third weight as the weight of the human body motion feature; The human body motion feature, the time interaction reconstruction feature and the space interaction reconstruction feature are weightedly calculated according to the first weight, the second weight and the third weight to obtain a weighted human body motion feature.

[0011] In one implementation, the first weight and the second weight are associated with a human body movement speed and a human body skeletal structure.

[0012] In one implementation, the human motion features are obtained based on the human historical motion sequence through a first network; the robot motion features are obtained based on the robot historical motion sequence through a second network; and the predicted trajectory of each joint of the human body is obtained based on the human motion features and the robot motion features through the first network; wherein, when training the first network, the second network is used as an auxiliary training task to train the first network.

[0013] In one implementation, a total loss function used when training the first network is composed of a loss function corresponding to the first network and a loss function corresponding to the second network.

[0014] In a second aspect, an embodiment of the present invention further provides a terminal device, wherein the terminal device comprises a memory, a processor, and a human motion prediction program for human-computer interaction scenarios stored in the memory and executable on the processor, and when the processor executes the human motion prediction program for human-computer interaction scenarios, the steps of the above-mentioned human motion prediction method for human-computer interaction scenarios are implemented.

[0015] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which is stored a human motion prediction program for human-computer interaction scenarios. When the human motion prediction program for human-computer interaction scenarios is executed by a processor, the steps of the above-mentioned human motion prediction method for human-computer interaction scenarios are implemented.

[0016] Beneficial effects: The present invention extracts human motion features and robot motion features from the historical motion sequences of the human body and the robot, and aligns the human motion features and robot motion features, and then extracts the features generated by the human body during the interaction process based on the motion features of the two after alignment, namely, the human body interaction features, and finally predicts the trajectory of each joint point of the human body based on the human body interaction features. From the above analysis, it can be seen that the present invention not only considers the impact of the robot on the human body's motion trajectory in the human-computer interaction scenario, but also aligns the motion features of the two to alleviate the inherent differences between the two, and finally improves the prediction accuracy of the human body's motion trajectory. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall flow chart of the present invention; Figure 2 A double-branch network structure diagram in an embodiment of the present invention; Figure 3 is a schematic diagram of a time-space interaction module in an embodiment of the present invention; Figure 4 is a schematic diagram of an adaptive weighted aggregation module in an embodiment of the present invention; Figure 5 A schematic diagram showing the comparison of prediction results in human and robotic arm scenarios; Figure 6 A schematic diagram showing the comparison of prediction results in the human and robot dog scenarios; Figure 7 This is a block diagram of the internal structure principle of the terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following is a clear and complete description of the technical solution of the present invention in combination with the embodiments and the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] Research has found that robots play an important role in the field of industrial production, and human-machine interaction is one of the important functions of robots. Therefore, understanding and analyzing human movements is of great significance to robots. Human motion prediction refers to predicting future human motion based on the historical motion skeleton sequence of human-robot interaction. The existing technology directly applies multi-person motion prediction to human motion prediction in human-machine interaction scenarios. That is, the existing technology uses the robot as one of the people in the multi-person motion to predict the motion of the interacting human body. Due to the difference in structure between the robot and the real human body, directly using the robot as a human body to predict the motion of the human in the human-machine interaction scenario using the multi-person motion prediction method will reduce the prediction accuracy.

[0020] To solve the above technical problems, the present invention provides a method, device and medium for predicting human motion in human-computer interaction scenarios, which solves the problem of low accuracy of human motion prediction in human-computer interaction scenarios in the prior art.

[0021] The human motion prediction method for human-computer interaction scenarios of this embodiment can be applied to a terminal device, which can be a terminal product with a data playback function, such as a computer. Figure 1 As shown in , the human motion prediction method for human-computer interaction scenarios specifically includes the following steps: S100, obtaining a human body historical motion sequence and a robot historical motion sequence in a human-machine interaction scene, and obtaining human body motion features according to the human body historical motion sequence, and obtaining robot motion features according to the robot historical motion sequence; S200, aligning the human motion feature and the robot motion feature to obtain human alignment motion features and robot alignment motion features having the same dimension; S300, extracting a human interaction feature between a human and a robot based on the human alignment motion feature and the robot alignment motion feature; S400, obtaining a predicted trajectory of each joint point of the human body according to the human body interaction characteristics.

[0022] In the first embodiment, steps S100 to S400 are performed on the first network. Figure 2 The upper part is used to predict the trajectory of each joint point of the human body, such as Figure 2 As shown, the first network includes a motion encoder, a spatiotemporal interaction module, an adaptive weighted aggregation module, a trajectory analysis module, a motion analysis module, and a motion decoder.

[0023] Figure 2 The lower part is a second network for predicting the trajectory of each joint point of the robot. The second network has the same structure as the first network, and the first network and the second network constitute a double-branch network.

[0024] The motion encoder of the first network consists of a discrete cosine transform module and a multilayer perceptron (MLP).

[0025] During prediction, the human body historical motion sequence in time domain is input into the discrete cosine transform of the first network, and the output of the discrete cosine transform is then input into the multi-layer perceptron to obtain the human body motion features in frequency domain. The robot historical motion sequence in time domain is input into the motion encoder of the second network to obtain the robot motion features. The spatiotemporal interaction module of the first network obtains the human body interaction features based on the human body motion features and the robot motion features. The human body interaction features are then processed in sequence by the adaptive weighted aggregation module, trajectory analysis module, motion analysis module, and motion decoder of the first network to obtain the predicted trajectory of each joint point (the trajectory is the time domain trajectory).

[0026] The human body historical motion sequence includes each frame of the human body point cloud skeleton, and each frame of the human body point cloud skeleton represents the position information of each joint point at each moment in the process of human body interaction with the robot (the coordinates of each joint represent the position information of each joint). The robot historical motion sequence also includes each frame of the robot point cloud skeleton, and each frame of the robot point cloud skeleton represents the position information of each joint point at each moment in the process of robot interaction with the human body.

[0027] The first network is trained so that it can be used to predict the trajectory of each joint of the human body. When training the first network, the second network is used as an auxiliary task, that is, training Figure 2During training, the human sample motion sequence and robot sample motion sequence generated by the human body and the robot in the interaction process are collected as samples, and the human sample motion sequence and robot sample motion sequence as training data sets are respectively input into the motion encoder of the first network and the motion encoder of the second network. The first network predicts the coordinates of each joint point of the human body at each subsequent moment based on the human sample motion sequence, that is, the first network outputs the predicted value of the coordinates of each joint point , according to the predicted value and the true value of the coordinates of each joint point corresponding to the human body at each subsequent moment , calculate the motion error of the first network ; The second network predicts the coordinates of each joint point of the robot at each subsequent moment based on the robot sample motion sequence, that is, the second network outputs the predicted value of the coordinates of each joint point , according to the predicted value And the true value of the coordinates of each joint point corresponding to the robot at each subsequent moment , calculate the motion error of the second network .

[0028] ; ; Represents the first frame of the human skeleton point cloud The predicted value of each joint point, Represents the first frame of the human skeleton point cloud The true value of the joint points, where the t-th frame corresponds to the t-th moment in each subsequent moment, L represents the number of joint points of the human body as a sample, and M is the number of frames included in the subsequent moments. It means human body; Represents the first point cloud of the robot skeleton in the tth frame The predicted value of each joint point, Represents the first point cloud of the robot skeleton in the tth frame The true value of each joint point, represents the robot, K represents the number of joints of the robot, Represents the Euclidean distance.

[0029] The motion error of the first network and the motion error of the second network Add together to get the total loss function of the two-branch network : ; According to the total loss function The parameters of the first network and the second network are adjusted by the value of to perform iterative training until the number of iterations reaches the set number or the total loss function If the value is less than the set value, the training is terminated and the trained first network is obtained. The above adjustment of the parameters of the two networks is to adjust Figure 2 Parameters of the motion encoders of the two networks, parameters of the spatiotemporal interaction module, parameters of the adaptive weighted aggregation module, parameters of the trajectory analysis module, parameters of the motion analysis module, and parameters of the motion decoder.

[0030] Embodiment 2 is based on embodiment 1. In this embodiment, step S100 of extracting human motion features includes the following specific steps: applying a motion encoder to the human historical motion sequence to obtain human motion features, wherein the motion encoder is composed of a discrete cosine transform and a multi-layer perceptron module.

[0031] That is, firstly, discrete cosine transform is applied to each frame of human body image in the human body historical motion sequence, and then the output of discrete cosine transform is input into multi-layer perceptron (multi-layer perceptron is MLP) to preliminarily obtain the characteristics of human body historical motion sequence, that is, human body motion characteristics. .

[0032] This embodiment performs the same processing on the robot's historical motion sequence to obtain the robot motion characteristics. .

[0033] Among them, the human body historical movement sequence is used The robot's historical motion sequence is represented by Indicates that represents the number of frames in the historical motion sequence, represents a set, J represents the number of predicted human joints, and K represents the number of robot joints.

[0034] Embodiment 3, based on embodiment 1 or embodiment 2, in this embodiment, by Figure 2 The spatiotemporal interaction module in the embodiment implements the alignment process in step S200. The function of the spatiotemporal interaction module is as follows: Figure 3 As shown in Figure 1, it includes three functions: alignment, interaction, and reconstruction. Alignment refers to the use of MLP (MLP is a multi-layer perceptron) to transform human motion features into and robot motion characteristics Align them to feature spaces of the same dimension, thereby solving the inherent differences between humans and robots, such as differences in skeletons and motion behaviors. The number of joint points and robot motion characteristics corresponding to the included features The number of joints corresponding to the included features is unified through deep learning so that the number of joint points of the two is the same.

[0035] In this embodiment, step S300 is implemented by the interaction function of the spatiotemporal interaction module of the first network, including the following specific steps S301, S302, and S303: S301, extracting spatial interaction features and temporal interaction features between a human and a robot based on the human body alignment motion features and the robot alignment motion features.

[0036] S302, mapping the spatial interaction features to various joints of the human body to obtain spatial interaction reconstruction features .

[0037] S303, mapping the time interaction feature to each joint of the human body to obtain a time interaction reconstruction feature and using the temporal interaction reconstruction feature and the spatial interaction reconstruction feature as human body interaction features.

[0038] The above-mentioned interactive functions include spatial interaction and temporal interaction. Based on the attention mechanism and residual connection layer, the spatiotemporal interaction relationship between the human body and the robot is captured. Attention interaction is performed in the temporal dimension and the spatial dimension respectively to capture a more comprehensive interaction relationship between the two.

[0039] The spatial interaction is to identify the close distance relationship between humans and robots through the coordinate similarity of the human body alignment motion feature and the robot alignment motion feature, so as to capture the interactive behavior characterized by close distance. The close distance is the distance when humans and robots are in contact, and the corresponding human features when humans and robots are in contact are the spatial interaction features, including touch features and kicking features.

[0040] Temporal interaction is to obtain temporal interaction features by detecting the interaction trends across multiple frames and the interaction behaviors that occur at a long distance between humans and machines (long distance means the distance when humans and machines are not in contact).

[0041] In this embodiment, if Figure 2 As shown, the spatiotemporal interaction module of the second network processes the robot motion features in the same way as the spatiotemporal interaction module of the first network. , get the robot contact reconstruction features when the robot and the human body are in contact (Robot contact reconstruction features are also the robot's spatial reconstruction features), as well as the robot's non-contact reconstruction features when the robot and the human body are not in contact (The robot's non-contact reconstruction feature is the robot's time reconstruction feature).

[0042] This embodiment enables the model to fully understand the contextual dynamic information in the human-machine system through temporal interaction and spatial interaction, so that it can skillfully deal with diverse interactive relationships. Reconstruction is to restore the motion features after interaction to the original skeleton structure of the human and robot.

[0043] Embodiment 4, based on embodiment 1 or embodiment 2 or embodiment 3, in this embodiment, step S400 includes the following specific steps S401 to S406: S401, setting the spatial interactive reconstruction feature The first weight ; Set the time interaction reconstruction feature The second weight ; S402, based on the first weight and the second weight , and get the third weight (the third weight is ), taking the third weight as the human motion feature The weight of S403, performing weighted calculation on the human motion feature, the temporal interaction reconstruction feature, and the spatial interaction reconstruction feature according to the first weight, the second weight, and the third weight to obtain a human weighted motion feature : ; Steps S401, S402, and S403 are performed as follows. Figure 4 The adaptive weighted aggregation module shown is implemented on and Associated with the human skeleton and the speed at which the human body moves.

[0044] S404, based on the weighted motion features of the human body , and obtain the global trajectory G of the human body in the entire prediction time.

[0045] like Figure 2 The trajectory analysis module shown in the figure uses an MLP-based spatial encoder to weight the motion features of the human body. The spatial dimension is encoded, and then the weighted motion features of the human body are encoded through the MLP-based temporal encoder The time dimension is encoded to obtain the global trajectory G of the human body in the entire predicted duration. For example, if the entire predicted duration includes the duration corresponding to two frames, then the global trajectory G is the trajectory of each joint point of the human body in the two-frame duration.

[0046] S405 , predicting the trajectory of the human body as a point based on the global trajectory G to obtain a pseudo trajectory Z.

[0047] A center point coordinate is generated for each time frame as the trajectory information of the frame, namely the pseudo trajectory Z. The pseudo trajectory Z is generated for the global trajectory G using MLP (MLP is a multi-layer perceptron).

[0048] S406: According to the global trajectory G and the pseudo trajectory Z, a predicted trajectory of each joint point of the human body is obtained.

[0049] S404 and S405 are Figure 2 The trajectory analysis module is implemented on the trajectory analysis module shown in the figure. After the trajectory analysis module generates a pseudo trajectory Z, it outputs a trajectory based on the pseudo trajectory Z and the global trajectory G. : ; like Figure 2 The motion analysis module shown is based on the trajectory Predict and infer the trajectory of each joint point, that is, Mapped into the trajectory of each joint point, so as to obtain the predicted trajectory of each joint point , ,N is the number of frames that need to be predicted.

[0050] This embodiment uses a pseudo trajectory as the global trajectory instead of the trajectory of the human hip joint as the global trajectory. This is because the assumption that the hip joint trajectory approximates the global motion is not always applicable, especially when the motion is mainly concentrated in the upper body. In addition, the trajectory analysis module indirectly assists in the conversion from global coordinates to relative coordinates, allowing the subsequent set of motion analysis modules to predict local postures. The motion analysis module stacks multiple layers of MLP modules to analyze and predict detailed future motion from both time and space perspectives.

[0051] The prediction method of the present invention and the prediction method of the prior art are respectively applied in a scene of interaction between a person and a robotic arm to predict the movement of the human body and the movement of the robotic arm; the prediction method of the present invention and the prediction method of the prior art are respectively applied in a scene of interaction between a person and a robotic dog to predict the movement of the human body and the movement of the robotic dog. Figure 5 Corresponding to the interaction scene between people and robotic arms, Figure 6 Corresponding to the interaction scene between humans and robot dogs. Figure 5 and Figure 6 siMLPe-H, siMLPe-HR, GCNext-H, and GCNext-HR are four baseline methods in the prior art (that is, four existing prediction methods), Ours-HR is the prediction method of the present invention, and GT is the true value. Figure 5 It can be seen that the method of the present invention predicts more reasonable human postures, especially in terms of the postures of the left leg and the right hand. Figure 6 It can be seen that the present invention provides more accurate predictions of the interaction position and timing of the hand touching the robot, showing significant improvement over the baseline model.

[0052] In summary, the present invention takes into account the impact of the robot on the human when predicting human motion, rather than analyzing human motion in isolation from the external environment, and has a more comprehensive understanding of human motion, thereby improving the prediction accuracy. In addition, the present invention takes into account that the interaction between the human body and the robot is dynamic, and proposes a new adaptive weighted aggregation module for adaptively fusing interactive features and non-interactive features. The context-based weights are organized in a frame-by-frame format, and can flexibly consider the impact of each time step, which enables the present invention to adaptively adapt to complex interactive scenarios.

[0053] Based on the above embodiments, the present invention further provides a terminal device, whose principle block diagram can be shown as follows: Figure 7 As shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a human body motion prediction method for a human-computer interaction scenario is implemented. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.

[0054] Those skilled in the art will understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal device to which the scheme of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0055] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a human motion prediction program for a human-computer interaction scenario stored in the memory and executable on the processor. When the processor executes the human motion prediction program for the human-computer interaction scenario, the following operation instructions are implemented: Obtain the historical motion sequence of the human body and the robot in the human-computer interaction scene, And according to the human body historical motion sequence, obtain the human body motion characteristics, and according to the robot historical motion sequence, obtain the robot motion characteristics; Performing alignment processing on the human motion feature and the robot motion feature to obtain human alignment motion features and robot alignment motion features having the same dimension; Extracting human interaction features between a human and a robot based on the human alignment motion features and the robot alignment motion features; According to the human body interaction characteristics, the predicted trajectory of each joint point of the human body is obtained.

[0056] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A human motion prediction method for human-computer interaction scenarios, characterized in that: include: Acquire a human body historical motion sequence and a robot historical motion sequence in a human-machine interaction scene, and obtain human body motion features based on the human body historical motion sequence, and obtain robot motion features based on the robot historical motion sequence; Performing alignment processing on the human motion feature and the robot motion feature to obtain human alignment motion features and robot alignment motion features having the same dimension; Extracting human interaction features between a human and a robot based on the human alignment motion features and the robot alignment motion features; According to the human body interaction characteristics, the predicted trajectory of each joint point of the human body is obtained.

2. The human motion prediction method for human-computer interaction scenarios according to claim 1, characterized in that: According to the human body historical motion sequence, the human body motion features are obtained, including: A motion encoder is applied to the human body historical motion sequence to obtain human body motion features, wherein the motion encoder is composed of a discrete cosine transform module and a multi-layer perceptron.

3. The human motion prediction method for human-computer interaction scenarios according to claim 1, characterized in that: Extracting the human interaction features between the human and the robot based on the human alignment motion features and the robot alignment motion features, including: Extracting spatial interaction features and temporal interaction features between humans and robots based on the human body alignment motion features and the robot alignment motion features; Mapping the spatial interaction features to various joints of the human body to obtain spatial interaction reconstruction features; The time interaction feature is mapped to each joint of the human body to obtain a time interaction reconstruction feature, and the time interaction reconstruction feature and the space interaction reconstruction feature are used as a human body interaction feature.

4. The human motion prediction method for human-computer interaction scenarios as claimed in claim 3, characterized in that: According to the human body interaction features, the predicted trajectory of each joint point of the human body is obtained, including: Performing weighted calculation on the human body motion feature, the time interaction reconstruction feature, and the space interaction reconstruction feature to obtain a human body weighted motion feature; Based on the weighted motion characteristics of the human body, a global trajectory of the human body within the entire prediction duration is obtained; According to the global trajectory, predict the trajectory of the human body as a point to obtain a pseudo trajectory; According to the global trajectory and the pseudo trajectory, the predicted trajectory of each joint point of the human body is obtained.

5. The human body motion prediction method for human-computer interaction scenarios as claimed in claim 4, characterized in that: The human body motion feature, the time interaction reconstruction feature, and the space interaction reconstruction feature are weightedly calculated to obtain a human body weighted motion feature, including: Setting a first weight of the spatial interactive reconstruction feature; Setting a second weight of the temporal interactive reconstruction feature; Obtaining a third weight according to the first weight and the second weight, and using the third weight as the weight of the human body motion feature; The human body motion feature, the time interaction reconstruction feature and the space interaction reconstruction feature are weightedly calculated according to the first weight, the second weight and the third weight to obtain a weighted human body motion feature.

6. The human body motion prediction method for human-computer interaction scenarios as claimed in claim 5, characterized in that: The first weight and the second weight are associated with a human body movement speed and a human body skeleton structure.

7. The human body motion prediction method for human-computer interaction scenarios according to claim 1, characterized in that: The human motion features are obtained through a first network based on the human body historical motion sequence; the robot motion features are obtained through a second network based on the robot historical motion sequence; and the predicted trajectory of each joint of the human body is obtained through the first network based on the human motion features and the robot motion features; wherein, when training the first network, the second network is used as an auxiliary training task to train the first network.

8. The human body motion prediction method for human-computer interaction scenarios as claimed in claim 7, characterized in that: The total loss function used when training the first network is composed of the loss function corresponding to the first network and the loss function corresponding to the second network.

9. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a human motion prediction program for human-computer interaction scenarios stored in the memory and executable on the processor. When the processor executes the human motion prediction program for human-computer interaction scenarios, the steps of the human motion prediction method for human-computer interaction scenarios as described in any one of claims 1-8 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a human motion prediction program for human-computer interaction scenarios. When the human motion prediction program for human-computer interaction scenarios is executed by a processor, the steps of the human motion prediction method for human-computer interaction scenarios as described in any one of claims 1-8 are implemented.

Citation Information

Patent Citations

  • Lightweight human skeleton interaction behavior reasoning network structure based on multilayer perceptron

    CN115862152A

  • Pedestrian trajectory prediction model, pedestrian trajectory prediction method and related device

    CN118247307A

  • Human body action prediction method and system based on space-time interaction attention mechanism

    CN119229540A

  • Systems and methods for collision-free trajectory planning in human-robot interaction through hand movement prediction from vision

    US20190143517A1

Cited By

  • A method and system for natural task perception and dynamic recognition in complex industrial spaces

    CN122574025A