Pose prediction method and device, equipment and medium

By combining historical pose sequences and historical IMU sequences, using pose prediction sub-models, pose compensation sub-models and fusion sub-models, the problem of inaccurate pose prediction in the prior art is solved, and more accurate pose prediction is achieved.

CN120233874APending Publication Date: 2025-07-01GOERTEK INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510264580.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing pose prediction models usually use linear networks, making it difficult to learn complex nonlinear relationships, resulting in inaccurate pose prediction.

Method used

A pose prediction method is proposed to generate more accurate predicted pose sequences by obtaining historical pose sequences and historical IMU sequences, combining pose prediction sub-models, pose compensation sub-models and fusion sub-models.

Benefits of technology

Based on historical pose data, a more accurate predicted pose sequence is generated, which improves the accuracy of pose prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233874A_ABST
    Figure CN120233874A_ABST
Patent Text Reader

Abstract

The invention discloses a pose prediction method and device, equipment and a medium, and relates to the technical field of data processing. The method comprises the steps that a historical pose sequence and a historical IMU sequence corresponding to the historical pose sequence are obtained, the historical pose sequence comprises multiple pieces of historical pose data arranged according to a time sequence, and the historical IMU data comprise multiple pieces of historical IMU data arranged according to the time sequence; generating a predicted pose sequence according to the historical pose sequence, the historical IMU sequence and a pose prediction model; the predicted pose sequence comprises a plurality of future pose data arranged according to a time sequence, the pose prediction model comprises a pose prediction sub-model, a pose compensation sub-model and a fusion sub-model, the pose prediction sub-model is used for receiving a historical pose sequence, the pose compensation sub-model is used for receiving a historical IMU sequence, and the fusion sub-model is used for fusing the historical IMU sequence. The outputs of the pose prediction sub-model and the pose compensation sub-model are the input of the fusion sub-model, and the fusion sub-model is used for outputting a predicted pose sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and more specifically, to a pose prediction method, apparatus, device, and medium. Background Art

[0002] Currently, in order to reduce the display latency and stuttering of a head-mounted display device, and thus enhance the user experience of the head-mounted display device, etc., an early rendering technique is usually adopted. Specifically, a pose prediction model is used to predict the pose of the head-mounted display device, and image rendering is completed according to the predicted pose, so as to early render the image to be displayed by the head-mounted display device.

[0003] However, existing pose prediction models usually adopt a linear network and only predict based on historical pose data, which makes it difficult for the pose prediction model to learn complex non-linear relationships, and thus unable to accurately predict the pose. Therefore, a new pose prediction method is urgently needed to be proposed. Summary of the Invention

[0004] An object of the present application is to provide a new technical solution for pose prediction.

[0005] According to a first aspect of the present application, there is provided a pose prediction method, including:

[0006] Obtain a historical pose sequence and a historical IMU sequence corresponding to the historical pose sequence, where the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence;

[0007] Generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and a pose prediction model;

[0008] Wherein, the predicted pose sequence includes a plurality of future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, the pose compensation sub-model is used to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence.

[0009] Optionally, any one of the pose prediction sub-model, the pose compensation sub-model, and the fusion sub-model is one of a linear layer model, a Bi-GRU model, a GRU model, and a Bi-LSTM model.

[0010] Optionally, the pose prediction sub-model and the pose compensation sub-model are linear layer models. Before generating the predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and the pose prediction model, the method further includes:

[0011] Obtain the prediction step number, the sample step number, the first sample set, and the second sample set. The first sample set includes a plurality of pose sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of pose sample data in sequence;

[0012] Set the pose prediction model to be trained. Setting the pose prediction model to be trained includes: setting the number of neurons in the input layer of the pose prediction sub-model and the input layer of the pose compensation sub-model to be the sample step number; setting the number of neurons in the last hidden layer of the pose prediction sub-model to be the same as the number of neurons in the last hidden layer of the pose compensation sub-model; setting the number of neurons in the output layer of the fusion sub-model to be the prediction step number, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model;

[0013] Generate a plurality of training samples according to the prediction step number, the sample step number, the first sample set, and the second training sample set. One training sample includes consecutive sample step number pose sample data as samples, consecutive sample step number IMU sample data as samples, and consecutive prediction step number pose label data as labels;

[0014] Train the pose prediction model to be trained according to the plurality of training samples to obtain a trained pose prediction model.

[0015] Optionally, in the case where the pose data is represented by three-dimensional position data, the obtaining of the historical pose sequence and the historical IMU sequence corresponding to the historical pose sequence includes:

[0016] Extract a position data component of the same dimension from the historical position data to obtain the historical pose sequence;

[0017] Extract the IMU component matching the position data component from the historical IMU data to obtain the historical IMU sequence.

[0018] Optionally, in the case where the pose data is represented by quaternion attitude data, the obtaining of the historical pose sequence and the historical IMU sequence corresponding to the historical pose sequence includes:

[0019] Convert the historical quaternion attitude data into historical three-dimensional Euler angle data;

[0020] Extract the unary Euler angle data of the same dimension from the historical ternary Euler angle data to obtain a historical pose sequence;

[0021] Extract the IMU components matching the unary Euler angle data from the historical IMU data to obtain a historical IMU sequence.

[0022] Optionally, the obtaining the prediction step number, the sample step number, the first sample set, and the second sample set includes:

[0023] Display a setting input interface;

[0024] Receive the prediction step number, the sample step number, the first sample set, and the second sample set input by the setting input interface.

[0025] According to the second aspect of the present application, there is provided a pose prediction device, including:

[0026] A first acquisition module, configured to acquire a historical pose sequence and a historical IMU sequence corresponding to the historical pose sequence, where the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence;

[0027] A prediction module, configured to generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and a pose prediction model;

[0028] Wherein, the predicted pose sequence includes a plurality of future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is configured to receive the historical pose sequence, the pose compensation sub-model is configured to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are inputs of the fusion sub-model, and the fusion sub-model is configured to output the predicted pose sequence.

[0029] Optionally, the device further includes:

[0030] A second acquisition module, configured to acquire a prediction step number, a sample step number, a first sample set, and a second sample set. The first sample set includes a plurality of pose sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of pose sample data in sequence;

[0031] Set up the pose prediction model to be trained. The setup of the pose prediction model to be trained includes: setting the number of neurons in the input layer of the pose prediction sub-model and the input layer of the pose compensation sub-model to be the number of sample steps; setting the number of neurons in the last hidden layer of the pose prediction sub-model to be the same as the number of neurons in the last hidden layer of the pose compensation sub-model; setting the number of neurons in the output layer of the fusion sub-model to be the number of prediction steps, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model;

[0032] Generate a plurality of training samples according to the number of prediction steps, the number of sample steps, the first sample set, and the second training sample set. One training sample includes consecutive pose sample data of the number of sample steps as samples, consecutive IMU sample data of the number of sample steps as samples, and consecutive pose label data of the number of prediction steps as labels;

[0033] Train the pose prediction model to be trained according to the plurality of training samples to obtain a trained pose prediction model.

[0034] According to the third aspect of the present application, an electronic device is provided. The electronic device includes any one of the devices in the second aspect;

[0035] Alternatively, the electronic device includes a memory and a processor. The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method according to any one of the first aspects.

[0036] According to the fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program, when executed by a processor, implements the method according to any one of the first aspects.

[0037] The present application provides a pose prediction method, which includes: obtaining a historical pose sequence and a corresponding historical IMU sequence, where the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence; generating a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and a pose prediction model; where the predicted pose sequence includes a plurality of future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, the pose compensation sub-model is used to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are the inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence. This method can realize combining historical IMU data on the basis of historical pose data to generate a more accurate predicted pose sequence, that is, it provides a new pose prediction method that can predict a more accurate predicted pose sequence.

[0038] Other features and advantages of the present application will become clear through the following detailed description of exemplary embodiments of the present application with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present application, and together with the description thereof are used to explain the principles of the present application.

[0040] Figure 1 is a block diagram of the hardware configuration of an electronic device for implementing a pose prediction method according to an embodiment of the present application Figure 1 ;

[0041] Figure 2 is a schematic flow chart of a method for implementing a pose prediction method according to an embodiment of the present application;

[0042] Figure 3 is a schematic structural diagram of a pose prediction model according to an embodiment of the present application;

[0043] Figure 4 is a schematic structural diagram of a device for implementing a pose prediction method according to an embodiment of the present application;

[0044] Figure 5 is a block diagram of the hardware configuration of an electronic device for implementing a pose prediction method according to an embodiment of the present application Figure 2 。 DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0046] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present application or its application or use.

[0047] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0048] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0049] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.

[0050] Figure 1 is a block diagram of the hardware configuration of an electronic device for implementing a pose prediction method according to an embodiment of the present application.

[0051] The electronic device 1000 can be a terminal or a server. Further, the terminal can be a head-mounted device (such as an AR device, an MR device, and a VR device), a portable computer, a tablet computer, a handheld computer, etc. The server can be a cloud server, etc.

[0052] The electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and so on. Among them, the processor 1100 can be a central processing unit CPU, a microprocessor MCU, etc. The memory 1200 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 can perform wired or wireless communication, for example. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The input device 1600 can include, for example, a touch screen, a keyboard, etc. The user can input / output voice information through the speaker 1700 and the microphone 1800.

[0053] Although in Figure 1Multiple devices are shown for the electronic device 1000, but this application may only relate to some of the devices. For example, the electronic device 1000 only relates to the memory 1200 and the processor 1100.

[0054] In the embodiments applied to this application, the memory 1200 of the electronic device 1000 is used to store instructions for controlling the processor 1100 to execute the pose prediction method provided in the embodiments of this application.

[0055] In the above description, those skilled in the art can design instructions according to the solutions disclosed in this application. How the instructions control the processor to operate is well known in the art, so it will not be described in detail here.

[0056] This application provides a pose prediction method applied to an electronic device as Figure 1 shown, and as Figure 2 shown, it includes the following steps S2100 and step S2200.

[0057] Step S2100, obtaining a historical pose sequence and a historical IMU sequence corresponding to the historical pose sequence.

[0058] Among them, the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence.

[0059] In this embodiment, the historical IMU data is collected by the IMU sensor at historical moments, and the historical IMU data includes three-axis acceleration data and three-axis angular velocity data corresponding to the historical moments.

[0060] Pose is usually represented by position data and attitude data. In one example, the pose can be represented as a seven-dimensional data: p_RS_R_x[m], p_RS_R_y[m], p_RS_R_z[m], q_RS_w[], q_RS_x[], q_RS_y[] and q_RS_z[]. Among them, p_RS_R_x[m], p_RS_R_y[m] and p_RS_R_z[m] represent three-dimensional position data, and q_RS_w[], q_RS_x[], q_RS_y[] and q_RS_z[] represent quaternion attitude data. On this basis, a historical pose data included in the historical pose sequence is composed of at least one of p_RS_R_x[m], p_RS_R_y[m], p_RS_R_z[m], q_RS_w[], q_RS_x[], q_RS_y[] and q_RS_z[] corresponding to the historical moment. That is to say, in this embodiment, the historical pose data may not represent a complete pose.

[0061] Moreover, the IMU sequence corresponding to the historical pose sequence refers to a sequence composed of historical IMU data that can compensate each piece of historical pose data in the historical pose sequence.

[0062] The historical pose sequence and the historical IMU sequence are used for a pose prediction model to predict a predicted pose sequence, where the predicted pose sequence includes multiple future pose data at future moments arranged in time sequence.

[0063] It should be noted that, in this embodiment, the number of historical pose data included in the historical pose sequence is the same as the number of historical IMU data included in the historical IMU sequence. For example, when the historical pose sequence is represented as [x1, x2, …, x m , the historical IMU sequence is represented as [IMU1, IMU2, …, IMU m .

[0064] Step S2200, generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and the pose prediction model.

[0065] Among them, the predicted pose sequence includes multiple future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, the pose compensation sub-model is used to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are the inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence.

[0066] In an embodiment of the present application, any one of the pose prediction sub-model, the pose compensation sub-model, and the fusion sub-model is one of a linear layer model, a Bi-GRU model, a GRU model, and a Bi-LSTM model.

[0067] Among them, for the Bi-GRU (Bidirectional Gated Recurrent Unit) model, the Bi-GRU model is a time series prediction method based on the gated recurrent unit (GRU). It combines a bidirectional model and a gating mechanism. The overall structure and the unit structure are the same as those of the GRU, so it can also effectively capture the temporal relationship in time series data. The overall structure of the Bi-GRU consists of two GRU networks in two directions. One network processes time series data from front to back, and the other network processes time series data from back to front. This bidirectional structure can capture both past and future information, thus more comprehensively modeling the temporal relationship in time series data. The GRU can be regarded as a simplified version of the LSTM, reducing the amount of computation but having a similar effect, which is beneficial for edge-side deployment.

[0068] Based on the above description of the pose prediction model, the structure of the pose prediction model provided by this application can be as follows Figure 3 shown. The traditional pose prediction model only generates a predicted pose sequence based on the historical pose sequence. Based on Figure 3 it can be known that the structure of the pose prediction model provided by this application is different from that of the traditional pose prediction model.

[0069] For the pose prediction model provided by this application, the pose prediction sub-model therein is used to generate a pose prediction intermediate quantity reflecting the predicted pose sequence according to the historical pose sequence. The pose compensation sub-model therein is used to generate a pose compensation quantity corresponding to each component in the aforementioned pose prediction intermediate quantity according to the historical IMU sequence corresponding to the historical pose sequence. The fusion sub-model therein is used to fuse the aforementioned pose prediction intermediate quantity reflecting the predicted pose sequence and the aforementioned pose compensation quantity to generate a predicted pose sequence. Specifically, each component in the pose prediction intermediate quantity is added to the corresponding component in the pose compensation quantity.

[0070] In an example, when the historical pose sequence is represented as [x1, x2, …, x m , and the historical IMU sequence is represented as [IMU1, IMU2, …, IMU m , the pose prediction intermediate quantity reflecting the predicted pose sequence is represented as [a1, a2, …, a k+1 , the pose compensation quantity is represented as [b1, b2, …, b k+1 , and the predicted pose sequence can be represented as [x m+i , x m+i+1 , …, x m+n , …, x m+i+k .

[0071] In this embodiment, the specific implementation of the above step S2200 is as follows: processing the historical pose sequence through the pose prediction sub-model to generate a pose prediction intermediate quantity reflecting the predicted pose sequence; processing the historical IMU sequence through the pose compensation sub-model to generate a pose compensation quantity corresponding to each component in the pose prediction intermediate quantity; through the fusion sub-model, in the way of late fusion (LateFusion), realizing the compensation of the corresponding component in the pose prediction intermediate quantity by the pose compensation quantity corresponding to each component in the pose prediction intermediate quantity to obtain an accurate predicted pose sequence. That is to say, the pose prediction method provided by this application can combine historical IMU data on the basis of historical pose data to generate a more accurate predicted pose sequence, that is, a new pose prediction method that can predict a more accurate predicted pose sequence is provided.

[0072] The present application provides a pose prediction method, which includes: obtaining a historical pose sequence and a corresponding historical IMU sequence, where the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence; generating a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and a pose prediction model; where the predicted pose sequence includes a plurality of future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, the pose compensation sub-model is used to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are the inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence. This method can realize combining historical IMU data on the basis of historical pose data to generate a more accurate predicted pose sequence, that is, it provides a new pose prediction method that can predict a more accurate predicted pose sequence.

[0073] In an embodiment of the present application, the pose prediction sub-model and the pose compensation sub-model are linear layer models. On this basis, the pose prediction method provided by the present application further includes a step of obtaining a pose prediction model before the above step S2200, and this step is implemented through the following steps S2210 to S2240.

[0074] Step S2210, obtain the prediction step number, the sample step number, the first sample set, and the second sample set.

[0075] Wherein, the first sample set includes a plurality of pose sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of pose sample data in sequence.

[0076] In this embodiment, the prediction step number is the number of future pose data that the pose prediction model to be trained can predict at one time, that is, the number of pose data included in the predicted pose sequence.

[0077] In an embodiment of the present application, for a single-step prediction scenario, the prediction step number can be set to 1, and for a multi-step prediction scenario, the prediction step number can be set to the corresponding number of steps for multiple steps.

[0078] The sample step number is the number of data used as samples in a training sample when training the pose prediction model to be trained. The sample step number is the same as the number of historical pose data included in the historical pose sequence in the above step S2100.

[0079] In an example, the first sample set can be represented as [x1, x2,..., x n , and the second sample set can be represented as [IMU1, IMU2,..., IMU n.

[0080] In an embodiment of the present application, the prediction steps, sample steps, first sample set, and second sample set in the above step S2210 can be obtained according to experimental acquisitions and input into the electronic device by the user. In this regard, the above step S2210 can also be implemented through the following step S2211 and step S2212.

[0081] Step S2211, display a setting input interface.

[0082] Step S2212, receive the prediction steps, sample steps, first sample set, and second sample set input through the setting input interface.

[0083] In this embodiment, in order to allow the user to input the prediction steps, sample steps, first sample set, and second sample set that meet their own needs, the electronic device provides a setting input interface. The user can input the prediction steps, sample steps, first sample set, and second sample set that meet their own needs through this setting input interface.

[0084] In one example, the setting input interface may specifically include a first interface, a second interface, a third interface, and a fourth interface. Among them, the first interface is used for the user to input the prediction steps that meet their own needs, the second interface is used for the user to input the sample steps that meet their own needs, the third interface is used for the user to input the first sample set that meet their own needs, and the fourth interface is used for the user to input the second sample set that meet their own needs.

[0085] Step S2220, set the pose prediction model to be trained.

[0086] Among them, setting the pose prediction model to be trained includes the following steps S2221 to S2223.

[0087] Step S2221, set the number of neurons in the input layer of the pose prediction sub-model and the input layer of the pose compensation sub-model to the sample steps.

[0088] In this embodiment, the input layer is the first layer of the model and is used to input the samples in the training samples during the model training process. The number of neurons in the input layer corresponds to the number of features included in the samples in a training sample. Based on this, in this embodiment, the number of neurons in the input layer of the pose prediction sub-model and the number of neurons in the input layer of the pose compensation sub-model are both the same as the sample steps.

[0089] Step S2222, set the number of neurons in the last hidden layer of the pose prediction sub-model and the number of neurons in the last hidden layer of the pose compensation sub-model to be the same.

[0090] In this embodiment, the pose prediction sub-model is used to generate pose prediction intermediate quantities reflecting the predicted pose sequence, and the pose compensation sub-model is used to generate pose compensation quantities corresponding to each component in the aforementioned pose prediction intermediate quantities, and one neuron outputs one feature. Therefore, in order to ensure that the pose compensation sub-model can generate pose compensation quantities corresponding to each component of the pose prediction intermediate quantities generated by the pose prediction sub-model, it is necessary to make the number of neurons in the last hidden layer of the pose prediction sub-model the same as the number of neurons in the last hidden layer of the pose compensation sub-model.

[0091] It should be noted that in this application, the number of hidden layers of the pose prediction sub-model and the pose compensation sub-model, as well as the number of neurons in non-last hidden layers, are not limited.

[0092] Step S2223: Set the number of neurons in the output layer of the fusion sub-model to the number of prediction steps, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model.

[0093] In this embodiment, the output layer is the last layer of the model and is used to output the pose predicted by the model during the model training process. The number of neurons in the output layer is the number of features included in the label in a training sample, that is, the number of future pose data that the pose prediction model can predict at one time. Based on this, in this embodiment, the number of neurons in the output layer is set to the number of prediction steps. For example, when the number of prediction steps is k + 1, the number of neurons in the output layer is k + 1.

[0094] Moreover, in order to completely receive the pose prediction intermediate quantities output by the pose prediction sub-model and the pose compensation quantities of the pose compensation sub-model, the neurons in the last hidden layer of the pose prediction sub-model and the neurons in the last hidden layer of the pose compensation sub-model are fully connected to the neurons in the input layer of the fusion sub-model. That is, the input layer of the fusion sub-model is a fully connected layer.

[0095] In an embodiment of this application, the fusion sub-model is a fully connected layer. Based on this, the input layer of the fusion sub-model is also the output layer, and the number of neurons in the fusion sub-model, the number of neurons in the last hidden layer of the pose prediction sub-model, and the number of neurons in the last hidden layer of the pose compensation sub-model are all the number of prediction steps.

[0096] Based on the above steps S2221 to S2223, the structure customization of the pose prediction model to be trained can be realized. Further, the training of the model is realized through the following steps S2230 and S2240.

[0097] Step S2230: Generate a plurality of training samples according to the number of prediction steps, the number of sample steps, the first sample set, and the second sample set.

[0098] Among them, a training sample includes pose sample data of the units digit of consecutive sample steps as samples, IMU sample data of consecutive sample steps as samples, and pose label data of consecutive predicted steps as labels.

[0099] In this embodiment, the specific implementation of the above step S2230 is as follows: select pose sample data of the units digit of consecutive sample steps from the first sample set, and obtain IMU sample data of consecutive sample steps corresponding to the aforementioned pose sample data of the units digit of consecutive sample steps from the second sample set to form the samples in the training sample; select pose sample data of consecutive predicted steps from the first sample set as the labels in the training sample, where the occurrence time of the pose sample data in the labels is after the occurrence time of the pose sample data in the samples.

[0100] In an example, when the pose sample data in the samples of the training sample is [x1, x2, x3, x4, x5, x6, x7], and the prediction requirement is to predict the pose data of the 10th, 11th, and 12th steps, the pose sample data in the labels of the training sample is [x 10 , x 11 , x 12 .

[0101] Step S2240, train the pose prediction model to be trained according to multiple training samples to obtain a trained pose prediction model.

[0102] In this embodiment, the pose sample data of the units digit of consecutive sample steps in the samples of the training sample is input into the pose prediction sub-model to be trained, and the IMU sample data of consecutive sample steps in the samples of the training sample is input into the fusion sub-model to be trained. Through traditional model iterative training, the weights and biases of each neuron in the pose prediction sub-model, pose compensation sub-model, and fusion sub-model are adjusted to implement the above step S2240.

[0103] In an embodiment of the present application, when the pose data is represented by three-dimensional position data, the above step S2100 is specifically implemented through the following step S2110 and step S2120.

[0104] Step S2110, extract a pose data component of the same dimension from the historical position data to obtain a historical pose sequence.

[0105] In this embodiment, when the pose data is characterized by the three-dimensional position data represented by p_RS_R_x[m], p_RS_R_y[m], and p_RS_R_z[m], one pose data component in the same dimension can specifically be any one of p_RS_R_x[m], p_RS_R_y[m], and p_RS_R_z[m]. That is, the historical pose sequence is composed of p_RS_R_x[m] at multiple historical moments, or composed of p_RS_R_y[m] at multiple historical moments, or composed of p_RS_R_z[m] at multiple historical moments.

[0106] Step S2120: Extract the IMU components that match the pose data components from the historical IMU data to obtain the historical IMU sequence.

[0107] When the historical pose sequence is composed of p_RS_R_x[m] at multiple historical moments, the IMU component that matches the pose data component is the x-axis acceleration at the corresponding historical moment. When the historical pose sequence is composed of p_RS_R_y[m] at multiple historical moments, the IMU component that matches the pose data component is the y-axis acceleration at the corresponding historical moment. When the historical pose sequence is composed of p_RS_R_z[m] at multiple historical moments, the IMU component that matches the pose data component is the z-axis acceleration at the corresponding historical moment.

[0108] In this embodiment, when characterizing the historical pose sequence through one position data component in the same dimension, the data processing amount of the pose prediction model is small, and the ability of the pose prediction model to learn non-linear relationships can be improved.

[0109] In an embodiment of the present application, when the pose data is characterized by quaternion attitude data, the above step S2100 is specifically implemented through the following steps S2130 to S2150.

[0110] Step S2130: Convert the historical quaternion attitude data into historical three-axis Euler angle data.

[0111] In this embodiment, according to the conversion formula between quaternion attitude and three-axis Euler angle, the historical quaternion attitude data is converted into historical three-axis Euler angle data. It can be understood that the historical three-axis Euler angle data is used to characterize the rotation angles around the X-axis, Y-axis, and Z-axis at historical moments respectively.

[0112] Step S2140: Extract the one-axis Euler angle data in the same dimension from the historical three-axis Euler angle data to obtain the historical pose sequence.

[0113] In this embodiment, the unary Euler angle data in the same dimension is one of the rotation angles about the X-axis, Y-axis, and Z-axis. That is, the historical pose sequence is composed of the rotation angles about the X-axis at multiple historical moments, or is composed of the rotation angles about the Y-axis at multiple historical moments, or is composed of the rotation angles about the Z-axis at multiple historical moments.

[0114] Step S2150: Extract the IMU components that match the unary Euler angle data from the historical IMU data to obtain a historical IMU sequence.

[0115] In this embodiment, when the historical pose sequence is composed of the rotation angles about the X-axis at multiple historical moments, the IMU component that matches the unary Euler angle data is specifically the angular velocity about the X-axis. When the historical pose sequence is composed of the rotation angles about the Y-axis at multiple historical moments, the IMU component that matches the unary Euler angle data is specifically the angular velocity about the Y-axis. When the historical pose sequence is composed of the rotation angles about the Z-axis at multiple historical moments, the IMU component that matches the unary Euler angle data is specifically the angular velocity about the Z-axis.

[0116] In this embodiment, when characterizing the historical pose sequence by the unary Euler angle data in the same dimension, the data processing amount of the pose prediction model is small, and the ability of the pose prediction model to learn non-linear relationships can be improved.

[0117] The present application also provides a pose prediction device 400, as Figure 4 shown, including:

[0118] A first acquisition module 410, configured to acquire a historical pose sequence and a historical IMU sequence corresponding to the historical pose sequence, where the historical pose sequence includes a plurality of historical pose data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence;

[0119] A prediction module 420, configured to generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence, and a pose prediction model;

[0120] Wherein, the predicted pose sequence includes a plurality of future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model, and a fusion sub-model. The pose prediction sub-model is configured to receive the historical pose sequence, the pose compensation sub-model is configured to receive the historical IMU sequence, the outputs of the pose prediction sub-model and the pose compensation sub-model are inputs of the fusion sub-model, and the fusion sub-model is configured to output the predicted pose sequence.

[0121] In one embodiment of the present application, any one of the pose prediction sub-model, the pose compensation sub-model, and the fusion sub-model is one of a linear layer model, a Bi-GRU model, a GRU model, and a Bi-LSTM model.

[0122] In one embodiment of the present application, the pose prediction device 400 provided by the present application further includes:

[0123] A second acquisition module, configured to acquire a prediction step number, a sample step number, a first sample set, and a second sample set, where the first sample set includes a plurality of pose sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of pose sample data in sequence;

[0124] Set a pose prediction model to be trained, and setting the pose prediction model to be trained includes: setting the number of neurons in the input layer of the pose prediction sub-model and the input layer of the pose compensation sub-model to be the sample step number; setting the number of neurons in the last hidden layer of the pose prediction sub-model to be the same as the number of neurons in the last hidden layer of the pose compensation sub-model; setting the number of neurons in the output layer of the fusion sub-model to be the prediction step number, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model;

[0125] Generate a plurality of training samples according to the prediction step number, the sample step number, the first sample set, and the second training sample set. One training sample includes consecutive sample step number pose sample data as samples, consecutive sample step number IMU sample data as samples, and consecutive prediction step number pose label data as labels;

[0126] Train the pose prediction model to be trained according to the plurality of training samples to obtain a trained pose prediction model.

[0127] In one embodiment of the present application, when the pose data is represented by three-dimensional position data, the first acquisition module 410 is specifically configured to: extract a position data component of the same dimension from the historical position data to obtain a historical pose sequence;

[0128] Extract an IMU component matching the position data component from the historical IMU data to obtain a historical IMU sequence.

[0129] In one embodiment of the present application, when the pose data is represented by quaternion attitude data, the first acquisition module 410 is specifically configured to:

[0130] Convert the historical quaternion attitude data into historical three-dimensional Euler angle data;

[0131] Extract the unary Euler angle data of the same dimension from the historical ternary Euler angle data to obtain a historical pose sequence;

[0132] Extract the IMU components matching the unary Euler angle data from the historical IMU data to obtain a historical IMU sequence.

[0133] In an embodiment of the present application, the second acquisition module is specifically configured to display a set input interface;

[0134] Receive the predicted number of steps, the sample number of steps, the first sample set, and the second sample set input through the set input interface.

[0135] The present application also provides an electronic device, which includes any one of the pose prediction devices 400 provided in the above device embodiment.

[0136] Or, as Figure 5 shown, the electronic device 500 includes a memory 510 and a processor 520. The memory 510 is used to store computer instructions, and the processor 520 is used to call the computer instructions from the memory 510 to execute any one of the pose prediction methods provided in the above method embodiment.

[0137] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the pose prediction methods provided in the above method embodiment.

[0138] The present application may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present application.

[0139] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0140] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0141] The computer program instructions for performing the operations of the present application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present application.

[0142] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0143] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that when these instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner. Thus, the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0144] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0145] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As is well known to those skilled in the art, implementations by hardware, by software, and by the combination of software and hardware are equivalent.

[0146] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present application is defined by the appended claims.

Claims

1. A posture prediction method, characterized in that: include: Acquire a historical posture sequence and a historical IMU sequence corresponding to the historical posture sequence, wherein the historical posture sequence includes a plurality of historical posture data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence; Generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence and the pose prediction model; Among them, the predicted pose sequence includes multiple future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, and the pose compensation sub-model is used to receive the historical IMU sequence. The outputs of the pose prediction sub-model and the pose compensation sub-model are the inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence.

2. The method according to claim 1, characterized in that Any one of the posture prediction sub-model, the posture compensation sub-model and the fusion sub-model is one of a linear layer model, a Bi-GRU model, a GRU model and a Bi-LSTM model.

3. The method according to claim 1, characterized in that The posture prediction submodel and the posture compensation submodel are linear layer models. Before generating a predicted posture sequence according to the historical posture sequence, the historical IMU sequence and the posture prediction model, the method further includes: Acquire the predicted number of steps, the number of sample steps, a first sample set and a second sample set, wherein the first sample set includes a plurality of pose sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of pose sample data in sequence; Setting a posture prediction model to be trained, wherein the setting of the posture prediction model to be trained includes: setting the number of neurons in the input layer of the posture prediction sub-model and the input layer of the posture compensation sub-model to the number of sample steps; setting the number of neurons in the last hidden layer of the posture prediction sub-model to be the same as the number of neurons in the last hidden layer of the posture compensation sub-model; setting the number of neurons in the output layer of the fusion sub-model to be the number of prediction steps, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model; Generate a plurality of training samples according to the predicted number of steps, the sample number of steps, the first sample set and the second training sample set, wherein one training sample includes pose sample data of the consecutive sample steps as samples, IMU sample data of the consecutive sample steps as samples, and pose label data of the consecutive predicted number of steps as labels; The posture prediction model to be trained is trained according to the multiple training samples to obtain a trained posture prediction model.

4. The method according to claim 1, characterized in that In the case where the pose data is represented by three-dimensional position data, the acquiring of the historical pose sequence and the historical IMU sequence corresponding to the historical pose sequence includes: Extract a position data component of the same dimension from the historical position data to obtain a historical posture sequence; An IMU component matching the position data component is extracted from the historical IMU data to obtain a historical IMU sequence.

5. The method according to claim 1, characterized in that In the case where the pose data is represented by quaternary pose data, the acquiring of the historical pose sequence and the historical IMU sequence corresponding to the historical pose sequence comprises: Convert historical quaternion posture data into historical ternary Euler angle data; Extracting one-dimensional Euler angle data of the same dimension from the historical three-dimensional Euler angle data to obtain a historical posture sequence; An IMU component matching the univariate Euler angle data is extracted from the historical IMU data to obtain a historical IMU sequence.

6. The method according to claim 3, characterized in that The obtaining of the prediction step number, the sample step number, the first sample set and the second sample set includes: Display setting input interface; The prediction step number, sample step number, first sample set and second sample set inputted from the setting input interface are received.

7. A posture prediction device, characterized in that: include: A first acquisition module is used to acquire a historical posture sequence and a historical IMU sequence corresponding to the historical posture sequence, wherein the historical posture sequence includes a plurality of historical posture data arranged in time sequence, and the historical IMU data includes a plurality of historical IMU data arranged in time sequence; A prediction module, used to generate a predicted pose sequence according to the historical pose sequence, the historical IMU sequence and the pose prediction model; Among them, the predicted pose sequence includes multiple future pose data arranged in time sequence, and the pose prediction model includes: a pose prediction sub-model, a pose compensation sub-model and a fusion sub-model. The pose prediction sub-model is used to receive the historical pose sequence, and the pose compensation sub-model is used to receive the historical IMU sequence. The outputs of the pose prediction sub-model and the pose compensation sub-model are the inputs of the fusion sub-model, and the fusion sub-model is used to output the predicted pose sequence.

8. The device according to claim 7, characterized in that The device also includes: A second acquisition module is used to acquire the predicted number of steps, the number of sample steps, a first sample set and a second sample set, wherein the first sample set includes a plurality of posture sample data arranged in time sequence, and the second sample set includes sample IMU data corresponding to the plurality of posture sample data in sequence; Setting a posture prediction model to be trained, wherein the setting of the posture prediction model to be trained includes: setting the number of neurons in the input layer of the posture prediction sub-model and the input layer of the posture compensation sub-model to the number of sample steps; setting the number of neurons in the last hidden layer of the posture prediction sub-model to be the same as the number of neurons in the last hidden layer of the posture compensation sub-model; setting the number of neurons in the output layer of the fusion sub-model to be the number of prediction steps, and any neuron in the last hidden layer is fully connected to the neurons in the input layer of the fusion sub-model; Generate a plurality of training samples according to the predicted number of steps, the sample number of steps, the first sample set and the second training sample set, wherein one training sample includes pose sample data of the consecutive sample steps as samples, IMU sample data of the consecutive sample steps as samples, and pose label data of the consecutive predicted number of steps as labels; The posture prediction model to be trained is trained according to the multiple training samples to obtain a trained posture prediction model.

9. An electronic device, characterized in that: The electronic device comprises the device according to claim 7 or 8; Alternatively, the electronic device includes a memory and a processor, the memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method as described in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Digital exhibition hall display method and equipment based on AR (Augmented Reality) technology

    CN121304992A