Robot state prediction model training method and device, medium and electronic equipment
By iterating the training of the state prediction model in cyclically and performing coordinate transformation, the accuracy problem of dynamic modeling of wheeled robots in the prior art under any coordinate system is solved, and the path planning and obstacle avoidance accuracy of the robot in a dynamic environment is improved.
Patent Information
- Application Number
- CN202311708603.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to accurately model the dynamics of wheeled robots under any coordinate system, resulting in low path planning and obstacle avoidance accuracy in dynamic environments.
By obtaining the robot's data set, including the actual instruction sequence and the actual state sequence, the state prediction model is trained using a loop iterative training method. The inputs and outputs related to state in the model are data under the local coordinate system. Through data transformation, the intermediate prediction state under the local coordinate system is converted into the target prediction state under the world coordinate system, and the state prediction model is updated to conform to physical symmetry.
The prediction accuracy of the state prediction model is improved, so that it is not affected by the translation and rotation of the coordinate system, and the robot's path planning and obstacle avoidance capabilities in dynamic environments are enhanced.
Smart Images

Figure CN120145796A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of robotics, and in particular, to a method, apparatus, medium, and electronic device for training a robot state prediction model. Background Art
[0002] Robots are being used more and more widely and have become important devices in various fields, such as wheeled mobile robots. Due to their versatility and flexibility in dynamic environments, wheeled robots play an important role in fields such as industry and agriculture. Wheeled robots use on-body perception data and a reference map for path planning and obstacle avoidance, and control drive wheel motors to achieve movements such as forward and turning.
[0003] Currently, usually by fitting data, a model that accurately represents the data is determined, and the obtained model is applied to the pose prediction scenario of the robot. Therefore, accurately modeling the dynamics of the wheeled robot is crucial. Summary of the Invention
[0004] This summary of the invention is provided to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. This summary of the invention is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides a method for training a robot state prediction model, including:
[0006] Obtain a data set of the robot, where the data set includes an actual instruction sequence and an actual state sequence;
[0007] Perform iterative training on the state prediction model according to the data set, and when the training is completed, output the trained state prediction model. In the process of one iterative training, perform the following operations:
[0008] Select an actual state subsequence from the actual state sequence, where each actual state in the actual state subsequence is data in the world coordinate system;
[0009] Process a target input using the state prediction model to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system. The target input is at least determined by the actual instruction sequence, and both the input and output related to the state in the state prediction model are data in the local coordinate system;
[0010] Update the state prediction model according to the prediction state subsequence composed of all the target prediction states and the actual state subsequence.
[0011] In a second aspect, the present disclosure provides a training device for a robot state prediction model, including:
[0012] An acquisition module, configured to acquire a data set of the robot, where the data set includes an actual instruction sequence and an actual state sequence;
[0013] A training module, configured to perform iterative training on the state prediction model according to the data set, and output the trained state prediction model when the training is completed. During one iterative training process, the following operations are performed:
[0014] Select an actual state subsequence from the actual state sequence, where each actual state in the actual state subsequence is data in the world coordinate system;
[0015] Process a target input by using the state prediction model to obtain intermediate prediction states corresponding to the respective actual states in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system. The target input is determined at least by the actual instruction sequence, and both the input and output related to the state in the state prediction model are data in the local coordinate system;
[0016] Update the state prediction model according to a prediction state subsequence composed of all the target prediction states and the actual state subsequence.
[0017] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect above are implemented.
[0018] In a fourth aspect, the present disclosure provides an electronic device, including:
[0019] A storage device, on which a computer program is stored;
[0020] A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect above.
[0021] Through the above technical solution, both the state-related inputs and outputs in the state prediction model are data in the local coordinate system, enabling the model to only focus on local motion. The intermediate prediction state in the local coordinate system output by the state prediction model is subjected to data transformation to obtain the target prediction state in the world coordinate system. On this basis, fitting of the state prediction model is performed based on the prediction state subsequence composed of the target prediction states in the world coordinate system and the actual state subsequence in the world coordinate system, ensuring that the fitted state prediction model can conform to physical symmetry, such that the prediction accuracy of the state prediction model is not affected by coordinate translation, rotation, etc., thereby improving the prediction accuracy of the state prediction model.
[0022] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In combination with the accompanying drawings and with reference to the following specific implementation manners, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings:
[0024] Figure 1 FIG. [FIG. NUMBER] is a schematic diagram of a process of cyclic iterative training shown according to an exemplary embodiment of the present disclosure.
[0025] Figure 2 FIG. [FIG. NUMBER] is a schematic diagram of a training process of a state prediction model shown according to an exemplary embodiment of the present disclosure.
[0026] Figure 3 FIG. [FIG. NUMBER] is a schematic diagram of a process of collecting each data in a data collection set shown according to an exemplary embodiment of the present disclosure.
[0027] Figure 4 FIG. [FIG. NUMBER] is another schematic diagram of a process of a training method of a robot state prediction model shown according to an exemplary embodiment of the present disclosure.
[0028] Figure 5 FIG. [FIG. NUMBER] is a block diagram of a training device of a robot state prediction model shown according to an exemplary embodiment of the present disclosure.
[0029] Figure 6 FIG. [FIG. NUMBER] is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0031] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0032] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0033] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0036] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0037] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0038] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way to send a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0039] It can be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation of the present disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0040] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0041] In related ways, the robot postures recorded in the data set are based on global coordinates, with a fixed coordinate origin and orientation. The dynamic model directly fitted based on such data does not conform to translational and rotational symmetries and cannot be applied to new scenarios under any coordinate system. In addition, the data set in the related technology is collected in a virtual environment and cannot reflect the complex dynamics in the real world, such as the slipping or sliding of the robot. Usually, the model fitted based on this data set needs to be manually adjusted to match the real world.
[0042] In view of this, the embodiments of the present disclosure provide a method, device, medium and electronic device for training a state prediction model.
[0043] The embodiments of the present disclosure provide a method for training a robot state prediction model. The method for training the robot state prediction model can be executed by an electronic device, specifically, it can be executed by a device for training a robot state prediction model, and the device can be implemented in a software and / or hardware manner. The training method can include: obtaining a data set of the robot, where the data set includes an actual instruction sequence and an actual state sequence; performing iterative training on the state prediction model according to the data set, and when the training is completed, outputting the trained state prediction model.
[0044] As an example, the robot in this embodiment can be a wheeled robot. The wheeled robot can use its own perception data and a reference map for path planning and obstacle avoidance, and control the drive wheel motors to achieve movements such as forward and turning.
[0045] Among them, the data set is data sampled in a real environment, that is, an actual application scenario, and it can reflect the interaction between the behavior of the robot and the environment in the actual application scenario.
[0046] Among them, the actual instruction sequence may include a plurality of actual instructions continuously executed by the robot and sorted in chronological order, and a time stamp corresponding to each actual instruction, which can be understood as the time stamp when each actual instruction starts to be executed. It can be understood that the time stamp when each actual instruction starts to be executed can be understood as the moment when the robot navigation module or controller issues the actual instruction, or the time stamp when each actual instruction starts to be executed can also be understood as the moment of the first action triggered when the robot starts to execute the actual instruction. As an example, the actual instruction can be a linear velocity instruction and an angular velocity instruction.
[0047] As can be seen from the above, each element in the actual instruction sequence is sorted in chronological order, and each element may include the corresponding actual instruction and the time stamp corresponding to the actual instruction. Further, the actual instruction may include a linear velocity instruction and an angular velocity instruction.
[0048] Among them, the actual state sequence may include a plurality of actual states of the robot continuously observed and sorted in the order of observation time, and a time stamp corresponding to each actual state, which can be understood as the time stamp when the robot starts to be in each actual state. As an example, the actual state can be the robot position and the robot orientation, and the actual state can also be the robot position, the robot orientation and the robot speed. Similar to the actual instruction sequence, each element in the actual state sequence is sorted in chronological order, and each element may include the corresponding actual state and the time stamp corresponding to the actual state. Further, the actual state may include the robot position and the robot orientation.
[0049] As an example, the data set can be represented by the following formula:
[0050]
[0051] Among them, D represents the data set, represents the actual instruction sequence, {tj, qj}n represents the actual state sequence, m represents and the number of, used to describe the time stamp of the i-th actual instruction in the actual instruction sequence, represents the i-th linear velocity instruction, represents the i-th angular velocity instruction, t j represents the time stamp of the i-th actual state in the actual state sequence, q j represents the i-th actual state in the actual state sequence, n represents q j and t j the number of, and i represents the arrangement positions in the actual instruction sequence and the actual state sequence respectively.
[0052] In addition, for the implementation of obtaining the dataset of the robot, reference may be made to the following related embodiments, which will not be elaborated herein.
[0053] In this embodiment, the state prediction model is an MLP (Multilayer Perceptron), and the MLP can also be referred to as an ANN (Artificial Neural Network).
[0054] As an example, the training being completed can be that the number of loop iterations reaches a preset number, the difference between the model input and output is less than a preset threshold, and so on.
[0055] Figure 1 is a schematic diagram of a process of loop iteration training shown according to an exemplary embodiment of the present disclosure. Refer to Figure 1 , the following steps are performed in one process of loop iteration training:
[0056] Step 110, select an actual state subsequence from the actual state sequence, and each actual state in the actual state subsequence is data in the world coordinate system.
[0057] Step 120, process the target input using the state prediction model to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system. The target input is determined at least by the actual instruction sequence, and both the input and output related to the state in the state prediction model are data in the local coordinate system;
[0058] Step 130, update the state prediction model according to the prediction state subsequence and the actual state subsequence composed of all target prediction states.
[0059] It should be noted that the sequence length of the actual state subsequence is less than the sequence length of the actual state sequence. As an example, for example, the actual state sequence includes [a1, a2, a3, a4, a5, a6, a7], where a1, a2, a3, a4, a5, a6, and a7 represent each actual state in the actual state sequence. Based on this actual state sequence, the actual state subsequence that can be sampled can be [a1, a2, a3] or [a4, a5, a6, a7], and so on. As an example, the actual state subsequence can also include the timestamps corresponding to each actual state.
[0060] Among them, the sequence lengths of the actual state subsequences selected during the cyclic iterative training process can be different. As an example, after the cyclic iterative training based on all the actual state subsequences with a sequence length of K selected from the dataset is completed, the sequence length K is updated in an increasing sequence length manner, and based on the updated sequence length K, actual state subsequences are continuously selected from the dataset for cyclic iterative training until the training is completed. As other examples, the sequence length K selected during the cyclic iterative training process can also be a fixed value.
[0061] It should be noted that each actual state in the actual state subsequence is data in the world coordinate system.
[0062] Among them, during one cyclic iterative training, a first loss function can be constructed based on the predicted state subsequence and the corresponding actual state subsequence, and the state prediction model can be updated based on this first loss function.
[0063] Among them, during one cyclic iterative training, multiple selected actual state subsequences can be included. Correspondingly, the predicted state subsequences corresponding to each actual state subsequence can be determined by the above method. On this basis, a second loss function can be constructed based on multiple predicted state subsequences and the actual state subsequences corresponding to each predicted state subsequence, and the state prediction model can be updated based on this second loss function.
[0064] As an example, the above loss function can be a loss function reflecting the mean square error between the actual state subsequence and the predicted state subsequence.
[0065] It should be noted that the sequence lengths of the predicted state subsequence and the actual state subsequence are equal, and the actual state and the target predicted state at the same sorting position in the predicted state subsequence and the actual state subsequence correspond to each other. On this basis, it can be understood that that is, each actual state at each sorting position also corresponds to an intermediate predicted state.
[0066] It should be noted that both the input and output related to the state in the state prediction model are data in the local coordinate system. The input related to the state here refers to the parameters used to characterize the state of the robot, such as the above-mentioned angular velocity and linear velocity, etc.
[0067] In the above manner, both the state-related inputs and outputs in the state prediction model are data in the local coordinate system, enabling the model to only focus on local motion. The intermediate prediction state in the local coordinate system output by the state prediction model is subjected to data transformation to obtain the target prediction state in the world coordinate system. On this basis, fitting of the state prediction model is performed based on the prediction state subsequence composed of the target prediction states in the world coordinate system and the actual state subsequence in the world coordinate system, ensuring that the fitted state prediction model can conform to physical symmetry, such that the prediction accuracy of the state prediction model is not affected by coordinate translation, rotation, etc., thereby improving the prediction accuracy of the state prediction model.
[0068] In a possible way, the steps of using the state prediction model to process the target input, obtaining the intermediate prediction states corresponding to the respective actual states in the actual state subsequence, and performing data transformation on the intermediate prediction states to obtain the target prediction state in the world coordinate system can be implemented in the following manner: For the intermediate prediction states corresponding to the respective actual states in the predicted actual state subsequence, initialize the coordinate offset vector and the target input; use the state prediction model to process the target input to obtain the intermediate prediction state corresponding to the actual state; perform data transformation on the intermediate prediction state according to the coordinate offset vector to obtain the target prediction state in the world coordinate system.
[0069] It should be noted that when dealing with the intermediate prediction states corresponding to the respective actual states in the predicted actual state subsequence, it is necessary to initialize the coordinate offset vector and the target input.
[0070] Among them, the coordinate offset vector is used to record the transformation from the current coordinate system (i.e., the local coordinate system) to the original data coordinate system (i.e., the world coordinate system). Before the start of prediction, the coordinate offset vector is initialized to a value such that the position and orientation of the state-related parameters used to predict the intermediate prediction state are both at the origin.
[0071] Among them, according to the type of the actual state corresponding to the predicted intermediate prediction state, the coordinate offset vector is obtained according to the initialization method of the corresponding type. Here, the type of the actual state can be the position of the actual state in the actual state subsequence.
[0072] As an example, the coordinate offset vector can be initialized in the following manner: When predicting the intermediate prediction state corresponding to the first actual state in the predicted actual state subsequence, the coordinate offset vector is initialized according to the previous actual state located in the actual state subsequence before the first actual state in the actual state sequence. The above-initialized coordinate offset vector can be represented by the following formula:
[0073]
[0074] Among them, qk is the initial state, which is the previous actual state of the first actual state in the actual state subsequence in the actual state sequence, Δq 0 is the coordinate offset vector obtained by initialization. Δq is the coordinate offset vector of the coordinate system where the actual state is located (i.e., the world coordinate system) relative to the coordinate system where the intermediate predicted state is located (i.e., the local coordinate system), that is, Δq and Δq 0 are parameters with the same physical meaning. Δx is the translation transformation offset of the coordinate system in the x-axis direction, Δy is the translation transformation offset of the coordinate system in the y-axis direction, and Δθ is the rotation transformation offset. Here, T represents the transpose symbol. The data transformation can be only for translation without rotation operation. In this case, Δθ is always 0.
[0075] As an example, when predicting the intermediate predicted state corresponding to a non-first actual state in the actual state subsequence, according to the previously initialized coordinate offset vector and the first intermediate predicted state predicted by the state prediction model for the actual state subsequence, the coordinate offset vector is initialized. The above initialized coordinate offset vector can be represented by the following formula:
[0076]
[0077] where, Δq i+1 is the updated coordinate offset vector, Δq i is the current coordinate offset vector, that is, the previously initialized coordinate offset vector, and R is the transformation matrix, is the first intermediate predicted state predicted by the state prediction model for the actual state subsequence. R is the transformation matrix for performing data transformation on the intermediate predicted state to obtain the target predicted state in the world coordinate system. R can be represented by the following formula:
[0078]
[0079] where the target input is the input of the state prediction model.
[0080] As an example, the target input can be determined only by the actual instruction sequence. In this case, as an example, the target input can be determined according to the target timestamp, the first preset duration, and the actual instruction sequence. The target timestamp is the timestamp carried by the actual state corresponding to the intermediate predicted state currently predicted by the state prediction model.
[0081] It should be noted that the target input in the above example includes target actual instructions within the first preset duration before the target timestamp determined from the actual instruction sequence, as well as the timestamps carried by each target actual instruction. All target actual instructions and the timestamps carried by each target actual instruction constitute a historical instruction sequence, and this historical instruction sequence is the target input. Among them, the target actual instructions can include instructions of types angular velocity and linear velocity, and the first preset duration can be set according to the actual situation, which will not be elaborated in this embodiment. The minimum sequence length of the historical instruction sequence can be 1, and the historical instruction sequence can be represented by the following formula:
[0082]
[0083] where C i historical instruction sequence, represents the actual instruction sequence, is the universal quantifier, representing any, t k+i represents the timestamp corresponding to the actual state corresponding to the currently predicted target prediction state, that is, t k+i represents the target timestamp, where T here represents the time span of the historical instruction sequence, that is, T represents the first preset duration.
[0084] As another example, on the basis that the target input includes what is determined by the actual instruction sequence, the target input can also be data determined by the actual instruction sequence. For example, according to the target timestamp, the first preset duration, and the actual instruction sequence, the first input is determined. On this basis, according to the first actual state in the actual state subsequence or the target prediction state corresponding to the intermediate prediction state already predicted by the state prediction model for the actual state subsequence, the second input is determined, and the target input is determined according to the first input and the second input.
[0085] It should be noted that the determination method of the first input can refer to the relevant embodiments above, which will not be elaborated in this embodiment.
[0086] It should be noted that for predicting the intermediate prediction state corresponding to the first actual state in the actual state subsequence, the second input is the actual state in the actual state sequence that is the previous actual state of the first actual state in the actual state subsequence; for predicting the intermediate prediction state corresponding to a non-first actual state in the actual state subsequence, the second input can be understood as a historical motion sequence, which includes at least one historical state, and the number of at least one historical state does not exceed the first preset quantity at most. At least one historical state can be determined based on the target prediction state corresponding to the intermediate prediction state predicted by the state prediction model for the actual state subsequence. For example, the target prediction states corresponding to the actual states within the second preset time period before the target time stamp are selected as at least one historical state. For example, the target prediction states corresponding to the first N actual states before the target time stamp are selected as at least one historical state. The explanation of the target time stamp can refer to the above related embodiments. As an example, the historical motion sequence can be represented by the following formula:
[0087]
[0088] where Q i represents the historical motion sequence, represents the target prediction state corresponding to the intermediate prediction state predicted by the state prediction model for each actual state in the actual state subsequence, and H is the sequence length of the historical motion sequence.
[0089] It should be noted that various state-related parameters in the historical motion sequence need to be transformed into the local coordinate system before being input into the prediction state model.
[0090] Figure 2 is a schematic diagram of the training process of a state prediction model shown according to an exemplary embodiment of the present disclosure. Referring to Figure 2 , the collected data set includes X and Y, where X can represent the actual instruction sequence and Y can represent the actual state sequence. The representations of X and Y can refer to the above related embodiments. Combining the above content, further explanation and illustration are made for Figure 2 .
[0091] For the actual state sequence, an actual state subsequence can be selected, and using the state prediction model, the prediction subsequence corresponding to the actual state subsequence is determined. First, when predicting each target prediction state in the prediction subsequence, it is necessary to initialize the coordinate offset vector and the target input.
[0092] For example, when predicting the first target prediction state in the prediction subsequence, that is, Figure 2 in the first-step prediction shown, first determine the initial state q k, this initial state can be understood as the previous actual state of the first actual state in the actual state subsequence within the selected actual state sequence, that is, it simultaneously represents the k-th actual state in the actual state sequence. Initialize the initial state to obtain a coordinate offset vector.
[0093] Next, initialize the target input. In this embodiment, the target input includes a historical instruction sequence (i.e., the first input) and a historical motion sequence (i.e., the second input). Among them, the historical instruction sequence can be represented by the following formula:
[0094]
[0095] Among them, C i historical instruction sequence, represents the actual instruction sequence, is a universal quantifier, indicating arbitrary, t k+i represents the timestamp corresponding to the actual state corresponding to the currently predicted target prediction state, that is, t k+i represents the target timestamp, T represents the time span of the historical motion sequence, that is, T represents the first preset duration;
[0096] Among them, the historical instruction sequence can be represented by the following formula:
[0097]
[0098] Among them, Q i represents the historical motion sequence, represents the target prediction state corresponding to the intermediate prediction state predicted by the state prediction model for each actual state in the actual state subsequence, H is the sequence length of the historical motion sequence. It should be noted that for the first-step prediction, the initial state is used to replace the historical motion sequence and input into the state prediction model. For other predictions except the first-step prediction, the historical motion sequence is determined by the above formula.
[0099] Next, for each step of prediction, the state prediction model receives the target input and calculates the next local motion state, that is, the intermediate prediction state. The intermediate prediction state can be represented by the following formula:
[0100]
[0101] Among them, represents the intermediate prediction state, F() represents the state prediction model.
[0102] Next, convert the locally predicted motion state predicted by the model into data in the world coordinate system to obtain the target prediction state, and simultaneously update the coordinate offset vector and the target input. Among them, the target prediction state can be represented by the following formula:
[0103]
[0104] Among them, represents the target prediction state, Δq i represents the currently initialized coordinate offset vector, and R is the transformation matrix.
[0105] Among them, the coordinate offset vector can be updated by the following formula, that is, the initialized coordinate offset vector in the next-step prediction can be characterized by the following formula:
[0106]
[0107] Among them, Δq i+1 is the updated coordinate offset vector, Δq i+1 that is, the coordinate offset vector used in the next-step prediction, Δq i is the current coordinate offset vector, Δq i that is, the coordinate offset vector used in the previous-step prediction, and R is the above transformation matrix, is the first intermediate prediction state predicted by the state prediction model.
[0108] In the above explanation of the example for Figure 2 the value of i increases gradually with a fixed step size of 1 until it reaches the sequence length of the actual state subsequence.
[0109] Finally, update the model parameters of the state prediction model based on the loss function characterized by the following formula:
[0110]
[0111] Among them, Loss represents the loss value, K is the sequence length of the prediction state subsequence and the actual state subsequence, is the i-th target prediction state in the prediction state subsequence, q i is the i-th actual state in the actual state subsequence.
[0112] In a possible way, the data set can be obtained as follows: when it is determined that the first preset condition is not satisfied, sample the target positions in sequence, and sample the actual instructions executed by the robot during the process of reaching the target positions, the timestamps when the actual instructions start to be executed, the actual states, and the timestamps when the robot is in the actual states; when it is determined that the first preset condition is satisfied, generate an actual instruction sequence according to all the sampled actual instructions and the timestamps when each actual instruction starts to be executed, and generate an actual state sequence according to all the sampled actual states and the timestamps when the robot is in each actual state.
[0113] It should be noted that, among two adjacent samplings of the target position, when the robot reaches the previous target position, the next target position is sampled.
[0114] As an example, the electronic device can randomly sample target points in the experimental area as the target position and send them to the robot, and the robot reaches the target position based on its own navigation module.
[0115] As an example, the first preset condition can be a condition indicating that the amount of collected data does not meet the standard. For example, the first preset condition can be that the number of sampled target positions is greater than the second preset number.
[0116] Through the above method, an automated acquisition process of the motion data of a planar mobile robot in a real environment is realized, so that the interaction between the robot and the environment can be captured. Compared with the data simulated based on the theoretical environment, the prediction accuracy of the model obtained by fitting the collected data is improved, and there is no need to manually adjust the model in the real environment.
[0117] In a possible way, among two adjacent samplings of the target position, during the process of the robot reaching the previous target position, when the second preset condition is met within the third preset duration, the next target position is sampled.
[0118] Among them, the second preset condition can be a condition indicating that the robot cannot reach the target position within the preset duration. Common situations are that the robot is stationary for a long time or the robot has not reached the target position for a long time.
[0119] Among them, when the second preset condition is a condition indicating that the robot has not reached the target position for a long time, the third preset duration can represent the duration from when the robot starts to navigate to the target position; when the second preset condition is a condition indicating that the robot is stationary for a long time, the third preset duration can represent the duration that the robot has been at a certain position.
[0120] Through the above method, when the robot is stationary for a long time or the robot has not reached the target position for a long time, the next target position is directly sampled. In this way, the acquisition efficiency of the data set is improved.
[0121] As another example, the above sampled target position can be replaced by sampling the actual instruction. During the process of the robot executing the sampled actual instruction, the executed actual instruction, the timestamp when the actual instruction starts to be executed, the actual state, and the timestamp when the actual state is in are sampled. Correspondingly, the first preset condition can be that the sampled actual instructions reach the third preset number. When it is determined that the first preset condition is met, an actual instruction sequence is generated according to all the sampled actual instructions and the timestamps when each actual instruction starts to be executed, and an actual state sequence is generated according to all the sampled actual states and the timestamps when each actual state is in.
[0122] Figure 3 It is a schematic flow chart of collecting each data in a data set shown according to an exemplary embodiment of the present disclosure. Referring to Figure 3 , the acquisition program refers to the computer program corresponding to the above-mentioned acquisition of the data set. After the acquisition program runs, the acquisition program can determine whether the first preset condition is met, that is, whether the amount of collected data does not meet the standard. In the case of determining that the first preset condition is not met, sample the next target position and send it to the robot. The robot tries to reach the target position. After that, every time a preset interval is waited, it is determined whether the robot has reached the target position. If the target position is reached, it is determined again whether the first preset condition is met; if the target position is not reached, it is further determined whether the robot is stationary for a long time. In the case of being stationary for a long time, sample the next target position again and send it to the robot. If the robot is not stationary for a long time, it is determined whether the robot has not reached the target position for a long time. If the robot has not reached the target position for a long time, sample the next target position again and send it to the robot. If the robot has not not reached the target position for a long time, wait. When the waiting duration reaches the preset interval, it is determined again whether the robot has reached the target position and other subsequent steps.
[0123] Through the above method, the automatic acquisition of the data set is realized while ensuring the acquisition efficiency and acquisition amount of the data set.
[0124] Figure 4 It is another schematic flow chart of a method for training a robot state prediction model shown according to an exemplary embodiment of the present disclosure. Referring to Figure 4 , the real machine represents the robot. By inputting a test instruction sequence to the real machine and repeating the real machine operation, motion capture / on-machine acquisition is performed to obtain an actual state sequence (i.e., the actual state sequence). Based on the test instruction sequence, a simulated instruction sequence (i.e., the actual instruction sequence) is obtained. Select the actual initial state (i.e., the actual state before the first actual state in the actual state subsequence in the actual state sequence) for initialization to obtain the simulated initial state. Use the dynamics model (i.e., the state prediction model) to process the simulated instruction sequence and the simulated initial state to obtain the simulated state (i.e., the intermediate prediction state). Use the simulated initial state to perform data conversion on the simulated state and record the data, and then obtain the simulated state sequence. The simulated state sequence contains different prediction state subsequences; for the simulated state sequence and the actual state sequence, data processing and feature selection are performed to obtain prediction state subsequences and actual state subsequences with matching sequence lengths, that is, feature sequence pairs. Index evaluation is performed to generate a feedback signal (i.e., the value corresponding to the loss function), and the feedback signal is sent to the parameter optimizer, so that the parameter optimizer optimizes / updates the model parameters based on the feedback information, thereby obtaining an updated dynamics model.
[0125] Based on the same concept, an embodiment of the present disclosure provides a training device for a robot state prediction model. Figure 5 It is a block diagram of a training device for a robot state prediction model shown according to an exemplary embodiment of the present disclosure. Refer to Figure 5 , the training device 500 for the robot state prediction model includes:
[0126] An acquisition module 501, configured to acquire a data set of the robot, where the data set includes an actual instruction sequence and an actual state sequence;
[0127] A training module 502, configured to perform cyclic iterative training on the state prediction model according to the data set, and output the trained state prediction model when the training is completed. During a cyclic iterative training process, the following operations are performed:
[0128] Select an actual state subsequence from the actual state sequence, where each actual state in the actual state subsequence is data in the world coordinate system;
[0129] Process a target input using the state prediction model to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system. The target input is at least determined by the actual instruction sequence, and the inputs and outputs related to the state in the state prediction model are all data in the local coordinate system;
[0130] Update the state prediction model according to the prediction state subsequence composed of all the target prediction states and the actual state subsequence.
[0131] In a possible manner, the training module 501 includes:
[0132] An initialization sub-module, configured to initialize a coordinate offset vector and a target input for predicting intermediate prediction states corresponding to each actual state in the actual state subsequence;
[0133] A processing sub-module, configured to process the target input using the state prediction model to obtain an intermediate prediction state corresponding to the actual state;
[0134] A prediction sub-module, configured to perform data transformation on the intermediate prediction state according to the coordinate offset vector to obtain a target prediction state in the world coordinate system.
[0135] In a possible manner, the initialization sub-module initializes the coordinate offset vector specifically through the following method:
[0136] When predicting the intermediate prediction state corresponding to the first actual state in the actual state subsequence, initialize the coordinate offset vector according to the previous actual state of the first actual state in the actual state sequence;
[0137] When predicting the intermediate prediction state corresponding to a non-first actual state in the actual state subsequence, initialize the coordinate offset vector according to the previously initialized coordinate offset vector and the first intermediate prediction state predicted by the state prediction model for the actual state subsequence.
[0138] In a possible way, the initialization sub-module initializes the target input in the following way:
[0139] Determine the target input according to the target timestamp, the first preset duration, and the actual instruction sequence, where the target timestamp is the timestamp carried by the actual state corresponding to the intermediate prediction state currently predicted by the state prediction model; or,
[0140] Determine the first input according to the target timestamp, the first preset duration, and the actual instruction sequence, determine the second input according to the target prediction state corresponding to the first actual state in the actual state subsequence or the intermediate prediction state already predicted by the state prediction model for the actual state subsequence, and determine the target input according to the first input and the second input.
[0141] In a possible way, the data set further includes the timestamps when each of the actual instructions starts to be executed, and the timestamps when each of the actual states starts to be in, and the obtaining module 501 is specifically configured to:
[0142] In the case of determining that the first preset condition is not satisfied, sequentially sample the target positions, and sample the actual instructions executed by the robot during the process of reaching the target positions, the timestamps when the actual instructions start to be executed, the actual states, and the timestamps when the robot is in the actual states, where in two adjacent samplings of the target positions, sample the next target position when the robot reaches the previous target position;
[0143] In the case of determining that the first preset condition is satisfied, generate an actual instruction sequence according to all the actually sampled instructions and the timestamps when each of the actual instructions starts to be executed, and generate an actual state sequence according to all the sampled actual states and the timestamps when the robot is in each of the actual states.
[0144] In a possible way, in two adjacent samplings of the target positions, during the process of the robot reaching the previous target position, sample the next target position when the second preset condition is satisfied within the third preset duration.
[0145] In possible ways, the sequence lengths of the actual state subsequences selected during the loop iteration training process are different.
[0146] Among them, the implementation manners of the various modules in the above device 500 may refer to the above related embodiments, and are not described herein again in this embodiment.
[0147] The embodiments of the present disclosure further provide a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the above training method are implemented.
[0148] The embodiments of the present disclosure further provide an electronic device, including:
[0149] A storage device, on which a computer program is stored;
[0150] A processing device, configured to execute the computer program in the storage device to implement the steps of the above training method.
[0151] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example, and should not impose any limitation on the functions and usage scopes of the embodiments of the present disclosure.
[0152] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0153] Typically, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0154] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are performed.
[0155] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0156] In some embodiments, the electronic device can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed network.
[0157] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.
[0158] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: obtain a data set of the robot, the data set including an actual instruction sequence and an actual state sequence; perform iterative training on a state prediction model according to the data set, and when the training is completed, output the trained state prediction model. During one iterative training process, the following operations are performed: select an actual state subsequence from the actual state sequence, where each actual state in the actual state subsequence is data in the world coordinate system; process a target input using the state prediction model to obtain intermediate prediction states respectively corresponding to the actual states in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system, where the target input is at least determined by the actual instruction sequence, and both the state-related inputs and outputs in the state prediction model are data in the local coordinate system; update the state prediction model according to a prediction state subsequence composed of all the target prediction states and the actual state subsequence.
[0159] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0161] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the module itself.
[0162] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and not limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0163] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0165] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0166] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.
Claims
1. A training method for a robot state prediction model, characterized in that, it includes: Obtain a data set of the robot, where the data set includes an actual instruction sequence and an actual state sequence; Perform iterative training on the state prediction model according to the data set, and when the training is completed, output the trained state prediction model. During one iterative training process, perform the following operations: Select an actual state subsequence from the actual state sequence, and each actual state in the actual state subsequence is data in the world coordinate system; Use the state prediction model to process the target input to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system. The target input is determined at least by the actual instruction sequence, and the inputs and outputs related to the state in the state prediction model are all data in the local coordinate system; Update the state prediction model according to the prediction state subsequence composed of all the target prediction states and the actual state subsequence.
2. The method according to claim 1, characterized in that, The step of using the state prediction model to process the target input to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and performing data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system includes: For predicting the intermediate prediction states corresponding to each actual state in the actual state subsequence, initialize a coordinate offset vector and a target input; Use the state prediction model to process the target input to obtain the intermediate prediction state corresponding to the actual state; Perform data transformation on the intermediate prediction state according to the coordinate offset vector to obtain a target prediction state in the world coordinate system.
3. The method according to claim 2, characterized in that, Initialize the coordinate offset vector in the following way: When predicting the intermediate prediction state corresponding to the first actual state in the actual state subsequence, initialize the coordinate offset vector according to the previous actual state of the first actual state in the actual state sequence; When predicting the intermediate prediction state corresponding to a non-first actual state in the actual state subsequence, initialize the coordinate offset vector according to the previously initialized coordinate offset vector and the first intermediate prediction state predicted by the state prediction model for the actual state subsequence.
4. The method according to claim 2, characterized in that, Initialize the target input in the following way: Determine the target input according to the target timestamp, the first preset duration, and the actual instruction sequence. The target timestamp is the timestamp carried by the actual state corresponding to the intermediate prediction state currently predicted by the state prediction model; Or, Determine a first input according to the target timestamp, the first preset duration, and the actual instruction sequence, determine a second input according to the first actual state in the actual state subsequence or the target prediction state corresponding to the intermediate prediction state predicted by the state prediction model for the actual state subsequence, and determine a target input according to the first input and the second input.
5. The method according to any one of claims 1-4, wherein, the data set further includes the timestamps when each of the actual instructions starts to be executed, and the timestamps when each of the actual states starts to be in, and the data set is obtained by the following method: In the case of determining that the first preset condition is not satisfied, successively sample target positions, and sample the actual instructions executed by the robot during the process of reaching the target positions, the timestamps when the actual instructions start to be executed, the actual states, and the timestamps when the robot is in the actual states, wherein in two adjacent samplings of target positions, the next target position is sampled when the robot reaches the previous target position; In the case of determining that the first preset condition is satisfied, generate an actual instruction sequence according to all the sampled actual instructions and the timestamps when each of the actual instructions starts to be executed, and generate an actual state sequence according to all the sampled actual states and the timestamps when the robot is in each of the actual states.
6. The method according to claim 5, wherein, In two adjacent samplings of target positions, during the process of the robot reaching the previous target position, in the case of satisfying the second preset condition within the third preset duration, sample the next target position.
7. The method according to claim 1, wherein, The sequence lengths of the actual state subsequences selected during the cyclic iterative training process are different.
8. A training device for a robot state prediction model, wherein, comprising: an acquisition module, configured to acquire a data set of a robot, the data set including an actual instruction sequence and an actual state sequence; a training module, configured to perform cyclic iterative training on a state prediction model according to the data set, and when the training is completed, output the trained state prediction model, and perform the following operations during one cyclic iterative training process: Select an actual state subsequence from the actual state sequence, and each actual state in the actual state subsequence is data in the world coordinate system; Process a target input by using the state prediction model to obtain intermediate prediction states corresponding to each actual state in the actual state subsequence, and perform data transformation on the intermediate prediction states to obtain target prediction states in the world coordinate system, the target input is at least determined by the actual instruction sequence, and both the input and output related to the state in the state prediction model are data in the local coordinate system; Update the state prediction model according to the prediction state subsequence composed of all the target prediction states and the actual state subsequence.
9. A computer-readable medium, on which a computer program is stored, wherein, when the program is executed by a processing device, the steps of the method according to any one of claims 1-7 are implemented.
10. An electronic device, characterized in that, it includes: a storage device on which a computer program is stored; a processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.