Robot action prediction method and device, equipment and storage medium
By introducing the Mamba module into the robot motion prediction model to extract the global and local features of the action sequence and fusion, the problem of insufficient utilization of historical action sequence information in the prior art is solved, and the reliability and robustness of action prediction are improved.
Patent Information
- Application Number
- CN202510506659.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, the information in historical action sequences is insufficiently utilized, resulting in poor reliability and robustness of robot action prediction.
A robot action prediction method is adopted. By introducing a global feature extraction module and a local feature extraction module in the reinforced learning action prediction model, the Mamba module is used to extract the global and local features of the action vector sequence, and fuse them to generate target features to predict the robot's action sequence.
By extracting and fusion of global and local features of action sequences, the performance of the action prediction model is improved, the selectivity and context perception of the model are enhanced, and the reliability and robustness of prediction are improved.
Smart Images

Figure CN120023835A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a robot motion prediction method, device, equipment and storage medium. Background Art
[0002] The decision-making process in robot reinforcement learning generally involves learning the mapping from state observation to action, so as to maximize the cumulative discounted reward. In offline reinforcement learning, a sequence modeling paradigm can be established through decision transformers to predict the robot's future action sequence based on the sequence of past rewards, states, and actions to achieve the expected reward.
[0003] However, the sequence modeling of robot reinforcement learning trajectories is different from traditional sequence modeling. Specifically, the next state in reinforcement learning problems depends only on the current state and action, but not on the past state, and due to the continuity of time steps, the characteristics of each step are related to long-term historical information. When performing sequence modeling of robot reinforcement learning trajectories based on existing technologies, the information in the historical action sequence is not fully utilized, and there is a problem of poor reliability and robustness in predicting actions. Summary of the invention
[0004] The purpose of the present application is to provide a robot motion prediction method, device, equipment and storage medium to address the deficiencies in the above-mentioned prior art, so as to solve the problems in the prior art of insufficient utilization of information in historical motion sequences and poor reliability and robustness of predicted motions.
[0005] To achieve the above objectives, the technical solutions adopted in this application are as follows: In a first aspect, the present application provides a robot motion prediction method, which is applied to a target robot deployed with a reinforcement learning motion prediction model, wherein the reinforcement learning motion prediction model includes: an encoder, a feature extractor, and a decoder, wherein the feature extractor includes at least: a global feature extraction module, a local feature extraction module, and a feature fusion module, wherein the global feature extraction module and the local feature extraction module each include at least one Mamba module, and the method includes: Check whether the target robot is started. If so, iterate and loop to execute the following steps: A. acquiring at least one historical action information from the historical action library of the target robot, inputting the historical action information into the encoder, and obtaining an action vector sequence, wherein the action vector sequence includes at least one action vector, and the action vector includes: a state element, an action element, and a reward element; B. Inputting the motion vector sequence into the feature extractor, extracting global features according to the motion vector sequence by the Mamba module in the global feature extraction module, and extracting local features of the motion vector sequence by the local feature extraction module based on the Mamba module in the local feature extraction module to obtain at least one local feature; C. Inputting the global feature and each of the local features into the feature fusion module to obtain the target feature; D. Input the target feature into the decoder, and the decoder predicts the predicted action sequence of the robot, controls the movement of the target robot according to the predicted action sequence, determines the state and reward according to the actual action sequence of the target robot when it moves, and adds the actual action sequence, state and reward as a historical action information to the historical action library.
[0006] Optionally, inputting the historical motion information of the target robot into the encoder to obtain a motion vector sequence includes: Sampling the historical action information to obtain a plurality of continuous triplet information of the target robot, wherein the triplet information includes: state information, action information and reward information; The plurality of continuous triplet information are input into the encoder, and the encoder performs vectorization processing on each triplet information to obtain the motion vector sequence.
[0007] Optionally, the local feature extraction module further includes: a segmentation module and a conversion module; The process of extracting at least one local feature from the motion vector sequence by the local feature extraction module includes: The segmentation module performs segmentation processing on the motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence; The Mamba module extracts features from each of the action vector subsequences to obtain at least one initial feature; The conversion module performs conversion processing on each of the initial features to obtain the at least one local feature.
[0008] Optionally, the segmentation module segments the motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence, including: The segmentation module fills the last bit of the motion vector sequence with a preset value until the length of the motion vector sequence reaches N, thereby obtaining a filled motion vector sequence; The segmentation module performs segmentation processing on the padded motion vector sequence based on a preset segmentation length to obtain at least one motion vector subsequence of the motion vector sequence.
[0009] Optionally, the local feature extraction module further includes: a regularization module; The step of inputting the global feature and each of the local features into the feature fusion module comprises: Inputting the at least one local feature into the regularization module to close neurons to obtain a processed local feature; The global features and each processed local feature are input into the feature fusion module.
[0010] Optionally, the feature fusion module includes: a first connection layer and a feedforward network layer; The step of inputting the global feature and each of the local features into the feature fusion module to obtain the target feature includes: Inputting the global feature and the local feature into the first connection layer, and performing connection processing by the first connection layer to obtain a connected feature; Inputting the concatenated features into the feedforward network layer, and performing transformation processing by the feedforward network layer to obtain transformed features; The transformed features and the connected features are concatenated to obtain the target features.
[0011] Optionally, the action prediction model further includes: a linear module, wherein the linear module includes: a normalization layer, a first linear layer, an activation layer, a second linear layer and a second connection layer; Inputting the target feature into the decoder comprises: The target feature is input into the linear module, and the linear module sequentially performs normalization processing, linearization processing, activation processing and connection processing to obtain a processed target feature, and the processed target feature is input into the decoder.
[0012] In a second aspect, the present application provides a robot motion prediction device, the device comprising: An encoding module, configured to obtain at least one historical action information from a historical action library of a target robot, input the historical action information into an encoder, and obtain an action vector sequence, wherein the action vector sequence includes at least one action vector, and the action vector includes: a state element, an action element, and a reward element; An extraction module is used to input the motion vector sequence into a feature extractor, and a Mamba module in a global feature extraction module extracts a global feature according to the motion vector sequence, and a local feature extraction module extracts a local feature of the motion vector sequence based on the Mamba module in the local feature extraction module to obtain at least one local feature; A fusion module, used for inputting the global feature and each of the local features into a feature fusion module to obtain a target feature; A prediction module is used to input the target feature into a decoder, and the decoder predicts the predicted action sequence of the robot, controls the movement of the target robot according to the predicted action sequence, determines the state and reward according to the actual action sequence of the target robot when it moves, and adds the actual action sequence, state and reward as a historical action information to the historical action library.
[0013] Optionally, the encoding module is used to: Sampling the historical action information to obtain a plurality of continuous triplet information of the target robot, wherein the triplet information includes: state information, action information and reward information; The plurality of continuous triplet information are input into the encoder, and the encoder performs vectorization processing on each triplet information to obtain the motion vector sequence.
[0014] Optionally, the local feature extraction module further includes: a segmentation module and a conversion module; The extraction module is also used for: The segmentation module performs segmentation processing on the motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence; The Mamba module extracts features from each of the action vector subsequences to obtain at least one initial feature; The conversion module performs conversion processing on each of the initial features to obtain the at least one local feature.
[0015] Optionally, the extraction module is further used for: The segmentation module fills the last bit of the motion vector sequence with a preset value until the length of the motion vector sequence reaches N, thereby obtaining a filled motion vector sequence; The segmentation module performs segmentation processing on the padded motion vector sequence based on a preset segmentation length to obtain at least one motion vector subsequence of the motion vector sequence.
[0016] Optionally, the local feature extraction module further includes: a regularization module; The fusion module is also used for: Inputting the at least one local feature into the regularization module to close neurons to obtain a processed local feature; The global features and each processed local feature are input into the feature fusion module.
[0017] Optionally, the feature fusion module includes: a first connection layer and a feedforward network layer; The fusion module is also used for: Inputting the global feature and the local feature into the first connection layer, and performing connection processing by the first connection layer to obtain a connected feature; Inputting the concatenated features into the feedforward network layer, and performing transformation processing by the feedforward network layer to obtain transformed features; The transformed features and the connected features are concatenated to obtain the target features.
[0018] Optionally, the action prediction model further includes: a linear module, wherein the linear module includes: a normalization layer, a first linear layer, an activation layer, a second linear layer and a second connection layer; The prediction module is also used to: The target feature is input into the linear module, and the linear module sequentially performs normalization processing, linearization processing, activation processing and connection processing to obtain a processed target feature, and the processed target feature is input into the decoder.
[0019] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of a robot motion prediction method as described in any one of the first aspects.
[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a robot motion prediction method as described in any one of the first aspects are executed.
[0021] The beneficial effects of the present application are: by extracting features from the global features and local features of the action sequence respectively, it is possible to pay attention to the temporal continuity of the historical action sequence while paying attention to the local correlation of the action sequence. By extracting features through the Mamba module, the selectivity and context-awareness of the model can be further enhanced, thereby improving the performance of the model. In addition, since the Mamba module has a significant lightweight advantage, the action prediction model improved based on the Mamba module in the present application also has the advantages of being lightweight and easy to deploy.
[0022] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 A schematic diagram of the architecture of a reinforcement learning action prediction model provided in an embodiment of the present application is shown; Figure 2 A flowchart of a robot motion prediction method provided by an embodiment of the present application is shown; Figure 3 A flowchart of obtaining an action vector sequence provided by an embodiment of the present application is shown; Figure 4 A flowchart of extracting local features provided by an embodiment of the present application is shown; Figure 5 A flowchart of segmenting an action vector sequence provided by an embodiment of the present application is shown; Figure 6 A schematic diagram of the architecture of another reinforcement learning action prediction model provided in an embodiment of the present application is shown; Figure 7 A flow chart of regularizing local features provided by an embodiment of the present application is shown; Figure 8 A flow chart of feature fusion provided by an embodiment of the present application is shown; Fig. 9 A schematic diagram of the structure of a robot motion prediction device provided in an embodiment of the present application is shown; Fig.10 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0026] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0027] After the decision transformer is introduced into reinforcement learning, it can predict the future action sequence based on the robot's historical action sequence, state and past rewards. The decision transformer can reinterpret the decision process as a sequence modeling problem, regard the robot reinforcement learning process as a sequence modeling task, learn a state-action mapping based on reward conditions, and construct the robot's action prediction model.
[0028] However, due to the particularity of robot reinforcement learning, there are many problems in directly treating reinforcement learning trajectories as a sequence and modeling sequences based on traditional methods. Specifically, there is a significant local correlation between the steps in the trajectory sequence, and the probability of transitioning to the next state depends only on the current state and action, so it is necessary to pay attention to the local correlation of the trajectory sequence. In addition, since the time steps are continuous, the characteristics of each step are related to long-term historical information, which shows that the reinforcement learning trajectory also has internal global correlation.
[0029] In summary, how to effectively capture the global and local features in the enhanced trajectory sequence has become the key to improving the performance of the action prediction model.
[0030] Based on this, this application proposes a robot motion prediction method, which extracts global features and local features of the trajectory sequence through the Mamba module, which can better understand the correlation within the reinforcement learning trajectory, thereby effectively improving the model performance of the motion prediction model.
[0031] The method of the present application can be applied to a target robot that is deployed with a reinforcement learning action prediction model, and the specific execution subject can be a control device in the target robot. The reinforcement learning action prediction model is used to predict the action based on the history of each part of the robot. 0 The action sequence within the time period is subjected to feature extraction and action prediction, thereby generating the action sequence of the robot within the future time period T.
[0032] Figure 1 This is a schematic diagram of the model architecture of a reinforcement learning action prediction model given in this application, refer to Figure 1The reinforcement learning action prediction model includes: an encoder, a feature extractor and a decoder. The feature extractor includes at least: a global feature extraction module, a local feature extraction module and a feature fusion module. The global feature extraction module and the local feature extraction module each include at least one Mamba module. Among them, the Mamba module has efficient sequence modeling capabilities and powerful multi-scale dependency capture capabilities, so it can ensure that both global features and local features in the input action trajectory can be fully utilized, thereby improving the performance of the model.
[0033] Next, combine Figure 2 , the robot motion prediction method of the present application is described, referring to Figure 2 , the method comprising: Check whether the target robot is started. If so, iterate and loop to execute the following steps: S201. Obtain at least one historical action information from a historical action library of the target robot, input the historical action information into an encoder, and obtain an action vector sequence. The action vector sequence includes at least one action vector. The action vector includes: a state element, an action element, and a reward element.
[0034] After the robot is started, the robot's actions, states, and responses can be collected in real time, and the actions, states, and responses at each moment can be stored in a historical action library as historical action information at that moment.
[0035] Among them, the robot's action includes the output signal of the robot actuator, and the robot's state can be the environment state and its own state under the robot's current action. The environment state includes surrounding obstacle information, map information, etc., and its own state includes position information and posture information, etc.
[0036] In the present application, k consecutive historical action information of the past time can be obtained, and the historical action information is input into the encoder, and the encoder performs encoding processing to obtain an action vector sequence. Exemplarily, the action vector sequence can be expressed as the following formula (1), which includes multiple action vectors, such as < >, Represents the reward element in the historical action information, Represents the state element in the historical action information, Represents the action element in the historical action information. represents the original action vector sequence, Indicates that after embedding function The sequence of motion vectors after processing.
[0037] (1) It is worth noting that the robot's actions at each moment have corresponding historical action information. In one possible implementation method, the actions, states, and rewards at each moment can be taken as a set of historical action information to obtain historical action information for multiple consecutive moments.
[0038] S202. Input the motion vector sequence into the feature extractor, and use the Mamba module in the global feature extraction module to extract global features based on the motion vector sequence. Use the local feature extraction module to extract local features of the motion vector sequence based on the Mamba module in the local feature extraction module to obtain at least one local feature.
[0039] Reference Figure 1 The feature extractor includes a global feature extraction module and a local feature extraction module. The motion vector sequence can be processed by the global feature extraction module and the local feature extraction module to extract the global feature and at least one local feature of the motion vector sequence respectively.
[0040] Among them, the global feature extraction module and the local feature extraction module are both composed of Mamba modules. The Mamba module in the global feature extraction module can model the motion vector sequence as a whole to obtain the global features of the motion vector sequence. The local feature extraction module can first segment the motion vector sequence to obtain multiple motion vector subsequences, and the Mamba module can model each motion vector subsequence separately to obtain the local features corresponding to each motion vector subsequence.
[0041] The feature extraction of the Mamba module is mainly based on the structured state space model and dynamic adjustment mechanism, combining linear computational complexity with global feature modeling capabilities. Among them, when processing action vector sequences, the Mamba module models the sequence data through the state space equation, thereby efficiently modeling the long-term dependencies between states. The dynamic convolution method in the Mamba module can dynamically generate convolution kernel weights based on the input features to enhance the perception of local texture and spatial differences. The channel attention method can adaptively adjust the channel importance through the SE module (Squeeze-and-Excitation) to suppress redundant features. The multi-scale feature fusion method can be combined with the cross-scale self-attention module to capture global contextual information at different resolutions. At the same time, for multimodal input, the Mamba module can also screen cross-modal correlation features through the gating mechanism, suppress noise information, and bidirectionally transfer hidden states in the temporal dimension to enhance the integration of complementary information between modalities.
[0042] S203: Input the global features and each local feature into a feature fusion module to obtain the target feature.
[0043] Optionally, the feature fusion module may perform feature fusion on the global features and the local features. The fusion method may be, for example, feature splicing on the global features and the local features, and standardizing the spliced features to obtain the target features.
[0044] As a possible implementation method, the feature fusion module can also combine the attention mechanism, use self-attention to adjust the feature weights, and cross-attention to fuse multimodal information to obtain the fused target features.
[0045] S204. Input the target features into the decoder, and the decoder predicts the predicted action sequence of the robot. The target robot moves according to the predicted action sequence, and the state and reward are determined according to the actual action sequence of the target robot when it moves. The actual action sequence, state and reward are added to the historical action library as a historical action information.
[0046] Optionally, the decoder may decode the target features to predict a predicted action sequence of the robot, wherein the predicted action sequence may be an executable action sequence of the robot within T time steps in the future.
[0047] After obtaining the predicted action sequence, the controller in the robot can control the corresponding parts of the robot to execute the predicted action sequence, and after executing each action, record the status and reward of the robot's current action, and add the robot's current action, current status and reward as a historical action information to the historical action library.
[0048] In the embodiment of the present application, by extracting features from the global features and local features of the action sequence respectively, it is possible to pay attention to the temporal continuity of the historical action sequence while paying attention to the local correlation of the action sequence. By extracting features through the Mamba module, the selectivity and context-awareness of the model can be further enhanced, thereby improving the performance of the model. In addition, since the Mamba module has a significant lightweight advantage, the action prediction model improved based on the Mamba module in the present application also has the advantages of being lightweight and easy to deploy.
[0049] The following is a further explanation of the above-mentioned input of the historical action information of the target robot into the encoder to obtain the action vector sequence, such as Figure 3 As shown, the above step S201 includes: S301. Sampling and processing historical action information to obtain a plurality of continuous triplet information of the target robot, the triplet information including: state information, action information and reward information.
[0050] Optionally, sampling the historical action information may be performed by obtaining a state and a reward corresponding to each action in the historical action information, and treating each action, the state corresponding to the action, and the reward corresponding to the action as a triplet of information.
[0051] As an optional implementation method, the historical action information of the same part of the robot can be sampled and processed in chronological order to obtain multiple continuous triplet information of each part.
[0052] S302 , input a plurality of continuous triplet information into an encoder, and the encoder performs vectorization processing on each triplet information to obtain a motion vector sequence.
[0053] The continuous triplets are input into the encoder in time sequence, and the encoder can encode each triplet and encode multiple continuous triplet information into a motion vector sequence.
[0054] As a possible implementation method, the encoder includes an autoencoder, an encoding layer, and a fully connected layer or a convolutional layer can be used as an encoder to extract features from continuous triplet information, generate a vector representation of a preset dimension, and obtain an action vector sequence.
[0055] Optionally, the local feature extraction module also includes: a segmentation module and a conversion module.
[0056] like Figure 4 As shown, the process of extracting at least one local feature according to the motion vector sequence by the local feature extraction module includes: S401 . A segmentation module performs segmentation processing on a motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence.
[0057] It should be noted that each time an action prediction is performed based on historical action information, the number of actions in the past T0 time step may be different, so the amount of historical action information may be different, which in turn causes the length of the action vector sequence to be different. By segmenting the action vector sequence based on the segmentation module, at least one action vector subsequence of the same length can be obtained.
[0058] S402: The Mamba module extracts features from each action vector subsequence to obtain at least one initial feature.
[0059] The number of Mamba modules may be one or more. When the number of Mamba modules is one, each action vector subsequence may be sequentially input into the Mamba module, and the Mamba module extracts features from the action vector subsequence to obtain initial features corresponding to the action vector subsequence. When the number of Mamba modules is multiple, the action vector subsequence may be divided into multiple groups, and the action vector subsequences of each group may be sequentially input into the corresponding Mamba module, and each Mamba module extracts features from the action vector subsequences of each group to obtain initial features corresponding to each action vector subsequence.
[0060] In a possible implementation, the action vector subsequence may be firstly subjected to layer normalization processing, and then modeled based on the Mamba module to obtain the initial features of the action vector subsequence.
[0061] S403: The conversion module performs conversion processing on each initial feature to obtain at least one local feature.
[0062] Among them, the conversion module can perform dimensional conversion on the initial features and ensure that the number of elements of the initial features remains unchanged, including changing the row and column arrangement of the initial features, so that the dimensions of the initial features meet the processing requirements of the next step module.
[0063] The following is a further description of the above step S401: Figure 5 As shown, the process of segmenting the motion vector sequence by the segmentation module to obtain at least one motion vector subsequence corresponding to the motion vector sequence includes: S501 , a segmentation module fills a preset value at the end of a motion vector sequence until the length of the motion vector sequence reaches N, thereby obtaining a filled motion vector sequence.
[0064] Optionally, the segmentation module can fill 0 at the end of the motion vector sequence until the length of the motion vector sequence reaches N, thereby obtaining a filled motion vector sequence. N is a positive integer greater than 0, and its value can be based on the past T 0 The maximum number of actions in a time step is determined.
[0065] S502: A segmentation module performs segmentation processing on the padded motion vector sequence based on a preset segmentation length to obtain at least one motion vector subsequence of the motion vector sequence.
[0066] The preset segmentation length may be expressed as L, and the value of L may be any integer greater than 0. In a possible implementation, the value of L may be a multiple of 3 to ensure that each action vector subsequence includes multiple complete historical action information.
[0067] Optional, such as Figure 6 As shown in Figure 1, it is a schematic diagram of the overall structure of a reinforcement learning action prediction model. Figure 6 , the local feature extraction module also includes: a regularization module. The regularization module can be a Dropout module.
[0068] like Figure 7 As shown, the process of inputting the global features and each local feature into the feature fusion module includes: S701: Input at least one local feature into a regularization module to perform regularization processing to obtain a processed local feature.
[0069] Among them, the regularization module is used to randomly discard neurons and perform output scaling compensation during training to prevent overfitting.
[0070] S702: Input the global features and each processed local feature into a feature fusion module.
[0071] In the embodiment of the present application, regularization processing is performed on local features through a regularization module, which can improve the robustness of the model while maintaining a high computational efficiency of the model, thereby effectively improving the performance of the model.
[0072] Reference Figure 6 , the feature fusion module includes: the first connection layer and the feedforward network layer. Figure 8 As shown, the process of inputting the global features and each local feature into the feature fusion module to obtain the target feature includes: S801, inputting global features and local features into the first connection layer, and the first connection layer performs connection processing to obtain connected features.
[0073] S802, input the concatenated features into the feedforward network layer, which performs transformation processing to obtain transformed features.
[0074] Optionally, the first connection layer can concatenate the global features and the local features into a connected feature, and input the connected feature into the feed-forward network layer. The feed-forward network layer can perform feature enhancement and linear transformation on the connected feature, specifically, by combining the linear transformation with the activation function to map the input data to a more complex feature space, thereby achieving efficient feature extraction and transformation of the data.
[0075] As a possible implementation method, the connected features can be firstly subjected to layer normalization processing to obtain processed connected features, and the processed connected features can be input into the feedforward network layer for transformation processing to obtain transformed features.
[0076] S803: Concatenate the transformed features and the connected features to obtain target features.
[0077] After obtaining the transformed features, the transformed features and the connected features can be concatenated through the add function to obtain the target features.
[0078] Optional, continue to refer to Figure 6 The reinforcement learning action prediction model also includes: a linear module, which includes: a normalization layer, a first linear layer, an activation layer, a second linear layer and a second connection layer.
[0079] The above process of inputting the target features into the decoder includes: The target features are input into the linear module, which performs normalization, linearization, activation and connection processing in sequence to obtain the processed target features, and the processed target features are input into the decoder.
[0080] Optionally, after the target features are input into the linear module, a normalization layer can perform normalization on the target features, a first linear layer can perform linearization on the normalized target features, an activation layer can perform activation on the linearized target features, a second linear layer can perform linearization on the activated target features, and a second connection layer can connect the target features processed by the second linear layer to obtain processed target features.
[0081] As a possible implementation manner, the second connection layer can connect the processed target features and the target features to obtain connected features.
[0082] It is worth noting that the model may include multiple linear modules. After being processed in sequence by each linear module, the final processed target features may be input into the encoder, and the encoder predicts the predicted action sequence for the next T time steps.
[0083] Based on the same inventive concept, a robot motion prediction device corresponding to the robot motion prediction method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned robot motion prediction method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0084] Fig. 9 A schematic diagram of the structure of a robot motion prediction device provided in an embodiment of the present application is shown.
[0085] The encoding module 901 is used to obtain at least one historical action information from the historical action library of the target robot, input the historical action information into the encoder, and obtain an action vector sequence, wherein the action vector sequence includes at least one action vector, and the action vector includes: a state element, an action element, and a reward element; An extraction module 902 is used to input the motion vector sequence into a feature extractor, and the Mamba module in the global feature extraction module extracts global features according to the motion vector sequence, and the local feature extraction module extracts local features of the motion vector sequence based on the Mamba module in the local feature extraction module to obtain at least one local feature; A fusion module 903 is used to input the global features and each local feature into a feature fusion module to obtain a target feature; The prediction module 904 is used to input the target features into the decoder, and the decoder predicts the predicted action sequence of the robot, controls the movement of the target robot according to the predicted action sequence, determines the state and reward according to the actual action sequence of the target robot when it moves, and adds the actual action sequence, state and reward as a historical action information to the historical action library.
[0086] Optionally, the encoding module 901 is used to: Sampling and processing the historical action information to obtain multiple continuous triplet information of the target robot, the triplet information includes: state information, action information and reward information; A plurality of continuous triplet information is input into the encoder, and the encoder vectorizes each triplet information to obtain an action vector sequence.
[0087] Optionally, the local feature extraction module further includes: a segmentation module and a conversion module; The extraction module 902 is also used for: The segmentation module performs segmentation processing on the motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence; The Mamba module extracts features from each action vector subsequence to obtain at least one initial feature; The conversion module performs conversion processing on each initial feature to obtain at least one local feature.
[0088] Optionally, the extraction module 902 is further used for: The segmentation module fills the last bit of the motion vector sequence with a preset value until the length of the motion vector sequence reaches N, thereby obtaining a filled motion vector sequence; The segmentation module performs segmentation processing on the padded motion vector sequence based on a preset segmentation length to obtain at least one motion vector subsequence of the motion vector sequence.
[0089] Optionally, the local feature extraction module further includes: a regularization module; The fusion module 903 is also used for: Inputting at least one local feature into a regularization module to turn off neurons to obtain a processed local feature; The global features and the processed local features are input into the feature fusion module.
[0090] Optionally, the feature fusion module includes: a first connection layer and a feedforward network layer; The fusion module 903 is also used for: The global features and the local features are input into the first connection layer, and the first connection layer performs connection processing to obtain the connected features; The connected features are input into the feed-forward network layer, which transforms them to obtain the transformed features; The transformed features and the connected features are concatenated to obtain the target features.
[0091] Optionally, the action prediction model further includes: a linear module, wherein the linear module includes: a normalization layer, a first linear layer, an activation layer, a second linear layer, and a second connection layer; The prediction module 904 is also used to: The target features are input into the linear module, which performs normalization, linearization, activation and connection processing in sequence to obtain the processed target features, and the processed target features are input into the decoder.
[0092] Fig.10 A structural schematic diagram of an electronic device provided in an embodiment of the present application is shown, including: a processor 1001, a storage medium 1002 and a bus 1003, wherein the storage medium 1002 stores machine-readable instructions executable by the processor 1001. When the electronic device runs a robot motion prediction method such as that in the embodiment, the processor 1001 communicates with the storage medium 1002 via the bus 1003, and the processor 1001 executes the machine-readable instructions, and the processor 1001 executes the preamble of the method item to execute the steps of the above-mentioned robot motion prediction method.
[0093] The present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor when it is running, and the processor performs the steps of the above-mentioned robot motion prediction method. In the present application, the computer program can also execute other machine-readable instructions when the processor is running to execute the method described in other embodiments. For the specific execution method steps and principles, please refer to the description of the embodiment, which will not be repeated in detail here.
[0094] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0095] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0097] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0098] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0099] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A robot motion prediction method, characterized in that: The invention is applied to a target robot deployed with a reinforcement learning action prediction model, wherein the reinforcement learning action prediction model comprises: an encoder, a feature extractor and a decoder, wherein the feature extractor comprises at least: a global feature extraction module, a local feature extraction module and a feature fusion module, wherein the global feature extraction module and the local feature extraction module respectively comprise at least one Mamba module, and wherein the method comprises: Check whether the target robot is started. If so, iterate and loop to execute the following steps: A. acquiring at least one historical action information from the historical action library of the target robot, inputting the historical action information into the encoder, and obtaining an action vector sequence, wherein the action vector sequence includes at least one action vector, and the action vector includes: a state element, an action element, and a reward element; B. Inputting the motion vector sequence into the feature extractor, extracting global features according to the motion vector sequence by the Mamba module in the global feature extraction module, and extracting local features of the motion vector sequence by the local feature extraction module based on the Mamba module in the local feature extraction module to obtain at least one local feature; C. Inputting the global feature and each of the local features into the feature fusion module to obtain the target feature; D. Input the target feature into the decoder, and the decoder predicts the predicted action sequence of the robot, controls the movement of the target robot according to the predicted action sequence, determines the state and reward according to the actual action sequence of the target robot when it moves, and adds the actual action sequence, state and reward as a historical action information to the historical action library.
2. The method according to claim 1, characterized in that: The step of inputting the historical action information into the encoder to obtain an action vector sequence includes: Sampling the historical action information to obtain a plurality of continuous triplet information of the target robot, wherein the triplet information includes: state information, action information and reward information; The plurality of continuous triplet information are input into the encoder, and the encoder performs vectorization processing on each triplet information to obtain the motion vector sequence.
3. The method according to claim 1, characterized in that The local feature extraction module also includes: a segmentation module and a conversion module; The process of extracting local features of the motion vector sequence by the local feature extraction module based on the Mamba module in the local feature extraction module to obtain at least one local feature includes: The segmentation module performs segmentation processing on the motion vector sequence to obtain at least one motion vector subsequence corresponding to the motion vector sequence; The Mamba module extracts features from each of the action vector subsequences to obtain at least one initial feature; The conversion module performs conversion processing on each of the initial features to obtain the at least one local feature.
4. The method according to claim 3, characterized in that The process of segmenting the motion vector sequence by the segmentation module to obtain at least one motion vector subsequence corresponding to the motion vector sequence includes: The segmentation module fills a preset value at the end of the action vector sequence until the length of the action vector sequence reaches N, thereby obtaining a filled action vector sequence, wherein N is a positive integer, and the value of N is determined based on the maximum value of the number of actions in the past T0 time step; The segmentation module performs segmentation processing on the padded motion vector sequence based on a preset segmentation length to obtain at least one motion vector subsequence of the motion vector sequence.
5. The method according to claim 3, characterized in that: The local feature extraction module also includes: a regularization module; The step of inputting the global feature and each of the local features into the feature fusion module comprises: Inputting the at least one local feature into the regularization module to close neurons to obtain a processed local feature; The global features and each processed local feature are input into the feature fusion module.
6. The method according to claim 1, characterized in that The feature fusion module includes: a first connection layer and a feedforward network layer; The step of inputting the global feature and each of the local features into the feature fusion module to obtain the target feature includes: Inputting the global feature and the local feature into the first connection layer, and performing connection processing by the first connection layer to obtain a connected feature; Inputting the concatenated features into the feedforward network layer, and performing transformation processing by the feedforward network layer to obtain transformed features; The transformed features and the connected features are concatenated to obtain the target features.
7. The method according to claim 6, characterized in that The action prediction model further includes: a linear module, wherein the linear module includes: a normalization layer, a first linear layer, an activation layer, a second linear layer and a second connection layer; Inputting the target feature into the decoder comprises: The target feature is input into the linear module, and the linear module sequentially performs normalization processing, linearization processing, activation processing and connection processing to obtain a processed target feature, and the processed target feature is input into the decoder.
8. A robot motion prediction device, characterized in that: include: An encoding module, configured to obtain at least one historical action information from a historical action library of a target robot, input the historical action information into an encoder, and obtain an action vector sequence, wherein the action vector sequence includes at least one action vector, and the action vector includes: a state element, an action element, and a reward element; An extraction module is used to input the motion vector sequence into a feature extractor, and a Mamba module in a global feature extraction module extracts a global feature according to the motion vector sequence, and a local feature extraction module extracts a local feature of the motion vector sequence based on the Mamba module in the local feature extraction module to obtain at least one local feature; A fusion module, used for inputting the global feature and each of the local features into a feature fusion module to obtain a target feature; A prediction module is used to input the target feature into a decoder, and the decoder predicts the predicted action sequence of the robot, controls the movement of the target robot according to the predicted action sequence, determines the state and reward according to the actual action sequence of the target robot when it moves, and adds the actual action sequence, state and reward as a historical action information to the historical action library.
9. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of a robot motion prediction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a robot motion prediction method as claimed in any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Reinforcement learning sequence decision-making method, system, equipment and medium
CN117972588A
Video sequence processing method and system, electronic equipment and storage medium
CN118864309A
Mama-based point cloud semantic segmentation method and device, equipment and medium
CN119006814A
Online reinforcement learning data increasing and expanding method based on diffusion model
CN119476372A
Method and system for constructing homography estimation model based on cross Mangban dynamic feature fusion optimization
CN119810605A