Robot motion generation method, device, robot, medium and program product
By generating action sequence annotation and verb extraction of robot joint angle data, and using preset language models and pretrained models to generate action instructions, the problem of large overhead of robot action generation and calculation is solved and the efficiency of action generation is improved.
Patent Information
- Application Number
- CN202510474598.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing robot action generation method has high calculation overhead, resulting in low action generation efficiency.
By generating the joint angle data of the robot, performing action sequence annotation and verb extraction, using the preset language model to generate action instructions, and outputting the model file to instruct the robot to complete the action.
The calculation overhead of robot action generation is reduced and the efficiency of action generation is improved.
Smart Images

Figure CN120002669B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a robot motion generation method, device, robot, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of artificial intelligence (AI), various simulated robots have emerged. These robots can execute actions according to robot motion instructions. A robot motion instruction is the angle change of each joint along the robot's route from the starting point to the destination of the task. This represents the position of each joint in the time series.
[0003] Traditionally, common robotic motion generation methods include: motion planning methods based on dynamic action primitives, motion planning methods based on reinforcement learning, and motion planning methods based on imitation learning. Methods based on dynamic action primitives are difficult to model and complex to solve; methods based on reinforcement learning suffer from time-consuming training, poor stability, and the potential for disruption to training scenarios; and motion planning methods based on imitation learning rely on learning from examples provided by a human instructor, such as learning optimal strategies from the decision-making data of human experts.
[0004] However, the above-mentioned robot motion generation method has high-dimensionality and complexity of joint motion data, which leads to high computational overhead of robot motion generation. Summary of the Invention
[0005] Based on this, it is necessary to provide a robot motion generation method, device, robot, computer-readable storage medium and computer program product that can reduce the computational overhead of robot motion generation and improve motion generation efficiency in response to the above technical problems.
[0006] In a first aspect, the present application provides a robot action generation method, comprising:
[0007] Generate action sequences based on the robot's joint angle data;
[0008] Label the generated action sequence to obtain the action sequence description;
[0009] Extracting verbs from the action sequence description, and inputting the extracted verbs into a preset language model, thereby generating action instructions corresponding to the action sequence description through the preset language model;
[0010] Inputting the action instructions into a pre-trained model, the pre-trained model outputting a model file based on the relationship between the action description corresponding to the action instructions and the motion data; the model file is used to instruct the robot to complete various actions;
[0011] Load and execute the model file.
[0012] In one embodiment, generating an action sequence based on the robot's joint angle data includes:
[0013] Collecting joint angle data of the robot; the joint angle data includes: joint angles of each joint;
[0014] The multiple joint angles corresponding to each joint are integrated in the order of execution time, wherein the joint angles corresponding to each joint at the same execution time are integrated into a set of motion data;
[0015] If the joint angles of some joints are missing in a set of motion data, the joint angles of the missing joints are searched forward or backward in the order of execution time, and the joint angles with the closest execution time are added to the motion data;
[0016] The obtained multiple groups of action data are sorted in the order of execution time to obtain an action sequence.
[0017] In one embodiment, the step of labeling the generated action sequence to obtain an action sequence description includes:
[0018] The generated action sequence is annotated by manual annotation and / or automatic annotation to obtain an action sequence description; the action sequence description includes: a text for describing the action.
[0019] In one embodiment, the pre-trained model includes: an encoder and a decoder;
[0020] The encoder is used to convert the action instruction into an action description;
[0021] The decoder is used to predict the motion data corresponding to the action description generated by the encoder, and output a model file based on the predicted motion data; wherein, the model file includes: the joint angles of each joint under a preset number of frames, the coordinates of each joint in a Cartesian coordinate system, and a timestamp; the timestamp is used to determine the execution order corresponding to the joint angles of each joint under a preset number of frames and the coordinates of each joint in a Cartesian coordinate system.
[0022] In one embodiment, after generating the motion sequence according to the joint angle data of the robot, the method further includes:
[0023] Perform linear interpolation on the generated action sequence;
[0024] Wherein, assuming that the action sequence includes N groups of action data, performing linear interpolation processing on the generated action sequence includes:
[0025] At least one set of transitional motion data is inserted into every interval of M sets of motion data, where M is a natural number greater than 1 and less than N.
[0026] In one embodiment, before loading and executing the model file, the method further includes:
[0027] Performing a computer test on the model file; the computer test means: executing the model file by a simulation robot and determining whether the action performed by the simulation robot is correct through visual model detection and / or manual review;
[0028] If the computer test passes, save the model file;
[0029] If the on-machine test fails, the model parameters in the pre-trained model are adjusted until the generated model file passes the on-machine test.
[0030] In a second aspect, the present application further provides a robot motion generation device, comprising:
[0031] An action sequence generation module is used to generate an action sequence based on the robot's joint angle data;
[0032] The annotation module is used to annotate the generated action sequence to obtain the action sequence description;
[0033] an action instruction generation module, configured to extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description through the preset language model;
[0034] A model file generation module is used to input the action instructions into a pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data; the model file is used to instruct the robot to complete various actions;
[0035] The execution module is used to load and execute the model file.
[0036] In a third aspect, the present application further provides a robot comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0037] Generate action sequences based on the robot's joint angle data;
[0038] Label the generated action sequence to obtain the action sequence description;
[0039] Extracting verbs from the action sequence description, and inputting the extracted verbs into a preset language model, thereby generating action instructions corresponding to the action sequence description through the preset language model;
[0040] Inputting the action instructions into a pre-trained model, the pre-trained model outputting a model file based on the relationship between the action description corresponding to the action instructions and the motion data; the model file is used to instruct the robot to complete various actions;
[0041] Load and execute the model file.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0043] Generate action sequences based on the robot's joint angle data;
[0044] Label the generated action sequence to obtain the action sequence description;
[0045] Extracting verbs from the action sequence description, and inputting the extracted verbs into a preset language model, thereby generating action instructions corresponding to the action sequence description through the preset language model;
[0046] Inputting the action instructions into a pre-trained model, the pre-trained model outputting a model file based on the relationship between the action description corresponding to the action instructions and the motion data; the model file is used to instruct the robot to complete various actions;
[0047] Load and execute the model file.
[0048] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0049] Generate action sequences based on the robot's joint angle data;
[0050] Label the generated action sequence to obtain the action sequence description;
[0051] Extracting verbs from the action sequence description, and inputting the extracted verbs into a preset language model, thereby generating action instructions corresponding to the action sequence description through the preset language model;
[0052] Inputting the action instructions into a pre-trained model, the pre-trained model outputting a model file based on the relationship between the action description corresponding to the action instructions and the motion data; the model file is used to instruct the robot to complete various actions;
[0053] Load and execute the model file.
[0054] The robot motion generation method, apparatus, robot, computer-readable storage medium, and computer program product described above generate motion sequences based on the robot's joint angle data, thereby converting the robot's joint angle data into a motion sequence containing a time sequence for execution. The generated motion sequence is annotated to obtain a motion sequence description, thereby increasing the description of the motion sequence to facilitate the subsequent generation of more accurate motion instructions using a preset language model. Verbs are extracted from the motion sequence description and input into a preset language model. The preset language model then generates motion instructions corresponding to the motion sequence description, thereby automatically generating a large number of motion instructions and improving the efficiency of motion instruction generation. The motion instructions are input into a pre-trained model, which outputs a model file based on the relationship between the motion description corresponding to the motion instruction and the motion data. The model file is used to instruct the robot to perform various motions. The model file is then loaded and executed. Consequently, a large number of motion instructions can be converted into model files that can be executed by the robot, reducing the computational overhead of robot motion generation and improving motion generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 A schematic diagram of the structure of a robot in one embodiment;
[0057] Figure 2 1 is a flow chart of a method for generating robot motions in one embodiment;
[0058] Figure 3 is a flowchart of a robot action generation method according to another embodiment;
[0059] Figure 4 1 is a flow chart of a method for generating robot motions in another embodiment;
[0060] Figure 5 is a structural block diagram of a robot motion generation device in one embodiment;
[0061] Figure 6 is a structural block diagram of a robot motion generating device in another embodiment;
[0062] Figure 7 is a structural block diagram of a robot motion generating device in yet another embodiment;
[0063] Figure 8 FIG. 4 is a diagram showing the internal structure of a processing system of a robot in one embodiment. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0065] The robot action generation method provided in the embodiment of the present application can be applied to Figure 1 The robot shown in FIG. The robot can be a humanoid robot or other motion robot with movable joints. Figure 1 As shown, the robot may include multiple joints 101 and a visual system 102. The visual system 102 is generally located at the head of the robot and is used to collect images of the external environment. In addition to the visual system 102, a speech processing system, a question-answering knowledge base, etc. may also be configured (in Figure 1 (not shown). The robot includes multiple joints, which are electrically connected to a drive mechanism to achieve a certain range of rotation. When the joint angles of the robot's joints change, the robot can achieve different posture changes. Therefore, the robot can be controlled to perform different actions by controlling the joint angles of each joint.
[0066] For example, when applied to a humanoid robot, it is desirable for the robot to be able to perform various actions according to motion commands, such as imitating human behaviors: waving, raising hands, nodding, shaking head, walking, jumping, etc. In interactive communication scenarios, it is also desirable for the robot to generate voice responses based on user questions, and to adapt the voice responses to appropriate actions, thereby making the robot more human-like and enhancing the interactive experience.
[0067] In an exemplary embodiment, Figure 2 As shown, a robot action generation method is provided, which is applied to Figure 1 The robot in FIG is taken as an example to illustrate the method, which includes the following steps 201 to 205. Among them:
[0068] Step 201: Generate an action sequence based on the joint angle data of the robot.
[0069] In this embodiment, different types of robots may have different numbers of joints. For example, a humanoid robot (which may have 32 joints) may have multiple joints distributed across its head, torso, limbs, and fingers. These joints, driven by their respective drive mechanisms, rotate through a certain angle, thereby generating a series of joint angle data. For example, this joint angle data can be obtained from an open-source database or collected during regular motion capture.
[0070] Exemplarily, the joint angle data of the robot is collected; the joint angle data includes: the joint angle of each joint; the multiple joint angles corresponding to each joint are integrated and processed in the order of execution time, wherein the joint angles corresponding to each joint with the same execution time are integrated into a set of motion data.
[0071] In one possible case, if the joint angles of some joints are missing in a set of action data, the joint angles of the missing joints are searched forward or backward in the order of execution time, and the joint angles with the closest execution time are added to the action data; the multiple sets of action data obtained are sorted in the order of execution time to obtain an action sequence.
[0072] In this example, let's assume the humanoid robot is waving its hands. This requires only the coordinated movements of the right arm and fingers, while the robot's other joints are not required to perform any movements. In this case, only the joint angles of the right arm and fingers may be captured, leaving the angles of the other joints unknown. To maintain the integrity of each group of motion data in the motion sequence, the search can begin with the motion data of the adjacent group, supplementing the missing joint angles with the corresponding joint angles from the temporally closest motion data.
[0073] For simplicity, let's assume a robot arm has joints A, B, and C. During execution time, the angles of joint A change from 45 degrees, 60 degrees, 75 degrees, and 90 degrees; the angles of joint B change from 35 degrees, 50 degrees, 90 degrees, and 105 degrees; and the angles of joint C change from 90 degrees, 75 degrees, 90 degrees, and 105 degrees. Generally speaking, each joint angle change is accompanied by a timestamp, meaning each joint angle has its own timestamp. Therefore, timestamps can be used to align joints, grouping joint angles with the same timestamp into a set of motion data. For example, the first set of motion data is (45 degrees, 35 degrees, 90 degrees), the second set is (60 degrees, 50 degrees, 75 degrees), the third set is (75 degrees, 90 degrees, 90 degrees), and the fourth set is (90 degrees, 105 degrees). These four sets of motion data constitute a motion sequence.
[0074] Step 202: annotate the generated action sequence to obtain an action sequence description.
[0075] In this embodiment, manual annotation and / or automatic annotation can be used to annotate the generated action sequence to obtain an action sequence description; the action sequence description includes: a text for describing the action.
[0076] For example, manual annotation involves staff annotating the action sequence with text. For example, they can use historical annotation records as a reference to annotate each action sequence individually. The annotated text can optionally include verbs, nouns, adjectives, and more. For example, "Raise your right hand and make an OK sign in front of your chest" or "Shake your head left and right three times."
[0077] An exemplary automatic annotation approach involves building a model capable of automatically annotating action sequences. For example, a database of annotated samples is constructed and a classification or prediction model is trained to output the text content category corresponding to the action sequence. Automatic annotation can significantly improve the efficiency of action sequence annotation.
[0078] Step 203 : extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description through the preset language model.
[0079] In this embodiment, verb extraction for an action sequence description involves extracting all verbs from a text based on part of speech. For example, if the action sequence description is "walk forward slowly and wave," verb extraction for this action sequence description yields the verbs "walk" and "wave."
[0080] In this embodiment, the preset language model may be a Large Language Model (LLM), which can generate grammatically and meaningful natural language text by learning from large amounts of text data. For example, the LLM can select action instructions corresponding to the action sequence description from a target action instruction set based on the extracted verbs. The target action instruction set is pre-established and can be constructed using an open source dataset.
[0081] In step 204 , the action instruction is input into the pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data.
[0082] The model file is used to instruct the robot to complete various actions.
[0083] In this embodiment, the pre-trained model can adopt a Transformer model with both encoding and decoding. The encoder-decoder architecture (such as the T5 model and the BART model) consists of two parts: an encoder and a decoder, which are linked and commonly used in machine translation, summary generation, etc. The encoder is responsible for receiving the input sequence (such as sentences and text) and converting it into a series of hidden representations. These hidden representations capture the semantic and contextual information in the input sequence. The decoder gradually generates the target output sequence based on the hidden representation generated by the encoder. The decoder predicts the next word by inputting the previously generated word and the encoder's hidden representation.
[0084] Exemplarily, the pre-trained model includes: an encoder and a decoder; the encoder is used to convert the action instruction into an action description; the decoder is used to predict the motion data corresponding to the action description generated by the encoder, and output a model file based on the predicted motion data; wherein, the model file includes: the joint angles of each joint under a preset number of frames, the coordinates of each joint in a Cartesian coordinate system, and a timestamp; the timestamp is used to determine the execution order corresponding to the joint angles of each joint under a preset number of frames and the coordinates of each joint in a Cartesian coordinate system.
[0085] In this embodiment, a preset language model can automatically generate a large number of action instructions (tokens) based on the action description corresponding to the action sequence, and the generated action instructions can be processed by a pre-trained model to obtain a model file that can be executed by the robot, so that a large number of action instructions can be converted into action data, which simplifies the action generation process and greatly improves the efficiency of action generation.
[0086] Step 205: Load and execute the model file.
[0087] In this embodiment, the robot can directly load the model file for use. After the robot loads the model file, it can choose to execute and then complete the action according to the action sequence in the model file.
[0088] In the above-mentioned robot motion generation method, an action sequence is generated based on the robot's joint angle data; thereby, the robot's joint angle data can be converted into an action sequence containing an execution time sequence. The generated action sequence is annotated to obtain an action sequence description; thereby, the description of the action sequence can be increased, so that more accurate action instructions can be generated through a preset language model in the future. Verbs are extracted from the action sequence description, and the extracted verbs are input into a preset language model. Action instructions corresponding to the action sequence description are generated through the preset language model; thereby, a large number of action instructions can be automatically generated, thereby improving the efficiency of action instruction generation. The action instructions are input into a pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data; the model file is used to instruct the robot to complete various actions; and the model file is loaded and executed. Thus, a large number of action instructions can be converted into model files that can be executed by the robot, reducing the computational overhead of robot action generation and improving action generation efficiency.
[0089] In another exemplary embodiment, Figure 3 As shown, a robot action generation method is provided, which is applied to Figure 1 The robot in the example is used to illustrate the process, including the following steps 301 to 306. Among them:
[0090] Step 301: Generate an action sequence based on the robot's joint angle data.
[0091] In this embodiment, the specific implementation process and technical effects of step 301 are shown in Figure 2 The description of step 201 in the illustrated method embodiment will not be repeated here.
[0092] Step 302: Perform linear interpolation processing on the generated action sequence.
[0093] In this embodiment, sometimes the amplitude of the motions between the motion sequences varies greatly, making the motions performed by the robot stiff. To address this issue, linear interpolation can be performed on the generated motion sequences.
[0094] Exemplarily, assuming that the action sequence includes N groups of action data, at least one group of transition action data may be inserted into every interval of M groups of action data, where M is a natural number greater than 1 and less than N.
[0095] In this embodiment, by performing linear interpolation processing on the generated action sequence, the actions can be made more coherent and smooth, and the robot's action performance can be made smoother and more natural.
[0096] Step 303: annotate the generated action sequence to obtain an action sequence description.
[0097] Step 304 : extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description through the preset language model.
[0098] In step 305 , the action instruction is input into the pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data.
[0099] Step 306: Load and execute the model file.
[0100] In this embodiment, the specific implementation process and technical effects of steps 303 to 306 are shown in Figure 2 The descriptions of steps 202 to 205 in the illustrated method embodiment are not repeated here.
[0101] In another exemplary embodiment, Figure 4 As shown, a robot action generation method is provided, which is applied to Figure 1 The robot in FIG is taken as an example to illustrate the process, including the following steps 401 to 407. Among them:
[0102] Step 401: Generate an action sequence based on the joint angle data of the robot.
[0103] Step 402: annotate the generated action sequence to obtain an action sequence description.
[0104] Step 403 : extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description through the preset language model.
[0105] In step 404, the action instruction is input into the pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data.
[0106] In this embodiment, the specific implementation process and technical effects of steps 401 to 404 are shown in Figure 2 The descriptions of steps 201 to 204 in the illustrated method embodiment are not repeated here.
[0107] Step 405 , perform a computer test on the model file and determine whether the computer test result passes. If so, execute step 406 ; if not, execute step 407 .
[0108] Step 406: Save the model file, load and execute the model file.
[0109] Step 407: Adjust the model parameters in the pre-trained model, and return to step 404.
[0110] In this embodiment, the on-machine test refers to: executing the model file by a simulation robot, and determining whether the action completed by the simulation robot is correct through visual model detection and / or manual review.
[0111] For example, testing a model file on a machine can involve generating an action sequence based on the model file, playing the generated action sequence on a simulation robot, and determining whether the action sequence is correct using a visual model or manual review. If correct, the model file is considered to have passed the test. If not, the model parameters of the pre-trained model are adjusted until the generated model file passes the machine test.
[0112] For example, consider the encoder-decoder model, a deep learning model framework widely used in fields such as natural language processing and computer vision. It typically consists of two parts: an encoder and a decoder. The encoder converts the input sequence into a fixed-length vector, while the decoder converts this vector into an output sequence. The model parameters of the encoder-decoder model may vary across different application scenarios and model architectures.
[0113] For example, in a sequence-to-sequence task, model parameters include: the length of the input sequence and the output sequence, the dimension of the intermediate representation, the batch size, the learning rate, the optimizer, etc.
[0114] For example, many advanced encoding / decoding models, such as neural machine translation models, use an attention mechanism to improve performance. This mechanism allows the decoder to focus on different parts of the input sequence when generating each output. Model parameters can include attention mechanism parameters, such as the calculation method and normalization method for the attention score.
[0115] In this embodiment, the model file is tested on a computer and the pre-trained model is continuously optimized according to the test results, so that the model file generated by the model is more accurate.
[0116] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0117] Based on the same inventive concept, embodiments of the present application also provide a robot motion generation device for implementing the aforementioned robot motion generation method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more embodiments of the robot motion generation device provided below can be found in the above-mentioned limitations on the robot motion generation method and will not be further elaborated here.
[0118] In an exemplary embodiment, Figure 5 As shown, a robot action generation device is provided, comprising: an action sequence generation module 501, a labeling module 502, an action instruction generation module 503, a model file generation module 504 and an execution module 505, wherein:
[0119] An action sequence generation module 501 is used to generate an action sequence according to the joint angle data of the robot;
[0120] Annotation module 502, for annotating the generated action sequence to obtain an action sequence description;
[0121] An action instruction generation module 503 is configured to extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description using the preset language model;
[0122] A model file generation module 504 is configured to input the motion instructions into a pre-trained model, and the pre-trained model outputs a model file based on the relationship between the motion description corresponding to the motion instruction and the motion data; the model file is used to instruct the robot to perform various actions;
[0123] The execution module 505 is used to load and execute the model file.
[0124] Exemplarily, the action sequence generation module 501 is specifically used to: collect joint angle data of the robot; the joint angle data includes: the joint angle of each joint; the multiple joint angles corresponding to each joint are integrated and processed in the order of execution time, wherein the joint angles corresponding to each joint with the same execution time are integrated into a set of action data; if the joint angles of some joints are missing in a set of action data, the joint angles of the missing joints are searched forward or backward in the order of execution time, and the joint angles with the closest execution time are added to the action data; the multiple sets of action data obtained are sorted in the order of execution time to obtain an action sequence.
[0125] Exemplarily, the annotation module 502 is specifically configured to: annotate the generated action sequence by manual annotation and / or automatic annotation to obtain an action sequence description; the action sequence description includes: a text for describing the action.
[0126] Exemplarily, the pre-trained model includes: an encoder and a decoder; the encoder is used to convert the action instruction into an action description; the decoder is used to predict the motion data corresponding to the action description generated by the encoder, and output a model file based on the predicted motion data; wherein, the model file includes: the joint angles of each joint under a preset number of frames, the coordinates of each joint in a Cartesian coordinate system, and a timestamp; the timestamp is used to determine the execution order corresponding to the joint angles of each joint under a preset number of frames and the coordinates of each joint in a Cartesian coordinate system.
[0127] In another exemplary embodiment, Figure 6 As shown, a robot action generation device is provided. Figure 5 On the basis of the device shown, it can also include: an interpolation processing module 506, which is used to perform linear interpolation processing on the generated action sequence; wherein, assuming that the action sequence includes N groups of action data, at least one group of transition action data is inserted in each interval of M groups of action data, and M is a natural number greater than 1 and less than N.
[0128] In another exemplary embodiment, Figure 7 As shown, a robot action generation device is provided. Figure 5On the basis of the device shown, a testing module 507 can also be included, and the testing module 507 is used to perform an on-machine test on the model file; the on-machine test means: the model file is executed by a simulation robot, and whether the action completed by the simulation robot is correct is judged through visual model detection and / or manual review; if the on-machine test passes, the model file is saved; if the on-machine test fails, the model parameters in the pre-trained model are adjusted until the generated model file passes the on-machine test.
[0129] It should be noted that you can also Figure 6 、 Figure 7 The device shown is combined, that is, it includes both the interpolation processing module 506 and the testing module 507.
[0130] Each module in the aforementioned robot motion generation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0131] In an exemplary embodiment, a processing system of a robot is provided, and the internal structure diagram of the processing system can be shown as follows: Figure 8 As shown. The processing system includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the processing system is used to provide computing and control capabilities. The memory of the processing system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the processing system is used to exchange information between the processor and external devices. The communication interface of the processing system is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements a method for generating robot motions. The display unit of the processing system is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the processing system can be a touch layer covering the display screen, or a button, trackball or touchpad set on the robot shell, or an external keyboard, touchpad or mouse.
[0132] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the processing system to which the solution of the present application is applied. The specific processing system may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0133] In an exemplary embodiment, a robot is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0135] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0137] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0138] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0139] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A robot motion generation method, characterized in that: The method comprises: Generate action sequences based on the robot's joint angle data; Label the generated action sequence to obtain the action sequence description; Extracting verbs from the action sequence description, and inputting the extracted verbs into a preset language model, thereby generating action instructions corresponding to the action sequence description through the preset language model; The action instruction is input into a pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data; the model file is used to instruct the robot to complete various actions; the pre-trained model includes: an encoder and a decoder; the encoder is used to convert the action instruction into an action description; the decoder is used to predict the motion data corresponding to the action description generated by the encoder, and output the model file based on the predicted motion data; wherein, the model file includes: the joint angle of each joint in a preset number of frames, the coordinates of each joint in a Cartesian coordinate system, and a timestamp; the timestamp is used to determine the execution order corresponding to the joint angle of each joint and the coordinates of each joint in a Cartesian coordinate system in a preset number of frames; Load and execute the model file.
2. The method according to claim 1, characterized in that The step of generating an action sequence based on the robot's joint angle data includes: Collecting joint angle data of the robot; the joint angle data includes: joint angles of each joint; The multiple joint angles corresponding to each joint are integrated in the order of execution time, wherein the joint angles corresponding to each joint at the same execution time are integrated into a set of motion data; If the joint angles of some joints are missing in a set of motion data, the joint angles of the missing joints are searched forward or backward in the order of execution time, and the joint angles with the closest execution time are added to the motion data; The obtained multiple groups of action data are sorted in the order of execution time to obtain an action sequence.
3. The method according to claim 1, characterized in that The step of labeling the generated action sequence to obtain an action sequence description includes: The generated action sequence is annotated by manual annotation and / or automatic annotation to obtain an action sequence description; the action sequence description includes: a text for describing the action.
4. The method according to any one of claims 1 to 3, characterized in that After generating the action sequence according to the joint angle data of the robot, the method further includes: Perform linear interpolation on the generated action sequence; Wherein, assuming that the action sequence includes N groups of action data, performing linear interpolation processing on the generated action sequence includes: At least one set of transitional motion data is inserted into every interval of M sets of motion data, where M is a natural number greater than 1 and less than N.
5. The method according to any one of claims 1 to 3, characterized in that Before loading and executing the model file, the method further includes: Performing a computer test on the model file; the computer test means: executing the model file by a simulation robot and determining whether the action performed by the simulation robot is correct through visual model detection and / or manual review; If the computer test passes, save the model file; If the on-machine test fails, the model parameters in the pre-trained model are adjusted until the generated model file passes the on-machine test.
6. A robot motion generation device, characterized in that: The device comprises: An action sequence generation module is used to generate an action sequence based on the robot's joint angle data; The annotation module is used to annotate the generated action sequence to obtain the action sequence description; an action instruction generation module, configured to extract verbs from the action sequence description, input the extracted verbs into a preset language model, and generate action instructions corresponding to the action sequence description through the preset language model; A model file generation module is used to input the action instructions into a pre-trained model, and the pre-trained model outputs a model file based on the relationship between the action description corresponding to the action instruction and the motion data; the model file is used to instruct the robot to complete various actions; the pre-trained model includes: an encoder and a decoder; the encoder is used to convert the action instructions into action descriptions; the decoder is used to predict the motion data corresponding to the action description generated by the encoder, and output the model file based on the predicted motion data; wherein, the model file includes: the joint angles of each joint in a preset number of frames, the coordinates of each joint in a Cartesian coordinate system, and a timestamp; the timestamp is used to determine the execution order corresponding to the joint angles of each joint in a preset number of frames and the coordinates of each joint in a Cartesian coordinate system; The execution module is used to load and execute the model file.
7. The device according to claim 6, characterized in that The device also includes an interpolation processing module, which is used to perform linear interpolation processing on the generated action sequence; wherein, assuming that the action sequence includes N groups of action data, at least one group of transition action data is inserted in every interval of M groups of action data, where M is a natural number greater than 1 and less than N.
8. A robot comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method for generating action sequence for driving virtual character to move according to text
CN116883555A
Natural language instruction analysis and operation mapping method for hot-line work robot, electronic equipment and storage medium
CN119691174A