Action generation method, device, robot and readable storage medium

The vector quantization variational autoencoder processes the robot joint angle data, generates and aligns the action sequence and description, reconstructs the motion data and generates instructions, solves the problem of large size and low generation efficiency of the robot's action data, and realizes efficient action generation and control.

CN120002671BActive Publication Date: 2025-08-15SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510474646.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the prior art, the original action data in the robot's action generation method is huge in size and difficult to efficiently encode, store and reconstruct, resulting in low generation efficiency.

Method used

A vector quantization variational autoencoder is used to process the joint angle data of the robot, generate action sequences and descriptions, and reconstruct the original motion data through alignment processing, generate action instructions, and input the trained action model to instruct the robot to complete the action.

Benefits of technology

Through quantitative encoding, the volume of the original action data is reduced, the storage and transmission efficiency is optimized, and the robot action generation efficiency is improved. It is suitable for robot action control in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120002671B_ABST
    Figure CN120002671B_ABST
Patent Text Reader

Abstract

The present application relates to an action generation method, device, robot, and readable storage medium. The method comprises: obtaining joint angle data of the robot; inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description; aligning the action sequence and the action description so that each action sequence corresponds to an action description; reconstructing the original motion data by the vector quantized variational autoencoder based on the aligned action sequence and action description, and generating action instructions based on the original motion data; inputting the action instructions into a trained action model, and having the trained action model output a model file, wherein the model file is used to instruct the robot to perform various actions. This enables the action data to be quantized and encoded, reducing the volume of the original action data, improving the efficiency of the robot action generation, and having strong applicability, making it easier for the robot to complete more complex actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an action generation method, device, robot, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of artificial intelligence (AI), robots are increasingly being used, particularly in simulation robots, which can perform various anthropomorphic actions according to motion instructions. A robot motion instruction is the angle change of each joint along the robot's route from its starting point to its destination. This represents the position of each joint in a time series.

[0003] Traditionally, robot motion commands are typically generated through reinforcement learning or imitation learning. Reinforcement learning is a method that autonomously interacts with the environment based on rewards and learns to generate motion commands. Imitation learning is a supervised learning approach whose core concept is to output motion commands consistent with the taught motion. Imitation learning methods first require collecting human-demonstrated teaching data and training a diffusion model. During the inference phase, the current observation image and the robot's joint poses are input, and a denoising method is used to output the robot's next action.

[0004] However, in the above-mentioned robot motion generation method, the original motion data is huge and difficult to encode, store and reconstruct in an efficient way, resulting in low efficiency of robot motion generation. Summary of the Invention

[0005] Based on this, it is necessary to provide a motion generation method, device, robot, computer-readable storage medium and computer program product that can quantize and encode motion data, reduce the volume of original motion data, and improve the efficiency of robot motion generation in response to the above technical problems.

[0006] In a first aspect, the present application provides an action generation method, comprising:

[0007] Get the robot's joint angle data;

[0008] Inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0009] Aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0010] According to the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs the original motion data, and generates action instructions based on the original motion data;

[0011] The action instructions are input into a trained action model, and the trained action model outputs a model file, which is used to instruct the robot to complete various actions.

[0012] In one embodiment, obtaining the robot's joint angle data includes:

[0013] Collect the robot's joint motion data through the motion capture system;

[0014] Splitting the joint motion data using a pre-built motion semantic segmentation model to obtain multiple motion data segments;

[0015] The motion data segments are redirected, and the redirected motion data segments are mapped to the joint angles of the robot to obtain the joint angle data of the robot.

[0016] In one embodiment, obtaining the robot's joint angle data includes:

[0017] Extracting bone hotspots from the robot motion video collected by the sensor, and redirecting the extracted bone hotspots to the robot's joint angles to obtain the robot's joint angle data; or

[0018] The robot's joint angle data can be directly obtained through virtual reality remote control operation.

[0019] In one embodiment, aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description includes:

[0020] Respectively obtaining a timestamp corresponding to the action sequence and a timestamp corresponding to the action description;

[0021] sorting the action sequence and the action description according to the time order indicated by the timestamp;

[0022] The action sequences and action descriptions corresponding to the same timestamp are aligned so that each action sequence corresponds to an action description.

[0023] In one embodiment, reconstructing original motion data by the vector quantized variational autoencoder according to the aligned motion sequence and motion description, and generating motion instructions based on the original motion data include:

[0024] constructing a motion vocabulary table according to the motion description by the vector quantized variational autoencoder, wherein each word in the motion vocabulary table corresponds to a motion feature vector;

[0025] Encoding the motion sequence by the vector quantized variational autoencoder to generate a continuous latent representation, and mapping the continuous latent representation to a motion feature vector in the motion vocabulary to obtain a motion label corresponding to the motion sequence;

[0026] The vector quantization variational autoencoder decodes the motion marker to reconstruct the original motion data, and generates a series of action instructions based on the original motion data.

[0027] In one embodiment, before inputting the motion instruction into the trained motion model and outputting the model file from the trained motion model, the method further includes:

[0028] Construct an initial motion model;

[0029] Constructing a training data set based on the motion instructions generated by the vector quantized variational autoencoder; wherein the motion labels corresponding to the motion instructions are determined by the vector quantized variational autoencoder in the process of generating the motion instructions;

[0030] Training the initial action model using the training data set, and evaluating the actions generated by the initial action model using a video language basic model to obtain an evaluation result;

[0031] The initial motion model is continuously optimized according to the evaluation results to obtain a trained motion model.

[0032] In one embodiment, before outputting the model file from the trained motion model, the method further includes:

[0033] In the trained action model, displaying the action generated according to the action instruction;

[0034] Determining whether the action generated according to the action instruction is correct through manual review and / or automatic review;

[0035] If it is incorrect, the model parameters of the trained action model are adjusted until the action generated according to the action instruction is correct.

[0036] In a second aspect, the present application further provides an action generation device, comprising:

[0037] Acquisition module, used to obtain the robot's joint angle data;

[0038] a processing module, configured to input the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0039] an alignment module, configured to align the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0040] an action instruction generation module, configured to reconstruct original motion data from the vector quantized variational autoencoder according to the aligned action sequence and action description, and generate action instructions based on the original motion data;

[0041] The model file generation module is used to input the action instructions into the trained action model, and the trained action model outputs a model file, and the model file is used to instruct the robot to complete various actions.

[0042] In a third aspect, the present application further provides a robot comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0043] Get the robot's joint angle data;

[0044] Inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0045] Aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0046] According to the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs the original motion data, and generates action instructions based on the original motion data;

[0047] The action instructions are input into a trained action model, and the trained action model outputs a model file, which is used to instruct the robot to complete various actions.

[0048] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0049] Get the robot's joint angle data;

[0050] Inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0051] Aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0052] According to the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs the original motion data, and generates action instructions based on the original motion data;

[0053] The action instructions are input into a trained action model, and the trained action model outputs a model file, which is used to instruct the robot to complete various actions.

[0054] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0055] Get the robot's joint angle data;

[0056] Inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0057] Aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0058] According to the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs the original motion data, and generates action instructions based on the original motion data;

[0059] The action instructions are input into a trained action model, and the trained action model outputs a model file, which is used to instruct the robot to complete various actions.

[0060] The above-mentioned motion generation method, device, robot, computer-readable storage medium, and computer program product obtain robot joint angle data; input the robot joint angle data into a vector quantized variational autoencoder for processing to obtain motion sequences and motion descriptions. The vector quantized variational autoencoder can then quantize and encode the joint angle data, reducing the volume of the original motion data and optimizing storage and transmission efficiency. The motion sequences and motion descriptions are then aligned so that each motion sequence corresponds to an action description. This allows for a one-to-one correspondence between the motion sequences and action descriptions, facilitating the decomposition of the original motion data into a series of motion instructions during the actual decoding process. Based on the aligned motion sequences and action descriptions, the vector quantized variational autoencoder reconstructs the original motion data and generates motion instructions based on the original motion data. The vector quantized variational autoencoder can then quickly convert motion data collected from various sources into minimum instructions that the robot can execute, improving the efficiency of robot motion instruction generation. The motion instructions are then input into a trained motion model, which then outputs a model file that instructs the robot to perform various actions. This enables quantitative encoding of motion data, reduces the volume of original motion data, improves the efficiency of robot motion generation, is suitable for robot motion control in various scenarios, and facilitates robots to complete more complex movements. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 A schematic diagram of the structure of a robot in one embodiment;

[0063] Figure 2 1 is a flow chart of an action generation method in one embodiment;

[0064] Figure 3 A schematic diagram of posture changes of an action model in one embodiment when it completes three action instructions continuously;

[0065] Figure 4 is a flowchart of an action generation method in another embodiment;

[0066] Figure 5 is a flowchart of an action generation method in yet another embodiment;

[0067] Figure 6 is a structural block diagram of an action generating device in one embodiment;

[0068] Figure 7 is a structural block diagram of an action generating device in another embodiment;

[0069] Figure 8 is a structural block diagram of an action generating device in yet another embodiment;

[0070] Figure 9 FIG. 4 is a diagram showing the internal structure of a processing system of a robot in one embodiment. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0072] To facilitate understanding of the technical solutions of various embodiments in this application, the following brief description of the professional terms appearing in this application is first made:

[0073] 1) An autoencoder is a neural network architecture consisting of two main components: an encoder and a decoder. The encoder compresses input data into a representation in a latent space, while the decoder reconstructs the original data from this representation.

[0074] 2) The Variational Autoencoder (VAE) introduces the concept of probabilistic modeling based on the autoencoder. By modeling the representation of the latent space as a probability distribution (usually a Gaussian distribution), it enables the generated new samples to be more diverse.

[0075] 3) Vector quantization (VQ) is the process of mapping continuous values to discrete values. In the Vector Quantized Variational AutoEncoder (VQ-VAE), vector quantization is used to replace the continuous representation in the latent space. This allows each point in the latent space to be mapped to a vector in a fixed discrete "vocabulary."

[0076] For example, Figure 1 The schematic diagram of the structure of the robot in one embodiment is shown. The motion generation method provided in the embodiment of the present application can be applied to the following examples: Figure 1 The robot shown in FIG. The robot can be a humanoid robot or other motion robot with movable joints. Figure 1 As shown, the robot may include: a visual system 101 and multiple joints 102 (such as arm joints, leg joints, head joints, etc.). The visual system 101 is generally located at the head of the robot and is used to collect images of the external environment. In addition to the visual system 101, the robot may also be equipped with a voice processing system, a question-answering knowledge base, etc. (in Figure 1 (not shown). The robot's multiple joints 102 are electrically connected to independently provided drive mechanisms to achieve a certain angle of rotation. When the rotation angles of the robot's joints change, the robot can achieve different posture changes. Accordingly, the robot can be controlled to perform different actions by controlling the joint angles of each joint.

[0077] For example, in humanoid robot applications, it's often desirable for the robot to be able to perform various actions based on motion commands, such as mimicking human behaviors: waving, raising hands, nodding, shaking heads, walking, jumping, and so on. In interactive communication scenarios, it's also desirable for the robot to generate spoken responses based on user questions, and to align the responses with appropriate actions, making the robot more human-like and enhancing the interactive experience.

[0078] In an exemplary embodiment, Figure 2 As shown, an action generation method is provided, which is applied to Figure 1 The robot in FIG is taken as an example to illustrate the method, which includes the following steps 201 to 205. Among them:

[0079] Step 201: Acquire the joint angle data of the robot.

[0080] In this embodiment, the robot's joint angle data may include: the joint angles of each joint, the coordinates of each joint in a Cartesian coordinate system, and a timestamp. The timestamp is a marker used to identify the order in which joints change. For example, in addition to the timestamp, the order in which joint angles change can be identified by encoding.

[0081] In an optional embodiment, the joint motion data of the robot can be collected through a motion capture system; the joint motion data is split through a pre-built motion semantic segmentation model to obtain multiple motion data segments; the motion data segments are redirected, and the redirected motion data segments are mapped to the joint angles of the robot to obtain the joint angle data of the robot.

[0082] In this embodiment, the motion capture system is a device that can accurately measure and record the motion trajectory and posture of an object during operation in real time. Therefore, the motion data of the robot can be collected through the motion capture system.

[0083] In this embodiment, the motion semantic segmentation model is a computer vision model designed to assign each pixel in an image to a specific semantic category while taking motion into account. This model not only provides fine-grained classification information for each pixel in the image, but also captures and analyzes motion patterns within the image, enabling more efficient and accurate segmentation in practical applications.

[0084] In this embodiment, redirecting the motion data segment refers to converting the information contained in the motion data segment into robot joint angles. For example, if the motion data segment includes a Cartesian coordinate change of joint A from M to N, redirecting the motion data segment converts this coordinate change into a joint angle change.

[0085] In another optional embodiment, when a motion video of a robot is obtained, bone hotspots can be extracted from the robot motion video collected by the sensor, and the extracted bone hotspots can be redirected to the joint angles of the robot to obtain the joint angle data of the robot.

[0086] In another optional embodiment, the robot's joint angle data is directly acquired through virtual reality remote control operation. In this case, no redirection process is required and step 202 is directly executed.

[0087] Step 202: Input the robot's joint angle data into a vector quantized variational autoencoder for processing to obtain an action sequence and action description.

[0088] In this embodiment, the encoder in the vector quantized variational autoencoder generates an action feature sequence based on the robot's joint angle data. The quantizer performs quantization processing on the generated action feature sequence as samples to obtain a quantized action feature sequence. The decoder then reconstructs the samples based on the quantized action feature sequence to obtain an action sequence and action description. The action description can be a text describing the action sequence.

[0089] Step 203: align the action sequences and action descriptions so that each action sequence corresponds to an action description.

[0090] In this embodiment, the vector quantized variational autoencoder outputs a large number of action sequences and action descriptions. At this time, there is no correspondence between the two. In order to generate more accurate action instructions later, it is necessary to find the action description corresponding to the action sequence.

[0091] Exemplarily, the timestamp corresponding to the action sequence and the timestamp corresponding to the action description are obtained respectively; the action sequence and the action description are sorted according to the time order indicated by the timestamps; and the action sequence and action description corresponding to the same timestamp are aligned so that each action sequence corresponds to an action description.

[0092] For example, a classification model can be used to predict the action description corresponding to an action sequence, and then a correspondence between the action sequence and the action description can be established. The classification model can include multiple classifiers, which are trained using a constructed dataset (action sequences with pre-labeled action descriptions), so that the trained classifiers can predict the action description corresponding to the input action sequence. Classifiers can include, but are not limited to, decision tree classifiers, support vector machine classifiers, naive Bayes classifiers, K-nearest neighbors classifiers, random forest classifiers, and neural network classifiers.

[0093] In step 204 , the vector quantized variational autoencoder is used to reconstruct the original motion data according to the aligned motion sequence and the motion description, and an action instruction is generated based on the original motion data.

[0094] In this embodiment, the vector quantized variational autoencoder works as follows: First, a vector vocabulary of motion features is learned, much like words in a natural language. Each word corresponds to a unique motion pattern. The input motion data is processed by the encoder, generating a continuous latent representation. These latent representations are then mapped to the nearest discrete vector (i.e., a vector in the motion vocabulary). During the quantization phase, the continuous vector output by the encoder is mapped to a vector in the discrete vocabulary. In this way, each motion data is converted into a motion "tag." The quantized motion tags are then fed into the decoder, which reconstructs the original motion data from these discrete tags.

[0095] Exemplarily, the vector quantized variational autoencoder constructs a motion vocabulary based on the action description, wherein each vocabulary in the motion vocabulary corresponds to a motion feature vector; the vector quantized variational autoencoder encodes the action sequence to generate a continuous latent representation, and maps the continuous latent representation to the motion feature vector in the motion vocabulary to obtain a motion tag corresponding to the action sequence; the vector quantized variational autoencoder decodes the motion tag to reconstruct the original motion data, and generates a series of action instructions based on the original motion data.

[0096] Step 205: input the motion instruction into the trained motion model, and the trained motion model outputs a model file.

[0097] The model file is used to instruct the robot to complete various actions.

[0098] In this embodiment, the action model may be a robot model loaded in computer software. When an action instruction is input, the robot model may be controlled to generate corresponding actions according to the action instruction.

[0099] For example, Figure 3 The figure shows the action model's posture changes as it completes three action commands in succession: first, the first action command (walk), then the second action command (sit), and finally the third action command (stand). When all three action commands are completed correctly, the action model can output a model file, which is a code that the robot can recognize and execute to control the robot to complete the actions indicated by the action commands.

[0100] In the above-mentioned action generation method, the robot's joint angle data is obtained; the robot's joint angle data is input into a vector quantized variational autoencoder for processing to obtain an action sequence and action description. The vector quantized variational autoencoder can then quantize and encode the joint angle data, reducing the volume of the original action data and optimizing storage and transmission efficiency. The action sequence and the action description are aligned so that each action sequence corresponds to an action description. This allows for a one-to-one correspondence between the action sequence and the action description, facilitating the decomposition of the original motion data into a series of action instructions during the actual decoding process. Based on the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs the original motion data and generates action instructions based on the original motion data. The vector quantized variational autoencoder can then quickly convert motion data collected from various channels into the minimum instructions that the robot can execute, thereby improving the efficiency of robot action instruction generation. The action instructions are input into a trained action model, which then outputs a model file that instructs the robot to perform various actions. This enables quantitative encoding of motion data, reduces the volume of original motion data, improves the efficiency of robot motion generation, is suitable for robot motion control in various scenarios, and facilitates robots to complete more complex movements.

[0101] In another exemplary embodiment, Figure 4 As shown, an action generation method is provided, which is applied to Figure 1 The robot in FIG is taken as an example to illustrate the process, including the following steps 401 to 409. Among them:

[0102] Step 401: Acquire the joint angle data of the robot.

[0103] Step 402: Input the robot's joint angle data into a vector quantized variational autoencoder for processing to obtain an action sequence and action description.

[0104] Step 403: align the action sequences and action descriptions so that each action sequence corresponds to an action description.

[0105] Step 404 : Reconstruct the original motion data using a vector quantized variational autoencoder according to the aligned motion sequence and motion description, and generate motion instructions based on the original motion data.

[0106] For the specific implementation process and technical effects of steps 401 to 404 in this embodiment, please refer to Figure 2 The descriptions of steps 201 to 204 in the illustrated method embodiment are not repeated here.

[0107] Step 405: construct an initial motion model.

[0108] In this embodiment, an initial motion model can be constructed in a computer. For example, different types of robots correspond to different motion models, and the motion model includes the same joint structure as a real robot and can demonstrate the same posture changes as a real robot.

[0109] Step 406: construct a training data set based on the action instructions generated by the vector quantized variational autoencoder.

[0110] In this embodiment, see Figure 2 In the specific implementation process of step 204 in the embodiment shown, the vector quantized variational autoencoder constructs a motion vocabulary based on the action description, wherein each vocabulary in the motion vocabulary corresponds to a motion feature vector; the vector quantized variational autoencoder encodes the action sequence to generate a continuous potential representation, and maps the continuous potential representation to the motion feature vector in the motion vocabulary to obtain the motion tag corresponding to the action sequence. Accordingly, a large number of action sequences with motion tags can be obtained, and after these action sequences are converted into action instructions, they also carry the motion tags corresponding to the action sequences synchronously. Therefore, a training data set can be constructed by the action instructions generated by the vector quantized variational autoencoder without the need to label the training data set, thereby improving the efficiency of action model training.

[0111] Step 407 : training the initial action model using the training data set, and evaluating the actions generated by the initial action model using the video language basic model to obtain an evaluation result.

[0112] In this embodiment, the video language base model (VLM base) can be combined to improve the training process of the action model. For example, video data of the action model executing an action instruction is obtained, and the VLM base processes the video data to output the action instruction execution result (e.g., accurate, inaccurate, error, etc.).

[0113] Step 408: Continuously optimize the initial motion model based on the evaluation results to obtain a trained motion model.

[0114] In this embodiment, the action model can be optimized based on the execution results of the action instructions, thereby improving the accuracy of the action model in generating large action codes.

[0115] Step 409: input the motion instruction into the trained motion model, and the trained motion model outputs a model file.

[0116] For the specific implementation process and technical effects of step 409 in this embodiment, please refer to Figure 2The description of step 205 in the illustrated method embodiment will not be repeated here.

[0117] In another exemplary embodiment, Figure 5 As shown, an action generation method is provided, which is applied to Figure 1 The robot in the example is used to illustrate the process, including the following steps 501 to 508. Among them:

[0118] Step 501: Acquire the joint angle data of the robot.

[0119] Step 502: Input the robot's joint angle data into a vector quantized variational autoencoder for processing to obtain an action sequence and action description.

[0120] Step 503: align the action sequences and action descriptions so that each action sequence corresponds to an action description.

[0121] Step 504 : Reconstruct the original motion data using a vector quantized variational autoencoder according to the aligned motion sequence and the motion description, and generate motion instructions based on the original motion data.

[0122] For the specific implementation process and technical effects of steps 501 to 504 in this embodiment, please refer to Figure 2 The descriptions of steps 201 to 204 in the illustrated method embodiment are not repeated here.

[0123] Step 505: input the action instruction into the trained action model, and display the action generated according to the action instruction in the trained action model.

[0124] In this embodiment, the motion instructions are executed and displayed through the trained motion model, so that the robot can preview the situation of completing the motion instructions in advance.

[0125] Step 506 , through manual review and / or automatic review, determine whether the action generated according to the action instruction is correct; if so, execute step 507 , if not, execute step 508 .

[0126] For example, in a manual review method, a staff member directly observes the action model's actions when executing the action instructions to determine whether the actions are correct.

[0127] For example, the automatic review method can be described in step 407 using the video language base model (VLM base). The VLM base evaluates the action model's execution of the action instruction. For example, video data of the action model executing the action instruction is obtained, processed by the VLM base, and the action instruction execution result (e.g., accurate, inaccurate, error, etc.) is output.

[0128] Step 507: Output a model file from the trained motion model.

[0129] Step 508 , adjust the model parameters of the trained motion model, and return to step 505 .

[0130] In this embodiment, if a trained motion model incorrectly executes motion instructions, the model parameters of the trained motion model can be adjusted. The model parameters are determined by the type of motion model itself. For example, they may include the weights, biases, learning rates, initial values, and the like of the neural network.

[0131] In this embodiment, before outputting the model file, the action instructions are first previewed by the trained action model, which can improve the accuracy of the model file so that when the real robot subsequently executes the model file, the corresponding action can be accurately generated.

[0132] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0133] Based on the same inventive concept, the present application also provides an action generation device for implementing the aforementioned action generation method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following action generation device embodiments can be found in the above-mentioned limitations on the action generation method and will not be further elaborated here.

[0134] In an exemplary embodiment, Figure 6 As shown, an action generation device is provided, comprising: an acquisition module 601, a processing module 602, an alignment module 603, an action instruction generation module 604 and a model file generation module 605, wherein:

[0135] An acquisition module 601 is used to acquire joint angle data of the robot;

[0136] A processing module 602 is configured to input the robot's joint angle data into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description;

[0137] an alignment module 603, configured to align the action sequences and the action descriptions so that each action sequence corresponds to an action description;

[0138] An action instruction generation module 604 is configured to reconstruct original motion data from the vector quantized variational autoencoder according to the aligned action sequence and action description, and generate action instructions based on the original motion data;

[0139] The model file generation module 605 is used to input the action instructions into the trained action model, and the trained action model outputs a model file, and the model file is used to instruct the robot to complete various actions.

[0140] Exemplarily, the acquisition module 601 is specifically used to: collect the joint motion data of the robot through a motion capture system; split the joint motion data through a pre-built motion semantic segmentation model to obtain multiple motion data segments; redirect the motion data segments, and map the redirected motion data segments to the joint angles of the robot to obtain the joint angle data of the robot.

[0141] Exemplarily, the acquisition module 601 is specifically used to: extract bone hotspots from the robot motion video collected by the sensor, and redirect the extracted bone hotspots to the joint angles of the robot to obtain the joint angle data of the robot; or directly obtain the joint angle data of the robot through virtual reality remote control operation.

[0142] Exemplarily, the alignment module 603 is specifically used to: respectively obtain the timestamp corresponding to the action sequence and the timestamp corresponding to the action description; sort the action sequence and the action description according to the time order indicated by the timestamp; align the action sequence and action description corresponding to the same timestamp so that each action sequence corresponds to an action description.

[0143] Exemplarily, the action instruction generation module 604 is specifically used to: construct a motion vocabulary according to the action description by the vector quantization variational autoencoder, wherein each vocabulary in the motion vocabulary corresponds to a motion feature vector; encode the action sequence by the vector quantization variational autoencoder to generate a continuous latent representation, and map the continuous latent representation to the motion feature vector in the motion vocabulary to obtain the motion label corresponding to the action sequence; decode the motion label by the vector quantization variational autoencoder to reconstruct the original motion data, and generate a series of action instructions based on the original motion data.

[0144] In another exemplary embodiment, Figure 7 As shown, an action generating device is provided. Figure 6 On the basis of the device shown, it can also include: a model construction and training module 606, which is used to: construct an initial action model; construct a training data set based on the action instructions generated by the vector quantization variational autoencoder; wherein the motion label corresponding to the action instruction is determined by the vector quantization variational autoencoder in the process of generating the action instruction; train the initial action model through the training data set, and evaluate the action generated by the initial action model through the video language basic model to obtain an evaluation result; continuously optimize the initial action model according to the evaluation result to obtain a trained action model.

[0145] In another exemplary embodiment, Figure 8 As shown, an action generating device is provided. Figure 6 On the basis of the device shown, it can also include: an adjustment module 607, and the adjustment module 607 is used to: display the action generated according to the action instruction in the trained action model; determine whether the action generated according to the action instruction is correct through manual review and / or automatic review; if it is incorrect, adjust the model parameters of the trained action model until the action generated according to the action instruction is correct.

[0146] Each module in the above-mentioned action generation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0147] In an exemplary embodiment, a processing system of a robot is provided, and the internal structure diagram of the processing system can be shown as follows: Figure 9As shown. The processing system includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the processing system is used to provide computing and control capabilities. The memory of the processing system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the processing system is used to exchange information between the processor and external devices. The communication interface of the processing system is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements a method for generating robot motions. The display unit of the processing system is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the processing system can be a touch layer covering the display screen, or a button, trackball or touchpad set on the robot shell, or an external keyboard, touchpad or mouse.

[0148] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0149] In an exemplary embodiment, a robot is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0150] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0151] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0153] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0154] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0155] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An action generation method, characterized in that: The method comprises: Get the robot's joint angle data; Inputting the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description; Aligning the action sequences and the action descriptions so that each action sequence corresponds to an action description; According to the aligned action sequence and action description, the vector quantized variational autoencoder reconstructs original motion data, and generates action instructions based on the original motion data, including: constructing a motion vocabulary table according to the action description by the vector quantized variational autoencoder, wherein each word in the motion vocabulary table corresponds to a motion feature vector; encoding the action sequence by the vector quantized variational autoencoder to generate a continuous latent representation, and mapping the continuous latent representation to the motion feature vector in the motion vocabulary table to obtain a motion tag corresponding to the action sequence; decoding the motion tag by the vector quantized variational autoencoder to reconstruct the original motion data, and generating a series of action instructions based on the original motion data; The action instructions are input into a trained action model, and the trained action model outputs a model file, which is used to instruct the robot to complete various actions.

2. The method according to claim 1, characterized in that The obtaining of the robot's joint angle data includes: Collect the robot's joint motion data through the motion capture system; Splitting the joint motion data using a pre-built motion semantic segmentation model to obtain multiple motion data segments; The motion data segments are redirected, and the redirected motion data segments are mapped to the joint angles of the robot to obtain the joint angle data of the robot.

3. The method according to claim 1, characterized in that The obtaining of the robot's joint angle data includes: Extracting bone hotspots from the robot motion video collected by the sensor, and redirecting the extracted bone hotspots to the robot's joint angles to obtain the robot's joint angle data; or The robot's joint angle data can be directly obtained through virtual reality remote control operation.

4. The method according to claim 1, wherein The aligning of the action sequences and the action descriptions so that each action sequence corresponds to an action description includes: Respectively obtaining a timestamp corresponding to the action sequence and a timestamp corresponding to the action description; sorting the action sequence and the action description according to the time order indicated by the timestamp; The action sequences and action descriptions corresponding to the same timestamp are aligned so that each action sequence corresponds to an action description.

5. The method according to claim 1, wherein Before inputting the motion instruction into the trained motion model and outputting the model file from the trained motion model, the method further includes: Construct an initial motion model; Constructing a training data set based on the motion instructions generated by the vector quantized variational autoencoder; wherein the motion labels corresponding to the motion instructions are determined by the vector quantized variational autoencoder in the process of generating the motion instructions; Training the initial action model using the training data set, and evaluating the actions generated by the initial action model using a video language basic model to obtain an evaluation result; The initial motion model is continuously optimized according to the evaluation results to obtain a trained motion model.

6. The method according to any one of claims 1 to 4, characterized in that Before outputting the model file from the trained motion model, the method further includes: In the trained action model, displaying the action generated according to the action instruction; Determining whether the action generated according to the action instruction is correct through manual review and / or automatic review; If it is incorrect, the model parameters of the trained action model are adjusted until the action generated according to the action instruction is correct.

7. An action generating device, characterized in that: The device comprises: Acquisition module, used to obtain the robot's joint angle data; A processing module, configured to input the joint angle data of the robot into a vector quantized variational autoencoder for processing to obtain an action sequence and an action description; an alignment module, configured to align the action sequences and the action descriptions so that each action sequence corresponds to an action description; An action instruction generation module is configured to reconstruct original motion data using the vector quantized variational autoencoder based on the aligned action sequence and action description, and generate action instructions based on the original motion data; specifically, the module is configured to: construct a motion vocabulary using the vector quantized variational autoencoder based on the action description, wherein each vocabulary in the motion vocabulary corresponds to a motion feature vector; encode the action sequence using the vector quantized variational autoencoder to generate a continuous latent representation, and map the continuous latent representation to the motion feature vector in the motion vocabulary to obtain a motion tag corresponding to the action sequence; decode the motion tag using the vector quantized variational autoencoder to reconstruct the original motion data, and generate a series of action instructions based on the original motion data; The model file generation module is used to input the action instructions into the trained action model, and the trained action model outputs a model file, and the model file is used to instruct the robot to complete various actions.

8. The device according to claim 7, characterized in that Also includes: A model building and training module, wherein the model building and training module is used to: build an initial motion model; A training data set is constructed based on the action instructions generated by the vector quantized variational autoencoder; wherein, the motion labels corresponding to the action instructions are determined by the vector quantized variational autoencoder in the process of generating the action instructions; the initial action model is trained using the training data set, and the actions generated by the initial action model are evaluated using a video language basic model to obtain an evaluation result; the initial action model is continuously optimized according to the evaluation result to obtain a trained action model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method for generating action sequence for driving virtual character to move according to text

    CN116883555A

  • Digital human action generation method based on text driving

    CN119579743A