Action sequence generation method and device, electronic equipment and storage medium
By using a diffusion model and a motion encoder generation model, the problem of unexpected performance in motion sequence reconstruction by machine learning algorithms is solved, achieving efficient and accurate motion sequence generation, which is applicable to fields such as virtual reality and robot control.
Patent Information
- Application Number
- CN202510942578.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing machine learning algorithms generate action sequences that do not meet expectations when processing complex instruction text, resulting in unsatisfactory reconstruction results.
A diffusion model is used to construct an action generation model. By extracting action instructions from the target instruction text, a sequence of process actions is generated and then concatenated based on the execution order of the actions in the target instruction text. Combined with a motion encoder and a local action generation model, a target action sequence that conforms to the real trajectory and is accurate is generated.
It enables efficient and accurate generation of complex action commands, applicable to fields such as virtual reality and robot control, improving user experience and operational precision.
Smart Images

Figure CN120429017B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a motion sequence generation method and device, electronic equipment and storage medium. BACKGROUND
[0002] In the fields of virtual reality and robot control, motion generation technology plays a crucial role. Application scenarios in these fields often require highly realistic and smooth motion performance to improve user experience or achieve precise robot operation. Traditional motion generation methods often rely on manual design and adjustment, which is inefficient and costly. With the development of artificial intelligence technology, it has gradually begun to explore the use of machine learning algorithms to automatically generate motion to improve efficiency and reduce costs.
[0003] However, existing machine learning algorithms often struggle when dealing with complex instruction texts, resulting in motion sequences that do not meet expectations, and thus the reconstruction effect of the motion sequence is not ideal. SUMMARY
[0004] The present application provides a motion sequence generation method, device, electronic equipment and storage medium to solve the defect that the motion does not meet the expectation when reconstructing the motion sequence based on the text in the prior art.
[0005] The present application provides a motion sequence generation method, comprising:
[0006] obtaining a target instruction text;
[0007] extracting motion instructions in the target instruction text, the motion instructions including movement instructions and / or local motion instructions;
[0008] inputting the motion instructions into a motion generation model to obtain a process motion sequence output by the motion generation model; the motion generation model is constructed based on a diffusion model;
[0009] splicing the process motion sequence according to the motion execution order of the motion instructions in the target instruction text to obtain a target motion sequence.
[0010] According to the motion sequence generation method provided by the present application, the motion generation model includes a movement motion generation model; the motion instructions include the movement instructions;
[0011] The inputting of the motion instructions into the motion generation model to obtain the process motion sequence output by the motion generation model comprises:
[0012] obtaining a movement motion noise sequence;
[0013] determining a target position corresponding to the movement motion instructions;
[0014] inputting the mobile action noise sequence, the target position and the mobile action instruction into the mobile action generation model to obtain a process action sequence corresponding to the mobile action instruction output by the mobile action generation model;
[0015] The target position is defined based on a three-dimensional coordinate system created by a pelvic node of a target mobile object.
[0016] According to the action sequence generation method provided by the application, the mobile action generation model comprises a motion encoder.
[0017] The inputting the mobile action noise sequence, the target position and the mobile action instruction into the mobile action generation model to obtain a process action sequence corresponding to the mobile action instruction output by the mobile action generation model comprises:
[0018] The target position and the mobile action instruction are encoded based on the motion encoder to obtain motion constraint features.
[0019] Based on the mobile action generation model, the mobile action noise sequence and the motion constraint features are applied to obtain a process action sequence corresponding to the mobile action instruction.
[0020] According to the action sequence generation method provided by the application, the action generation model further comprises a local action generation model; and the action instruction comprises a local action instruction.
[0021] The inputting the action instruction into the action generation model to obtain a process action sequence output by the action generation model comprises:
[0022] A local action noise sequence is determined.
[0023] The local action noise sequence and the local action instruction are inputted into the local action generation model to obtain a process action sequence corresponding to the local action instruction output by the local action generation model.
[0024] According to the action sequence generation method provided by the application, the process action sequence is spliced according to the action execution order of the target instruction text to obtain a target action sequence.
[0025] The process action sequence is processed based on an interpolation function to obtain a smooth action sequence.
[0026] The smooth action sequence is spliced according to the action execution order to obtain a target action sequence.
[0027] The application further provides an action sequence generation device, comprising:
[0028] An acquisition unit acquires a target instruction text;
[0029] An instruction splitting unit extracts action instructions in the target instruction text, wherein the action instructions comprise a moving action instruction and / or a local action instruction;
[0030] A process action sequence generation unit inputs the action instructions into an action generation model to obtain a process action sequence output by the action generation model, wherein the action generation model is constructed based on a diffusion model;
[0031] A sequence splicing unit splices the process action sequence according to an action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0032] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the action sequence generation method according to any one of the above or the action sequence generation method in a surgical room scenario when executing the program.
[0033] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the action sequence generation method according to any one of the above or the action sequence generation method in a surgical room scenario.
[0034] The application further provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the action sequence generation method according to any one of the above or the action sequence generation method in a surgical room scenario.
[0035] The action sequence generation method, device, electronic device, and storage medium provided by the application extract a moving action instruction and / or a local action instruction in a target instruction text, input the action instructions one by one into an action generation model to obtain a plurality of process action sequences output by the action generation model, splice the plurality of process action sequences according to an action execution order in the target instruction text to obtain a target action sequence, and thus generate an action sequence that conforms to a real trajectory and is accurate for a complex action instruction. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0037] Figure 1 is a flowchart of the action sequence generation method provided by the present application;
[0038] Figure 2 is one of the flowcharts of the action sequence generation method in the operating room scenario provided by the present application;
[0039] Figure 3 is another of the flowcharts of the action sequence generation method in the operating room scenario provided by the present application;
[0040] Figure 4 is a structural schematic diagram of the action sequence generation device provided by the present application;
[0041] Figure 5 is a structural schematic diagram of the action sequence generation device in the operating room scenario provided by the present application;
[0042] Figure 6 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0043] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application. In order to facilitate understanding, the embodiments of the present application select the practical application of the action sequence generation method in the operating room scenario for description.
[0044] In view of the above problems, the present application provides an action sequence generation method to realize accurate generation of complex action sequences. Figure 1 is a flowchart of the action sequence generation method provided by the present application, as shown in Figure 1 the method comprises:
[0045] Step 110, obtaining a target instruction text;
[0046] Here, the target instruction text refers to an action task described by natural language, which can include at least one of a moving task and a task of in-place operation.
[0047] Specifically, the target instruction text can be obtained by user input, API interface, database or file system. The target instruction text here can be "the doctor walks to the operating table and disinfects the patient". It should be noted that the action task described by the target instruction text here is usually a complex task containing multiple stages and continuous actions, rather than a single static action.
[0048] Step 120: Extract the action instructions from the target instruction text, the action instructions including movement action instructions and / or local action instructions;
[0049] Here, movement commands refer to commands that involve trajectory movement, affecting the global positional change of the target moving object, such as "forward" or "backward". Additionally, local movement commands refer to non-movement actions performed at a specific location, affecting the local actions of the target moving object, such as "disinfect" or "grab".
[0050] Specifically, the extracted prompt text can be obtained by adding the target instruction text into the constructed base prompt text. Then, the extracted prompt text can be input into a large language model, which outputs the action instructions from the target instruction text. These action instructions include movement action instructions and / or local action instructions.
[0051] It should be noted that the basic prompt text here can provide extraction examples and specify the format of the output action instructions, such as specifying JSON format. For example, the basic prompt text here could be: "You are a smart task parser responsible for converting [target instruction text] into structured JSON format. Your task is to analyze a series of instructions describing character movement and actions and output structured JSON data. Please follow these rules: categorize movement-related actions as movement; categorize stationary actions as stationary; maintain the original order of actions in the input target instruction text; if the target instruction text explicitly mentions the executor, include the "target movement object" field in the output JSON; do not add any additional details not mentioned in the input."
[0052] It should be noted that the constructed basic prompt text adheres to three key principles, enabling it to parse complex target instruction text. These key principles are: First, it can accurately identify meta-actions related to movement, i.e., accurately identify movement action instructions; second, it can accurately identify local actions at specific locations, i.e., accurately identify local action instructions; and third, the output action instructions must maintain the same order as the process actions in the input target instruction text, i.e., the execution order of movement actions and local actions must be consistent with the order in the target instruction text.
[0053] It should also be noted that by decomposing the target instruction text containing complex actions, movement action instructions and local action instructions are obtained, thereby dividing the movement trajectory and local actions, reducing the complexity of the actions, and making it easier to generate the process action sequences corresponding to the movement action instructions and local action instructions respectively, thus ensuring that an accurate target action sequence is obtained through the process action sequence.
[0054] Step 130: Input the action command into the action generation model to obtain the process action sequence output by the action generation model; the action generation model is constructed based on the diffusion model.
[0055] Here, the action generation model can be built based on the diffusion model to generate action sequences with a time dimension, i.e., process action sequences.
[0056] Specifically, action commands can be input into an action generation model, which then outputs a sequence of actions corresponding to each command. It's understood that action commands include multiple types, namely movement commands and local commands. For any type of action command, it can contain multiple commands of the same type; for example, it can contain two movement commands or three local commands.
[0057] Therefore, for any input action command, the action generation model can gradually reduce the noise of random Gaussian noise according to the instructions of the action command, such as reducing noise for t time steps, until a clean process action sequence is finally generated.
[0058] It should be noted that by constructing an action generation model using a diffusion model, during the training phase, noise is continuously added to the training sample data. Through this learning process, the model grasps the underlying patterns in the data, specifically the amount of noise added at each time step. In the application phase, the random noise is progressively reduced to generate a coherent sequence of actions that conforms to the distribution of real-world actions. Furthermore, by inputting action commands as input conditions, the generation target of the actions is controlled, making the generated action sequence even more closely resemble realistic action trajectories. Understandably, the high-quality trajectories generated based on complex target command text can be applied to tasks such as medical training, significantly improving the applicability of trajectory generation.
[0059] Step 140: Concatenate the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain the target action sequence.
[0060] Here, the action execution order refers to the execution order of all action instructions in the target instruction text. For example, if the target instruction text is "The doctor walks to the operating table and disinfects the patient", then the action execution order is to execute "The doctor walks to the operating table" first, and then execute "Disinfect the patient".
[0061] Specifically, the target instruction text can be input into a large language model, which then outputs the action execution sequence. The sequence of actions can then be concatenated according to this sequence to obtain the target action sequence. For example, this could be the complete action sequence of "a doctor walks to the operating table and disinfects the patient."
[0062] When the action command contains only a single movement action command or a single local action command, meaning the target command text contains only one action command, the process action sequence output by the action model can be directly used as the target action sequence. It can be understood that the target action sequence here is the action sequence containing the keyframes corresponding to the action command.
[0063] It should be noted that generating target action sequences based on complex target instruction text can be applied to action sequence generation in various scenarios, providing efficient and accurate operation process planning solutions for different industries and powerfully promoting the automation and intelligent development of various fields. For example, it can be applied to action sequence generation in operating rooms, manufacturing plants, and logistics warehouses. In medical scenarios, it can provide a new solution for medical training. That is, it can generate corresponding action trajectories based on text descriptions. Then, Unreal Engine can be used to create a realistic operating room environment with digital humans to achieve an immersive medical training experience, avoiding the high risks and high costs of traditional training processes, and adapting to complex task requirements.
[0064] For example, on an automated assembly line in an automobile manufacturing plant, various parts need to be picked up from a conveyor belt and placed in designated positions for subsequent assembly. This can be achieved by obtaining the user-inputted action command, "Pick up the red circular part from conveyor belt A and place it in the circular position on assembly table B," and applying the action sequence generation method provided in this embodiment of the invention to generate the corresponding action sequence. Then, by transmitting the generated action sequence to the robot control system, the robotic arm is controlled to perform the picking and placing actions, thus achieving automated assembly.
[0065] For example, in the warehousing area of a large manufacturing plant, raw materials, semi-finished products, and finished products need to be moved and stored between different warehouses and production lines. This can be achieved by obtaining the user-inputted action command "Move 10 boxes of raw materials from warehouse D to the designated location on production line E," and applying the action sequence generation method provided in this embodiment of the invention to generate the corresponding action sequence. Then, the handling equipment is controlled to execute the handling task according to the generated action sequence, ensuring the safe and efficient transportation of materials.
[0066] The method provided in this invention extracts movement commands and / or local commands from the target command text; inputs each command into the command generation model to obtain multiple process command sequences output by the command generation model; and concatenates the multiple process command sequences according to the execution order of the commands in the target command text to obtain the target command sequence. This method achieves the generation of accurate command sequences that conform to real trajectories for complex command commands.
[0067] To further enhance the realism of the action sequence in the process of generating mobile action instructions, based on any of the above embodiments, the action generation model includes a mobile action generation model; the action instruction includes the mobile action instruction.
[0068] Step 130 includes:
[0069] Obtain the noise sequence of the movement action;
[0070] Determine the target position corresponding to the movement action command;
[0071] The motion noise sequence, the target position, and the motion command are input into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model.
[0072] The target position is defined based on a three-dimensional coordinate system created from the pelvic nodes of the target moving object.
[0073] Here, the target position refers to the final location that the move action command wants the target object to move to.
[0074] Specifically, this can be achieved by acquiring a motion noise sequence, such as the noise sequence obtained by adding noise during the training phase of the motion generation model. Alternatively, a three-dimensional coordinate system can be created with the pelvic node of the moving object as the origin, providing the spatial location of the target within this system, which can then be used to define the target's position. For instance, the target's position in the pelvic node coordinate system can be calculated using the map coordinates of the target location and the positional transformation relationship between the map coordinate system and the coordinate system created with the pelvic node as the origin.
[0075] Then, by inputting the motion noise sequence, target position, and motion command into the motion generation model, and using the target position and motion command as generation constraints, the process motion sequence corresponding to the motion command output by the motion generation model can be obtained. The process motion sequence corresponding to the motion command can be calculated using the following formula, as shown below:
[0076]
[0077] In the formula, Represents the motion generation model The output movement command corresponds to the sequence of process actions, where m represents the movement action. The model parameters represent the motion generation model; This represents a sequence of moving motion noise. This represents the total time steps for denoising the noisy sequence of motion actions. , This indicates the generation of constraints. Indicates a movement action command. The target location can be defined based on the pelvic nodes in the SMPL-H model.
[0078] Understandably, the denoising process is repeated iteratively t times until a final clean sequence of movement actions is obtained, i.e., the sequence of process actions corresponding to the movement action command, denoted as . .
[0079] The method provided in this embodiment of the invention uses the target position as a conditional input to the motion generation model to constrain the process motion sequence generated by the motion generation model, so that the generated process motion sequence is more in line with the final position indicated by the motion action command, thereby improving the accuracy of the process motion sequence.
[0080] Based on any of the above embodiments, the motion generation model includes a motion encoder;
[0081] The step of inputting the motion noise sequence, the target position, and the motion command into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model includes:
[0082] Based on the motion encoder, the target position and the movement command are encoded to obtain motion constraint features;
[0083] Based on the mobile motion generation model, the process motion sequence corresponding to the mobile motion command is obtained by applying the mobile motion noise sequence and the motion constraint features.
[0084] It should be noted that, to enhance spatial control of the generated motion, a trainable motion encoder can be integrated into the zero-shot module of the motion generation model. This allows the motion encoder to accurately control the trajectory of the generated motion sequence, ensuring that the final position of the generated motion sequence aligns with the target position expected by the motion command.
[0085] Specifically, the target position and movement commands can be input into the motion encoder of the motion generation model. The motion encoder encodes the target position and movement commands to obtain motion constraint features. It should be noted that these motion constraint features can be considered as understandable latent variables, used to guide the diffusion model in generating a sequence of movements that conforms to the constraints. Then, the motion constraint features can guide the diffusion model in the motion generation model to denoise the noisy sequence of movements, generating a sequence of movements that conforms to the constraints, greatly improving the accuracy of the sequence of movements.
[0086] In addition, the training process of the motion encoder includes: first, obtaining the initial motion encoder, sample action description text, and the real trajectory labels corresponding to the sample action description text; inputting the sample action description text into the diffusion model containing the initial motion encoder to obtain the sample trajectory output by the initial motion encoder; applying the L2Loss loss function based on the sample trajectory and the real trajectory labels to calculate the movement trajectory loss; and iterating the initial motion encoder based on the movement trajectory loss to obtain the motion encoder.
[0087] The sample action description text includes the sample movement action command and the sample target location. In the sample data acquisition phase, firstly, an untrained motion encoder can be obtained as the initial motion encoder. Secondly, the sample action description text and the corresponding ground truth trajectory labels can be obtained as the labels for the training samples. The sample action description text and the corresponding ground truth trajectory labels can be obtained from the existing HumanML3D dataset.
[0088] Additionally, during the training phase, the model parameters of the diffusion model can be frozen, allowing adjustments to be made only to the motion encoder's model parameters. The trajectory loss here can be calculated using the following formula, as shown below:
[0089]
[0090] In the formula, Indicates the loss of movement trajectory; This represents the model parameters of the motion encoder in the motion generation model; Represents the sample trajectory; This indicates the actual trajectory label.
[0091] The method provided in this invention adds a motion encoder with a guidance mechanism to the motion generation model to encode the constraint text of the motion sequence generation, thereby obtaining understandable latent variables. This guides the diffusion model to generate motion sequences that meet the constraints, ensuring that the generated process motion sequence highly overlaps with the expected target position, thus improving the accuracy of the process motion sequence.
[0092] Based on any of the above embodiments, the action generation model further includes a local action generation model; the action instruction includes a local action instruction;
[0093] The action command is input into the action generation model to obtain the process action sequence output by the action generation model, including:
[0094] Determine the local motion noise sequence;
[0095] The local motion noise sequence and the local motion command are input into the local motion generation model to obtain the process motion sequence corresponding to the local motion command output by the local motion generation model.
[0096] Here, the local action generation model can be constructed using a UNet-based diffusion method to provide better performance in action generation and motion quality.
[0097] Specifically, firstly, a local action noise sequence can be obtained from the action sequence noise obtained during the training phase of the local action generation model. Then, the local action noise sequence and local action commands can be input into the local action generation model. The local action commands instruct the model to progressively denoise the local action noise sequence, resulting in the process action sequence corresponding to the local action commands, which can be denoised as the local action sequence. Here, the local action sequence can be calculated using the following formula, as shown below:
[0098]
[0099] In the formula, This represents the sequence of process actions corresponding to a local action instruction, i.e., the local action sequence. This represents a local action generation model, where, This represents the model parameters of the local action generation model. Indicates a partial action; Represents a localized action noise sequence; This represents the total time steps for denoising a local motion noise sequence. This indicates a local action command.
[0100] It should be noted that during the training phase of the local action generation model, a medical action dataset can be constructed for training the model. In the sample data construction phase, medical exam training videos can be crawled from video websites. The crawled videos are sliced using YOLO and RMTPose to obtain sample labels for the training samples. Then, the open-source method AiOS is used to estimate the SMPL-H parameters of the characters in the videos. Then, unreasonable action estimates can be manually adjusted, and textual descriptions of the actions can be provided as training samples. Finally, approximately 1000 pairs of sample data labels from medical scenarios can be obtained. Furthermore, the training process of the local action generation model can be supervised using the following objective function, as shown in the following equation:
[0101]
[0102] In the formula, This represents the objective function, or loss function, of the local action generation model; Represents the model parameters of the local action generation model; This represents the sample local action output by the local action generation model during training; This represents real data from a medical motion dataset.
[0103] It should be noted that the process action sequence is a segmented trajectory generated by the action generation model, such as the intermediate result of a movement action or a local action, which may contain discontinuities or abrupt changes in action. Therefore, in order to further improve the quality of the target action sequence, based on any of the above embodiments, step 140 includes:
[0104] Based on the interpolation function, the process action sequence is processed to obtain a smooth action sequence;
[0105] The smooth action sequence is spliced together according to the action execution order to obtain the target action sequence.
[0106] Here, the interpolation function can be a linear interpolation function, a Bézier curve, or other functions.
[0107] Specifically, a smooth motion sequence can be obtained by performing linear interpolation on the process motion sequences, such as the process motion sequences corresponding to movement motion commands and / or local motion commands, ensuring that the difference between adjacent keyframes is within an allowable range, such as the rate of change of velocity. Then, the smooth motion sequences can be spliced together according to the execution order of the actions to obtain a higher quality target motion sequence.
[0108] It should be noted that by interpolating the generated process action sequence, the transition frames reduce mechanical vibrations or animation stutters during action switching, so that the final target action sequence can adapt to the requirements of complex tasks.
[0109] Based on any of the above embodiments Figure 2 This is one of the flowcharts illustrating the action sequence generation method in an operating room setting provided by the present invention, such as... Figure 2 As shown, the method includes:
[0110] Step 210: Obtain the operating room simulation scene and the target motion sequence of the target moving object, wherein the target motion sequence is obtained based on the motion sequence generation method described in any of the above embodiments;
[0111] Step 220: In the operating room simulation scenario, drive the target moving object to move according to the target action sequence.
[0112] Specifically, firstly, a realistic operating room environment with a digital human can be created using Unreal Engine to achieve an immersive medical training experience. Here, the digital human refers to the target moving object. Secondly, the target moving object's target action sequence can be generated using the action sequence generation method in any of the above embodiments. It should be noted that in the operating room scenario, the target moving object's target action sequence involves interaction with the operating room environment and professional and complex medical actions. By decomposing the complex target instruction text using a large language model, executable movement action instructions and local action instructions are obtained, reducing the complexity of the actions and improving the accuracy of subsequent target action sequence generation based on text instructions. Furthermore, by using an action generation model based on a diffusion model to generate a process action sequence, the accuracy of the process action sequence is ensured.
[0113] Then, in the operating room simulation scenario, the target moving object is driven to move according to the target action sequence, realizing a virtual operating room based on static and dynamic digital assets. For example, using Unreal Engine plugins, motion parameter sequences can be used to drive virtual characters in the operating room simulation scenario.
[0114] In practical applications, due to the scarcity of motion data in medical surgery, the motion sequence generation method provided in this invention can construct a medical motion dataset containing 1,000 text-action pairs, covering 160,000 frames of medical motions, to achieve an immersive medical training experience.
[0115] The method provided in this invention extracts movement commands and / or local commands from the target command text; inputs each command into a command generation model to obtain multiple process command sequences output by the command generation model; and concatenates the multiple process command sequences according to the execution order of the commands in the target command text to obtain the target command sequence. This method enables the generation of accurate command sequences that conform to real trajectories for complex commands, thereby achieving the generation of trajectories in realistic and high-quality operating room scenarios.
[0116] Figure 3 This is the second flowchart illustrating the action sequence generation method for an operating room scenario provided by the present invention, as shown below. Figure 3 As shown, the method includes: First, in the task planning phase, receiving target instruction text input by the user. Then, inputting the target instruction text into a large language model, which outputs movement action instructions and local action instructions. Finally, a pelvic coordinate system is created with the pelvic node of the target moving object as the origin. The coordinate transformation relationship between the world coordinate system and the pelvic coordinate system is calculated and used to define the target position of the moving object in the pelvic coordinate system in the movement action command, that is, to define the spatial target.
[0117] Furthermore, this can be achieved by combining the movement command, target position, and movement noise sequence. The input is fed into the motion generation model to obtain the sequence of motion actions corresponding to the motion action commands output by the model. The zero-shot module in the motion generation model integrates a trainable motion encoder, i.e. Figure 3 The "zero conv" in the text. Additionally, this can be achieved by combining local motion commands and local motion noise sequences. The input is fed into a local action generation model to obtain the process action sequence corresponding to the local action instructions output by the model. Next, linear interpolation can be performed on the generated process action sequence to obtain a smooth action sequence. Then, the smooth action sequence can be concatenated according to the action execution order in the target instruction text to obtain the target action sequence. Finally, in an operating room simulation scenario, the target moving object can be driven to move according to the target action sequence, completing the movement trajectory and local actions.
[0118] Based on any of the above embodiments Figure 4 This is a schematic diagram of the action sequence generation device provided by the present invention, as shown below. Figure 4 As shown, the device includes:
[0119] Acquire unit 410 to acquire the target instruction text;
[0120] The instruction splitting unit 420 extracts action instructions from the target instruction text, the action instructions including movement action instructions and / or local action instructions;
[0121] The process action sequence generation unit 430 inputs the action instruction into the action generation model to obtain the process action sequence output by the action generation model; the action generation model is constructed based on the diffusion model.
[0122] The sequence splicing unit 440 splices the process action sequence according to the action execution order of the action instruction in the target instruction text to obtain the target action sequence.
[0123] The apparatus provided in this invention extracts movement commands and / or local commands from a target command text; inputs each command into a command generation model to obtain multiple process command sequences output by the command generation model; and concatenates these multiple process command sequences according to the execution order of the commands in the target command text to obtain a target command sequence. This achieves the generation of accurate command sequences that conform to real trajectories for complex command commands.
[0124] Based on any of the above embodiments, the action generation model includes a movement action generation model; the action instruction includes the movement action instruction;
[0125] The process action sequence generation unit is specifically used for:
[0126] Obtain the noise sequence of the movement action;
[0127] Determine the target position corresponding to the movement action command;
[0128] The motion noise sequence, the target position, and the motion command are input into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model.
[0129] The target position is defined based on a three-dimensional coordinate system created from the pelvic nodes of the target moving object.
[0130] Based on any of the above embodiments, the motion generation model includes a motion encoder;
[0131] The process action sequence generation unit is also specifically used for:
[0132] Based on the motion encoder, the target position and the movement command are encoded to obtain motion constraint features;
[0133] Based on the mobile motion generation model, the process motion sequence corresponding to the mobile motion command is obtained by applying the mobile motion noise sequence and the motion constraint features.
[0134] Based on any of the above embodiments, the action generation model further includes a local action generation model; the action instruction includes a local action instruction;
[0135] The process action sequence generation unit is also specifically used for:
[0136] Determine the local motion noise sequence;
[0137] The local motion noise sequence and the local motion command are input into the local motion generation model to obtain the process motion sequence corresponding to the local motion command output by the local motion generation model.
[0138] Based on any of the above embodiments, the sequence splicing unit is specifically used for:
[0139] Based on the interpolation function, the process action sequence is processed to obtain a smooth action sequence;
[0140] The smooth action sequence is spliced together according to the action execution order to obtain the target action sequence.
[0141] Based on any of the above embodiments Figure 5 This is a schematic diagram of the action sequence generation device for an operating room scenario provided by the present invention, as shown below. Figure 5 As shown, the device includes:
[0142] The operating room scene acquisition unit 510 acquires the operating room simulation scene and the target action sequence of the target moving object, wherein the target action sequence is obtained based on the action sequence generation method in any of the above embodiments;
[0143] The drive unit 520 drives the target moving object to move according to the target action sequence in the operating room simulation scenario.
[0144] The apparatus provided in this invention extracts movement commands and / or local commands from the target command text; inputs each command into the command generation model to obtain multiple process command sequences output by the command generation model; and concatenates the multiple process command sequences according to the execution order of the commands in the target command text to obtain the target command sequence. This achieves the generation of accurate command sequences that conform to real trajectories for complex command commands, thereby realizing the generation of trajectories in realistic and high-quality operating room scenarios.
[0145] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an action sequence generation method. This method includes: acquiring a target instruction text; extracting action instructions from the target instruction text, the action instructions including movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and concatenating the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0146] Alternatively, an action sequence generation method for an operating room scenario can be executed to obtain the operating room simulation scenario and the target action sequence of the target moving object. The target action sequence is obtained based on the action sequence generation method in any of the above embodiments. In the operating room simulation scenario, the target moving object is driven to move according to the target action sequence.
[0147] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0148] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the action sequence generation method provided by the above methods. The method includes: obtaining a target instruction text; extracting action instructions from the target instruction text, the action instructions including movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and concatenating the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0149] Alternatively, an action sequence generation method for an operating room scenario can be executed to obtain the operating room simulation scenario and the target action sequence of the target moving object. The target action sequence is obtained based on the action sequence generation method in any of the above embodiments. In the operating room simulation scenario, the target moving object is driven to move according to the target action sequence.
[0150] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the action sequence generation method provided by the above methods. The method includes: acquiring a target instruction text; extracting action instructions from the target instruction text, the action instructions including movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and concatenating the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0151] Alternatively, an action sequence generation method for an operating room scenario can be executed to obtain the operating room simulation scenario and the target action sequence of the target moving object. The target action sequence is obtained based on the action sequence generation method in any of the above embodiments. In the operating room simulation scenario, the target moving object is driven to move according to the target action sequence.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0153] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating action sequences, characterized in that, include: Obtain the target instruction text; Extract the action instructions from the target instruction text. The action instructions include movement action instructions and / or local action instructions. The movement action instructions include global position changes of the target moving object. The local action instructions include non-movement actions performed at specific locations. The action command is input into the action generation model to obtain the process action sequence output by the action generation model; the action generation model is constructed based on the diffusion model. The process action sequence is concatenated according to the action execution order of the action instructions in the target instruction text to obtain the target action sequence; The motion generation model includes a movement motion generation model; the motion command includes the movement motion command. The process of inputting the action command into the action generation model to obtain the process action sequence output by the action generation model includes: Obtain the noise sequence of the movement action; Determine the target position corresponding to the movement action command; The motion noise sequence, the target position, and the motion command are input into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model. The target position is defined based on a three-dimensional coordinate system created from the pelvic nodes of the target moving object; The motion generation model includes a motion encoder; The step of inputting the motion noise sequence, the target position, and the motion command into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model includes: Based on the motion encoder, the target position and the movement command are encoded to obtain motion constraint features; Based on the mobile motion generation model, the mobile motion noise sequence and the motion constraint features are applied to obtain the process motion sequence corresponding to the mobile motion command; The action generation model further includes a local action generation model; the action instructions include local action instructions. The process of inputting the action command into the action generation model to obtain the process action sequence output by the action generation model includes: Determine the local motion noise sequence; The local motion noise sequence and the local motion command are input into the local motion generation model to obtain the process motion sequence corresponding to the local motion command output by the local motion generation model.
2. The action sequence generation method according to claim 1, characterized in that, The step of concatenating the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain the target action sequence includes: Based on the interpolation function, the process action sequence is processed to obtain a smooth action sequence; The smooth action sequence is spliced together according to the action execution order to obtain the target action sequence.
3. An action sequence generation device, characterized in that, include: The acquisition unit acquires the target instruction text. The instruction splitting unit extracts action instructions from the target instruction text. The action instructions include movement action instructions and / or local action instructions. The movement action instructions include global position changes of the target moving object. The local action instructions include non-movement actions performed at specific locations. A process action sequence generation unit inputs the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model. The sequence splicing unit splices the process action sequence according to the action execution order of the action instruction in the target instruction text to obtain the target action sequence; The motion generation model includes a movement motion generation model; the motion command includes the movement motion command. The process action sequence generation unit is specifically used for: Obtain the noise sequence of the movement action; Determine the target position corresponding to the movement action command; The motion noise sequence, the target position, and the motion command are input into the motion generation model to obtain the process motion sequence corresponding to the motion command output by the motion generation model. The target position is defined based on a three-dimensional coordinate system created from the pelvic nodes of the target moving object; The motion generation model includes a motion encoder; The process action sequence generation unit is also specifically used for: Based on the motion encoder, the target position and the movement command are encoded to obtain motion constraint features; Based on the mobile motion generation model, the mobile motion noise sequence and the motion constraint features are applied to obtain the process motion sequence corresponding to the mobile motion command; The action generation model further includes a local action generation model; the action instructions include local action instructions. The process action sequence generation unit is also specifically used for: Determine the local motion noise sequence; The local motion noise sequence and the local motion command are input into the local motion generation model to obtain the process motion sequence corresponding to the local motion command output by the local motion generation model.
4. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the action sequence generation method as described in any one of claims 1 to 2.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the action sequence generation method as described in any one of claims 1 to 2.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the action sequence generation method as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Multi-person interaction action generation method and device and electronic equipment
CN119024971A
Human body action sequence generation method and device, equipment and storage medium
CN120088867A