Action sequence generation method and device, electronic equipment and storage medium
By extracting and splicing action instructions in the action generation model, combining the diffusion model and the motion encoder, the problem that the action sequence in the prior art does not meet expectations is solved, and accurate action sequence generation is achieved, which is suitable for virtual reality and robot control.
Patent Information
- Application Number
- CN202510942578.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-09
AI Technical Summary
When existing machine learning algorithms process complex action instruction text, the generated action sequences do not meet expectations, resulting in unsatisfactory reconstruction results.
By extracting the moving action instructions and local action instructions in the target instruction text, input them into the action generation model based on the diffusion model, generating a process action sequence, and splicing them in the order of action execution, using the motion encoder and the local action generation model for precise control, and processing the action sequence with the interpolation function.
It realizes the accurate generation of complex action instructions, the generated action sequence conforms to the real trajectory, is suitable for fields such as virtual reality and robot control, improving user experience and operation accuracy.
Smart Images

Figure CN120429017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an action sequence generation method, device, electronic device and storage medium. Background Art
[0002] Motion generation technology plays a crucial role in fields such as virtual reality and robotic control. Applications in these fields often demand highly realistic and smooth motion performance to enhance user experience or enable precise robotic operation. Traditional motion generation methods often rely on manual design and adjustment, resulting in low efficiency and high costs. With the development of artificial intelligence (AI), the use of machine learning algorithms to automatically generate motion is increasingly being explored to improve efficiency and reduce costs.
[0003] However, existing machine learning algorithms often seem unable to cope with complex instruction texts, resulting in the generated action sequences not meeting expectations, and the reconstruction effect of the action sequences is not ideal. Summary of the Invention
[0004] The present invention provides an action sequence generation method, device, electronic device and storage medium, which are used to solve the defect in the prior art that the actions do not meet expectations when the action sequence is reconstructed based on text.
[0005] The present invention provides an action sequence generation method, comprising: Get the target instruction text; Extracting action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; Inputting the action instruction into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; According to the action execution order of the action instructions in the target instruction text, the process action sequence is spliced to obtain the target action sequence.
[0006] According to an action sequence generation method provided by the present invention, the action generation model includes a movement action generation model; the action instruction includes the movement action instruction; The step of inputting the action instruction into the action generation model to obtain the process action sequence output by the action generation model includes: Obtain a moving action noise sequence; Determining a target position corresponding to the movement action instruction; Inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model; The target position is defined based on a three-dimensional coordinate system created by the pelvic node of the target moving object.
[0007] According to a method for generating an action sequence provided by the present invention, the mobile action generation model includes a motion encoder; The step of inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model includes: Encoding the target position and the movement action instruction based on the motion encoder to obtain a motion constraint feature; Based on the movement action generation model, the movement action noise sequence and the motion constraint feature are applied to obtain a process action sequence corresponding to the movement action instruction.
[0008] According to an action sequence generation method provided by the present invention, the action generation model further includes a local action generation model; the action instructions include local action instructions; The step of inputting the action instruction into the action generation model to obtain the process action sequence output by the action generation model includes: Determine the local motion noise sequence; The local action noise sequence and the local action instruction are input into the local action generation model to obtain a process action sequence corresponding to the local action instruction output by the local action generation model.
[0009] According to an action sequence generation method provided by the present invention, the process action sequence is spliced according to the action execution order of the action instructions in the target instruction text to obtain the target action sequence, including: Based on the interpolation function, the process action sequence is processed respectively to obtain a smooth action sequence; The smooth action sequences are spliced according to the action execution order to obtain a target action sequence.
[0010] The present invention also provides an action sequence generating device, comprising: Acquisition unit, obtains the target instruction text; An instruction splitting unit extracts action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; a process action sequence generating unit, which inputs the action instruction into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; The sequence splicing unit splices the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements any of the above-mentioned action sequence generation methods, or the action sequence generation method in an operating room scenario.
[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating an action sequence as described above or the method for generating an action sequence in an operating room scenario is implemented.
[0013] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned action sequence generation methods, or the action sequence generation method in an operating room scenario.
[0014] The action sequence generation method, device, electronic device and storage medium provided by the present invention extract movement action instructions and / or local action instructions from the target instruction text; input the action instructions one by one into the action generation model to obtain multiple process action sequences output by the action generation model; splice the multiple process action sequences according to the action execution order in the target instruction text to obtain the target action sequence, thereby achieving the generation of action sequences that conform to real trajectories and are accurate for complex action instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 It is a flowchart of the action sequence generation method provided by the present invention; Figure 2 This is one of the flow charts of the method for generating an action sequence in an operating room scenario provided by the present invention; Figure 3 This is the second flow chart of the method for generating an action sequence in an operating room scenario provided by the present invention; Figure 4 It is a structural diagram of the action sequence generating device provided by the present invention; Figure 5It is a structural diagram of the action sequence generation device in the operating room scenario provided by the present invention; Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In order to make it easier to understand, the embodiment of the present invention selects the actual application of the action sequence generation method in the operating room scenario for description.
[0018] In view of the above problems, the present invention provides an action sequence generation method to achieve accurate generation of complex action sequences. Figure 1 It is a flow chart of the action sequence generation method provided by the present invention, such as Figure 1 As shown, the method includes: Step 110, obtaining the target instruction text; Here, the target instruction text refers to an action task described in natural language, which may include at least one of a moving task and an in-place operation task.
[0019] Specifically, the target instruction text can be obtained through user input, an API interface, a database, or a file system. Here, the target instruction text might be "The doctor walks to the operating table and disinfects the patient." It should be noted that the action task described by the target instruction text is typically a complex task involving multiple stages and continuous actions, rather than consisting of a single static action.
[0020] Step 120: extracting action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; Here, movement instructions refer to instructions involving trajectory movement, involving global position changes of the target moving object, such as "forward" and "backward." In addition, local action instructions refer to non-movement actions performed at specific locations, involving local movements of the target moving object, such as "disinfect" and "grab."
[0021] Specifically, the extracted prompt text can be obtained by inserting the target instruction text into the constructed basic prompt text. The extracted prompt text can then be input into a large language model, which then outputs the action instructions in the target instruction text. The action instructions include movement action instructions and / or local action instructions.
[0022] It should be noted that the basic prompt text here can provide extraction examples and limit the format of the output action instructions, such as limiting the output to JSON format. For example, the basic prompt text here can be "You are an intelligent task parser, responsible for converting [target instruction text] into a structured JSON format. Your task is to analyze a series of instructions describing the character's movement and actions, and output structured JSON data. Please follow the following rules: classify actions related to movement as movement; classify actions in a stationary state as stationary; and maintain the original order of actions in the input target instruction text; if the executor is explicitly mentioned in the target instruction text, please include the "target movement object" field in the output JSON; do not add additional details not mentioned in the input."
[0023] It should be noted that the constructed basic prompt text adheres to three key principles, enabling it to parse complex target instruction text. These individual key principles include: first, accurate identification of meta-actions related to movement, i.e., accurate identification of movement action instructions; second, accurate identification of local actions at specific locations, i.e., accurate identification of local action instructions; and third, the output action instructions must maintain the same sequence of process actions in the input target instruction text, i.e., the execution order of movement actions and local actions must be consistent with the order in the target instruction text.
[0024] It should also be noted that by decomposing the target instruction text containing complex actions, moving action instructions and local action instructions are obtained, which realizes the division of moving trajectories and local actions, reduces the complexity of the actions, and facilitates the subsequent generation of process action sequences corresponding to moving action instructions and local action instructions, thereby ensuring that an accurate target action sequence is obtained through the process action sequence.
[0025] Step 130: input the action instruction into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; Here, the action generation model can be constructed based on the diffusion model and used to generate an action sequence with a time dimension, that is, a process action sequence.
[0026] Specifically, the action instructions can be input into the action generation model, and the action generation model can output the process action sequence corresponding to each action instruction. It is understood that the action instructions include multiple types of action instructions, that is, movement action instructions and local action instructions. For any type of action instruction, it can include multiple action instructions of the same type, such as two movement action instructions or three local action instructions.
[0027] Therefore, for any input action instruction, the action generation model can gradually reduce the random Gaussian noise according to the instructions of the action instruction, such as performing noise reduction for t time steps until a clean process action sequence is finally generated.
[0028] It should be noted that the action generation model is constructed using a diffusion model. During the training phase, the model continuously adds noise to the training sample data. Through this learning process, the model grasps the underlying patterns of the data, namely the amount of noise added by the action diffusion model at each time step. During the application phase, the random noise is gradually reduced to generate a coherent action sequence that conforms to the distribution of real actions. Furthermore, by inputting action instructions and using these instructions as input conditions, the generated action targets are controlled, making the generated action sequences more consistent with real action trajectories. It can be understood that the high-quality trajectories generated based on complex target instruction text can be applied to tasks such as medical training, greatly improving the applicability of trajectory generation.
[0029] Step 140 : splicing the process action sequence according to the action execution order of the action instruction in the target instruction text to obtain a target action sequence.
[0030] Here, the action execution order refers to the execution order of all action instructions in the target instruction text. For example, if the target instruction text is "The doctor walks to the operating table and disinfects the patient", the action execution order is to execute "The doctor walks to the operating table" first, and then execute "Disinfect the patient".
[0031] Specifically, the target instruction text can be input into a large language model, which then outputs the action execution sequence. Then, the process action sequence can be spliced together according to the action execution order to obtain the target action sequence. For example, the action sequence for the entire process of "a doctor walks to the operating table and disinfects the patient" can be obtained.
[0032] When the action instruction only contains a single movement action instruction or a single local action instruction, it means that the target instruction text only contains one action instruction. Then, the process action sequence output by the action model can be directly used as the target action sequence. It is understood that the target action sequence here is the action sequence containing the key frames corresponding to the action instruction.
[0033] It's important to note that generating target action sequences based on complex target instruction text can be applied to action sequence generation in various scenarios, providing efficient and accurate operational process planning solutions for different industries and significantly promoting the development of automation and intelligentization in various fields. For example, it can be applied to action sequence generation in application scenarios such as operating rooms, manufacturing plants, and logistics warehouses. For example, in medical scenarios, it can provide a new solution for medical training. Specifically, it can generate corresponding action trajectories based on text descriptions. Then, using Unreal Engine, a realistic operating room environment with digital humans can be created, enabling an immersive medical training experience, avoiding the high risks and costs of traditional training processes and adapting to complex task demands.
[0034] For example, on an automated assembly line in an automobile manufacturing plant, various parts need to be grabbed from a conveyor belt and placed in designated locations for subsequent assembly. By obtaining the user-input action instruction "grab the red circular part from conveyor belt A and place it in the circular position on assembly table B," the action sequence generation method provided by an embodiment of the present invention can be applied to generate an action sequence corresponding to the action instruction. The generated action sequence is then transmitted to a robot control system, controlling the robotic arm to execute the grabbing and placing actions, thereby achieving automated assembly.
[0035] For example, in the storage area of a large manufacturing plant, raw materials, semi-finished products, and finished products need to be moved and stored between different warehouses and production lines. By obtaining the user-input action instruction "Move 10 boxes of raw materials from warehouse D to a specified location on production line E," the action sequence generation method provided by an embodiment of the present invention can be applied to generate an action sequence corresponding to the action instruction. The handling equipment is then controlled to execute the handling task according to the generated action sequence, ensuring the safe and efficient transportation of materials.
[0036] The method provided by the embodiment of the present invention extracts movement action instructions and / or local action instructions from the target instruction text; inputs the action instructions one by one into the action generation model to obtain multiple process action sequences output by the action generation model; splices the multiple process action sequences according to the action execution order in the target instruction text to obtain the target action sequence, thereby realizing the generation of accurate action sequences that conform to real trajectories for complex action instructions.
[0037] In order to further improve the authenticity of the action sequence in the process of generating the mobile action instruction, based on any of the above embodiments, the action generation model includes a mobile action generation model; the action instruction includes the mobile action instruction; Step 130 includes: Obtain a moving action noise sequence; Determining a target position corresponding to the movement action instruction; Inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model; The target position is defined based on a three-dimensional coordinate system created by the pelvic node of the target moving object.
[0038] Here, the target position refers to the position where the target moving object is expected to move to in the final step of the movement action instruction.
[0039] Specifically, a mobile action noise sequence can be obtained, for example, a noise sequence obtained by adding noise to a mobile action generation model during the training phase, as a mobile action noise sequence. In addition, a three-dimensional coordinate system can be created with the pelvic node of the target moving object as the origin of the coordinate system to obtain the spatial position of the target position in the three-dimensional coordinate system for defining the target position. For example, the target position in the pelvic node coordinate system can be calculated using the map coordinates of the target position in the map coordinate system and the position conversion relationship between the map coordinate system and the coordinate system created with the pelvic node as the origin.
[0040] Then, by inputting the movement noise sequence, target position, and movement instruction into the movement generation model, and using the target position and movement instruction as generation constraints, the process action sequence corresponding to the movement instruction output by the movement generation model can be obtained. The process action sequence corresponding to the movement instruction can be calculated using the following formula, as shown below: Where, Represents a mobile action generation model The process action sequence corresponding to the output movement action instruction, m represents the movement action, represents the model parameters of the mobile action generation model; represents a sequence of moving action noises; represents the total time steps for denoising the mobile action noise sequence; , represents the generation constraints, Indicates a move action instruction. Indicates the target position, which can be defined based on the pelvic node in the SMPL-H model.
[0041] It can be understood that the denoising process is repeated t times until the final clean movement action sequence is obtained, that is, the process action sequence corresponding to the movement action instruction, which is recorded as .
[0042] The method provided by an embodiment of the present invention uses the target position as the conditional input of the mobile action generation model to constrain the process action sequence generated by the mobile action generation model, so that the generated process action sequence is more consistent with the final position indicated by the mobile action instruction, thereby improving the accuracy of the process action sequence.
[0043] Based on any of the above embodiments, the movement action generation model includes a motion encoder; The step of inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model includes: Encoding the target position and the movement action instruction based on the motion encoder to obtain a motion constraint feature; Based on the movement action generation model, the movement action noise sequence and the motion constraint feature are applied to obtain a process action sequence corresponding to the movement action instruction.
[0044] It should be noted that to enhance spatial control of the generated motion, a trainable motion encoder can be integrated into the zero-shot module of the mobile motion generation model. This allows the mobile motion generation model to accurately control the trajectory of the generated process motion sequence, ensuring that the final position of the generated process motion sequence is aligned with the target position expected by the mobile motion instruction.
[0045] Specifically, the target position and movement instructions can be input into the motion encoder in the movement generation model, which then encodes the target position and movement instructions to obtain motion constraint features. It should be noted that the motion constraint features here can be considered as understandable latent variables, which are used to guide the diffusion model to generate a motion sequence that meets the constraints. The motion constraint features can then be used to guide the diffusion model in the movement generation model to denoise the noise sequence of the movement and generate a process motion sequence that meets the constraints, greatly improving the accuracy of the process motion sequence.
[0046] In addition, the training process of the motion encoder includes: first, obtaining the initial motion encoder, sample action description text and the real trajectory label corresponding to the sample action description text; inputting the sample action description text into the diffusion model containing the initial motion encoder to obtain the sample trajectory output by the initial motion encoder; based on the sample trajectory and the real trajectory label, applying the L2Loss loss function to calculate the movement trajectory loss; based on the movement trajectory loss, iterating the initial motion encoder to obtain the motion encoder.
[0047] The sample action description text includes sample movement instructions and sample target positions. During the sample data acquisition phase, an untrained motion encoder can be first obtained as the initial motion encoder. Furthermore, the sample action description text is obtained as a training sample, and the ground truth trajectory labels corresponding to the sample action description text are obtained as labels for the training samples. The sample action description text and the ground truth trajectory labels corresponding to the sample action description text can be obtained from the existing HumanML3D dataset.
[0048] In addition, during the training phase, the model parameters of the diffusion model can be frozen and only the model parameters of the motion encoder can be adjusted. The motion trajectory loss here can be calculated using the following formula, as shown below: Where, represents the loss of moving trajectory; Represents the model parameters of the motion encoder in the mobile action generation model; represents the sample trajectory; represents the true trajectory label.
[0049] The method provided by an embodiment of the present invention adds a motion encoder with a guidance mechanism to the mobile action generation model to encode the constraint text generated by the action sequence, obtains understandable latent variables, and guides the diffusion model to generate an action sequence that meets the constraints, ensuring that the generated process action sequence is highly consistent with the expected target position, thereby improving the accuracy of the process action sequence.
[0050] Based on any of the above embodiments, the action generation model further includes a local action generation model; the action instruction includes a local action instruction; Inputting the action instruction into the action generation model to obtain the process action sequence output by the action generation model includes: Determine the local motion noise sequence; The local action noise sequence and the local action instruction are input into the local action generation model to obtain a process action sequence corresponding to the local action instruction output by the local action generation model.
[0051] Here, the local motion generation model can be constructed using a UNet-based diffusion method to provide better performance in motion generation and motion quality.
[0052] Specifically, first, a local action noise sequence can be obtained from the action sequence noise obtained in the training phase of the local action generation model. Then, the local action noise sequence and the local action instruction can be input into the local action generation model. The local action instruction can be used to instruct the local action generation model to gradually reduce the noise of the local action noise sequence, thereby obtaining the process action sequence corresponding to the local action instruction, which can be recorded as the local action sequence. Here, the local action sequence can be calculated using the following formula, as shown below: Where, It represents the process action sequence corresponding to the local action instruction, that is, the local action sequence; represents the local action generation model, where represents the model parameters of the local action generation model, Indicates local action; represents a local action noise sequence; represents the total time steps for denoising the local action noise sequence; Indicates a local action instruction.
[0053] It should be noted that in the training stage of the local action generation model, a medical action dataset can be constructed to train the local action generation model. In the sample data construction stage, physician examination training videos can be crawled from video websites. The crawled videos are sliced by Yolo and RMTPose to obtain sample labels for training samples. Then, the open source method AiOS is used to estimate the SMPL-H parameters of the characters in the video. Then, unreasonable action estimates can be manually adjusted and given text descriptions of the actions as training samples. Finally, about 1,000 pairs of sample data label pairs from medical scenes can be obtained. In addition, the training process of the local action generation model can be supervised by the following objective function, as shown in the following formula: Where, represents the objective function of the local action generation model, i.e., the loss function; represents the model parameters of the local action generation model; represents the sample local actions output by the local action generation model during training; Represents real data from the medical action dataset.
[0054] It should be noted that the process action sequence is a segmented trajectory generated by the action generation model, such as the intermediate results of the moving action or local action, and may contain action discontinuities or sudden changes. Therefore, in order to further improve the quality of the target action sequence, based on any of the above embodiments, step 140 includes: Based on the interpolation function, the process action sequence is processed respectively to obtain a smooth action sequence; The smooth action sequences are spliced according to the action execution order to obtain a target action sequence.
[0055] Here, the interpolation function may be a linear interpolation function, a Bezier curve, or the like.
[0056] Specifically, a smoothed action sequence can be obtained by linearly interpolating the process action sequences, such as those corresponding to movement action instructions and / or local action instructions, to ensure that the difference between adjacent keyframes is within an acceptable range, such as the velocity change rate. The smoothed action sequences can then be spliced together according to the order in which the actions are executed to obtain a higher-quality target action sequence.
[0057] It should be noted that by interpolating the generated process action sequence, the transition frames reduce the mechanical vibration or animation freeze of the action switching, so that the final target action sequence can adapt to the requirements of complex tasks.
[0058] Based on any of the above embodiments, Figure 2 This is one of the flow charts of the method for generating an action sequence in an operating room scenario provided by the present invention. Figure 2 As shown, the method includes: Step 210: Acquire the target action sequence of the operating room simulation scene and the target moving object, wherein the target action sequence is obtained based on the action sequence generation method described in any of the above embodiments; Step 220: In the operating room simulation scene, drive the target moving object to move according to the target action sequence.
[0059] Specifically, first, a realistic operating room environment with a digital human can be created through Unreal Engine to achieve an immersive medical training experience. The digital human here refers to the target moving object. In addition, the target action sequence of the target moving object can be generated through the action sequence generation method in any of the above embodiments. It should be noted that in the operating room scenario, the target action sequence of the target moving object will involve interaction with the operating room scene, as well as professional and complex medical actions. The complex target instruction text is split by a large language model to obtain executable movement action instructions and local action instructions, which reduces the complexity of the action and improves the accuracy of the subsequent generation of target action sequences based on text instructions. In addition, the process action sequence is generated by the action generation model constructed based on the diffusion model, which ensures the accuracy of the process action sequence.
[0060] Then, in an operating room simulation, the target moving object is driven to move according to the target action sequence, realizing a virtual operating room generated based on static and dynamic digital assets. For example, the action parameter sequence can be used to drive the virtual character in the operating room simulation scene in conjunction with the Unreal Engine plug-in.
[0061] In practical applications, due to the scarcity of motion data in medical surgery, the motion sequence generation method provided by the embodiment of the present invention can construct a medical motion dataset, which contains 1,000 text-action pairs and covers 160,000 frames of medical actions to achieve an immersive medical training experience.
[0062] The method provided by the embodiment of the present invention extracts movement action instructions and / or local action instructions from the target instruction text; inputs the action instructions one by one into the action generation model to obtain multiple process action sequences output by the action generation model; splices the multiple process action sequences according to the action execution order in the target instruction text to obtain the target action sequence, thereby realizing the generation of accurate action sequences that conform to real trajectories for complex action instructions, and thus realizing the generation of real, high-quality trajectories in operating room scenarios.
[0063] Figure 3 This is the second flow chart of the method for generating an action sequence in an operating room scenario provided by the present invention. Figure 3 As shown, the method includes: first, in the task planning stage, the target instruction text input by the user can be received. Then, the target instruction text can be input into a large language model, and the large language model can output movement action instructions and local action instructions. Then, the pelvic coordinate system can be created by taking the pelvic node of the target moving object as the coordinate origin. , the coordinate transformation relationship between the world coordinates in the world coordinate system and the pelvic coordinate system is calculated to define the target position of the target moving object in the pelvic coordinate system in the movement action instruction, that is, to define the spatial target.
[0064] Furthermore, the movement action instruction, target position and movement action noise sequence can be , is input into the mobile action generation model to obtain the process action sequence corresponding to the mobile action instruction output by the mobile action generation model. Among them, the zero-shot module in the mobile action generation model integrates a trainable motion encoder, namely Figure 3 In addition, the local action command and the local action noise sequence can be , which is input into the local action generation model to obtain the process action sequence corresponding to the local action instructions output by the local action generation model. Next, the generated process action sequence can be linearly interpolated to obtain a smooth action sequence. The smooth action sequence can then be spliced according to the action execution order in the target instruction text to obtain the target action sequence. Furthermore, in an operating room simulation scenario, the target moving object can be driven to move according to the target action sequence, completing the movement trajectory and local actions.
[0065] Based on any of the above embodiments, Figure 4 Schematic diagram of the structure of the action sequence generating device provided by the present invention. Figure 4 As shown, the device includes: An acquisition unit 410 acquires a target instruction text; An instruction splitting unit 420 extracts action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; The process action sequence generating unit 430 inputs the action instruction into the action generation model to obtain the process action sequence output by the action generation model; the action generation model is constructed based on the diffusion model; The sequence splicing unit 440 splices the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0066] The device provided by the embodiment of the present invention extracts movement action instructions and / or local action instructions from the target instruction text; inputs the action instructions one by one into the action generation model to obtain multiple process action sequences output by the action generation model; splices the multiple process action sequences according to the action execution order in the target instruction text to obtain the target action sequence, thereby realizing the generation of accurate action sequences that conform to real trajectories for complex action instructions.
[0067] Based on any of the above embodiments, the action generation model includes a movement action generation model; the action instruction includes the movement action instruction; The process action sequence generation unit is specifically used to: Obtain a moving action noise sequence; Determining a target position corresponding to the movement action instruction; Inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model; The target position is defined based on a three-dimensional coordinate system created by the pelvic node of the target moving object.
[0068] Based on any of the above embodiments, the movement action generation model includes a motion encoder; The process action sequence generation unit is further specifically used for: Encoding the target position and the movement action instruction based on the motion encoder to obtain a motion constraint feature; Based on the movement action generation model, the movement action noise sequence and the motion constraint feature are applied to obtain a process action sequence corresponding to the movement action instruction.
[0069] Based on any of the above embodiments, the action generation model further includes a local action generation model; the action instruction includes a local action instruction; The process action sequence generation unit is further specifically used for: Determine the local motion noise sequence; The local action noise sequence and the local action instruction are input into the local action generation model to obtain a process action sequence corresponding to the local action instruction output by the local action generation model.
[0070] Based on any of the above embodiments, the sequence splicing unit is specifically used for: Based on the interpolation function, the process action sequence is processed respectively to obtain a smooth action sequence; The smooth action sequences are spliced according to the action execution order to obtain a target action sequence.
[0071] Based on any of the above embodiments, Figure 5 This is a schematic diagram of the structure of the action sequence generation device in the operating room scenario provided by the present invention. Figure 5 As shown, the device includes: An operating room scene acquisition unit 510 acquires an operating room simulation scene and a target motion sequence of a target moving object, wherein the target motion sequence is obtained based on the motion sequence generation method in any of the above embodiments; The driving unit 520 drives the target moving object to move according to the target action sequence in the operating room simulation scene.
[0072] The device provided by the embodiment of the present invention extracts movement action instructions and / or local action instructions from the target instruction text; inputs the action instructions one by one into the action generation model to obtain multiple process action sequences output by the action generation model; splices the multiple process action sequences according to the action execution order in the target instruction text to obtain the target action sequence, thereby realizing the generation of accurate action sequences that conform to real trajectories for complex action instructions, and thus realizing the generation of real, high-quality trajectories in operating room scenarios.
[0073] Figure 6An example of a physical structure diagram of an electronic device is shown below. Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communications bus 640. The processor 610 may call logic instructions in the memory 630 to execute an action sequence generation method, which includes: obtaining a target instruction text; extracting action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and splicing the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0074] The action sequence generation method in the operating room scene can also be executed to obtain the target action sequence of the operating room simulation scene and the target moving object, and the target action sequence is obtained based on the action sequence generation method in any of the above embodiments; in the operating room simulation scene, the target moving object is driven to move according to the target action sequence.
[0075] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0076] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the action sequence generation method provided by the above methods, which includes: obtaining a target instruction text; extracting action instructions in the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and splicing the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0077] The action sequence generation method in the operating room scene can also be executed to obtain the target action sequence of the operating room simulation scene and the target moving object, and the target action sequence is obtained based on the action sequence generation method in any of the above embodiments; in the operating room simulation scene, the target moving object is driven to move according to the target action sequence.
[0078] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the action sequence generation method provided by the above-mentioned methods, the method comprising: obtaining a target instruction text; extracting action instructions in the target instruction text, the action instructions comprising movement action instructions and / or local action instructions; inputting the action instructions into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; and splicing the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
[0079] The action sequence generation method in the operating room scene can also be executed to obtain the target action sequence of the operating room simulation scene and the target moving object, and the target action sequence is obtained based on the action sequence generation method in any of the above embodiments; in the operating room simulation scene, the target moving object is driven to move according to the target action sequence.
[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0081] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating an action sequence, characterized in that: include: Get the target instruction text; Extracting action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; Inputting the action instruction into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; According to the action execution order of the action instructions in the target instruction text, the process action sequence is spliced to obtain the target action sequence.
2. The method for generating an action sequence according to claim 1, wherein: The action generation model includes a movement action generation model; the action instruction includes the movement action instruction; The step of inputting the action instruction into the action generation model to obtain the process action sequence output by the action generation model includes: Obtain a moving action noise sequence; Determining a target position corresponding to the movement action instruction; Inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model; The target position is defined based on a three-dimensional coordinate system created by the pelvic node of the target moving object.
3. The action sequence generation method according to claim 2, characterized in that: The movement action generation model includes a motion encoder; The step of inputting the movement action noise sequence, the target position, and the movement action instruction into the movement action generation model to obtain a process action sequence corresponding to the movement action instruction output by the movement action generation model includes: Encoding the target position and the movement action instruction based on the motion encoder to obtain a motion constraint feature; Based on the movement action generation model, the movement action noise sequence and the motion constraint feature are applied to obtain a process action sequence corresponding to the movement action instruction.
4. The method for generating an action sequence according to any one of claims 1 to 3, characterized in that: The action generation model further includes a local action generation model; the action instructions include local action instructions; The step of inputting the action instruction into the action generation model to obtain the process action sequence output by the action generation model includes: Determine the local motion noise sequence; The local action noise sequence and the local action instruction are input into the local action generation model to obtain a process action sequence corresponding to the local action instruction output by the local action generation model.
5. The method for generating an action sequence according to any one of claims 1 to 3, characterized in that: The process action sequence is spliced according to the action execution order of the action instructions in the target instruction text to obtain the target action sequence, including: Based on the interpolation function, the process action sequence is processed respectively to obtain a smooth action sequence; The smooth action sequences are spliced according to the action execution order to obtain a target action sequence.
6. An action sequence generating device, characterized in that: include: Acquisition unit, obtains the target instruction text; An instruction splitting unit extracts action instructions from the target instruction text, wherein the action instructions include movement action instructions and / or local action instructions; a process action sequence generating unit, which inputs the action instruction into an action generation model to obtain a process action sequence output by the action generation model; the action generation model is constructed based on a diffusion model; The sequence splicing unit splices the process action sequence according to the action execution order of the action instructions in the target instruction text to obtain a target action sequence.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the action sequence generation method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the action sequence generation method according to any one of claims 1 to 5 is implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the action sequence generation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Motion generation method and virtual character animation generation method
CN116392812A
Human body action generation method, device and product based on natural language description
CN118196242A
Long sequence dance generation method based on global and local action optimization
CN118552674A
Multi-person interaction action generation method and device and electronic equipment
CN119024971A
Human body action sequence generation method and device, equipment and storage medium
CN120088867A