A method and system for training a robot autonomous operation model, and an electronic device
Patent Information
- Application Number
- CN202610896696.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0007]为了克服现有超声机器人作业不成熟,常规VLA架构存在决策短视、语义动作表征断层、动态自适应校正能力缺失,难以满足高精度、长时序、高鲁棒性的超声自主作业需求,本发明提供了一种机器人自主操作模型的训练方法、系统及电子设备
[0041]1、本发明通过构建多模态联合训练体系,实现语言指令、超声图像与机器人动作的智能化映射,摆脱了传统超声作业高度依赖操作人员临床经验的现状。摒弃了传统人工操作的主观差异性问题,能够在各类复杂扫描与介入任务中保持稳定的作业逻辑与成像效果,大幅提升超声诊疗作业的流程标准化程度与图像重复性;同时克服了传统离线规划方案仅适用于静态场景的局限,具备实时场景感知与在线推理能力,可适应作业过程中的动态扰动工况,避免预设路径失效导致的操作偏差与安全隐患,提升作业安全性与场景适配性;
Smart Images

Figure CN122433837B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot intelligent control technology, and more specifically, to a training method, system, and electronic device for a robot autonomous operation model. Background Technology
[0002] With the rapid development of embodied intelligence technology, the Vision-Language-Motion (VLA) architecture has become the mainstream solution for robots to achieve intelligent autonomous operation. This architecture relies on a vision-language model to complete scene understanding and command parsing, enabling end-to-end mapping from external perceived information to robot action commands, providing a technological foundation for automated ultrasound robot operations. Currently, clinical ultrasound operations are still mainly performed manually, with operational quality highly dependent on operator experience, resulting in significant subjective differences. Image consistency in complex scans and interventional tasks is poor, making it difficult to establish standardized operating procedures. Furthermore, existing automation solutions based on preoperative 3D modeling and offline trajectory planning can only adapt to fixed, static scenes and lack real-time environmental perception and online reasoning capabilities. Once dynamic disturbances occur in the operational scene, the preset operational path is prone to failure, leading to operational deviations and posing certain operational safety hazards. In addition, the traditional VLA framework has further inherent technical defects when applied to long-term, high-precision dynamic ultrasound operation scenarios.
[0003] First, traditional VLA models generally suffer from short-sighted decision-making and lack long-term task planning capabilities. Existing models mostly generate independent actions based on instantaneous observations at a single moment, failing to effectively utilize historical temporal state information for experience accumulation or predict subsequent scenario evolution trends. For tasks like continuous ultrasound scanning, which involve multiple coupled steps and strong temporal correlations, a single instantaneous decision-making approach easily leads to fragmented action sequences and disjointed execution logic, further reducing the effectiveness of standardized operations and the quality of task completion.
[0004] Secondly, traditional VLA architecture suffers from a disconnect between high-level semantics and low-level action representation. The semantic features output by the visual-language model focus on scene and instruction understanding, which differs significantly from the action feature dimensions and representation logic required for robot motion control. This results in a lack of efficient cross-modal fusion and feature bridging mechanisms. Semantic information is prone to feature loss and bias during the conversion to action policies, failing to accurately match real-world operational needs. Ultimately, this leads to inconsistencies between robot actions and human intentions, and limited control precision.
[0005] Finally, traditional VLA models lack dynamic robustness and adaptive correction capabilities. Real-world ultrasonic operation scenarios are constantly changing, with ongoing environmental disturbances and target state variations. Existing models lack explicit future state prediction and dynamic deviation feedback mechanisms, making it impossible to anticipate scenario changes and correct action strategies online. Fixed prediction and execution logic struggles to adapt to dynamic operating conditions, exhibits weak anti-interference capabilities, and cannot cope with dynamic uncertainties during operations, making it difficult to maintain stable and standardized autonomous operation results over long periods.
[0006] In summary, existing ultrasonic robot operation schemes suffer from high reliance on manual labor, low standardization, poor adaptability to offline planning, and technical problems such as weak long-term planning capabilities of VLA architecture, failure of cross-modal feature mapping, and insufficient dynamic adaptive capabilities. They cannot meet the requirements of ultrasonic robots for high-precision, continuous, and robust long-term autonomous operation. Therefore, it is urgent to design a new robot autonomous operation model training scheme. Summary of the Invention
[0007] To overcome the immaturity of existing ultrasonic robot operations, the shortcomings of conventional VLA architectures such as short-sighted decision-making, fragmented semantic action representation, and lack of dynamic adaptive correction capabilities, which make it difficult to meet the requirements of high-precision, long-term, and highly robust ultrasonic autonomous operations, this invention provides a training method, system, and electronic equipment for a robot autonomous operation model.
[0008] The technical solution of this invention is as follows:
[0009] A method for training an autonomous operation model for a robot includes the following steps:
[0010] Image sequence samples containing multiple ultrasound images and language instruction samples containing clinical operation scripts are collected and paired to form a first training dataset. Based on the first training dataset, the pre-trained visual-language model VLM is fine-tuned and trained with instructions, and the annotated scenario description text is used as a supervision label to train and generate a scenario multimodal perception module.
[0011] Historical state sequences are selected as input samples, and multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data are selected as labels. A second training dataset is constructed based on the input samples and labels. The video diffusion model (VDM) is fine-tuned based on the second training dataset to obtain the scenario dynamics perception module. The historical state sequence is any one or a combination of two of the following: a continuous ultrasonic image frame sequence and a continuous state vector sequence of the robot body.
[0012] The context representation samples, dynamic prediction samples, robot real-time state samples, and historical action sequence samples are combined into input features, and continuous action blocks of preset duration are used as action block labels to construct a third training dataset; all network parameters of the context multimodal perception module and the context dynamics perception module are fixed and frozen, and the action generation module is obtained by training the policy network to be trained based on the third training dataset.
[0013] The training steps of the action generation module include: feeding four types of input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a bias reward term representing the organization tracking error.
[0014] Preferably, the process of training and generating a multimodal perception module for a given context includes:
[0015] Load the pre-trained VLM base weights and freeze the parameters of the VLM backbone visual encoding layer, open text interaction layer and multimodal mapping layer;
[0016] The ultrasound images and language instructions of a single pair of samples in the first training dataset are input into the model encoder, and the scenario description text bound to the samples is used as the ground truth label to construct a multimodal cross-entropy loss function.
[0017] The parameters of the open layer are iteratively optimized using mini-batch gradient descent. After the preset convergence condition is met, the weights are exported to form a scenario multimodal perception module.
[0018] Preferably, the future scene trajectory sample is derived sequentially from the end of the historical state time series. The image and pose of the frame are jointly predicted; the action block labels are continuous. A set of robotic arm joint motion commands for step length. >1.
[0019] Preferably, the process of fine-tuning the video diffusion model (VDM) based on the second training dataset to obtain the scene dynamics perception module includes:
[0020] Load the weights of the pre-trained video diffusion model and fix the underlying feature extraction parameters and open the temporal prediction branch parameters;
[0021] Input the historical state sequence consisting of consecutive frames in the second training dataset into the VDM encoder, and construct the diffusion loss function using the future scene trajectory of the time-sequentially delayed τ frames as the ground value label.
[0022] The parameters of the time series prediction branch are updated iteratively using a batch optimization algorithm, and the model weights are exported after convergence to form the scenario dynamics perception module.
[0023] Preferably, the process of multi-source feature fusion through cross-modal attention mechanisms includes:
[0024] The four types of features—contextual representation, dynamic prediction representation, robot current state, and historical action sequence—are dimensional alignment mappings. Based on multi-head cross-modal attention, the correlation weight coefficients between different features are calculated. Based on the weight coefficients, the modal features are weighted and aggregated to obtain a fused feature vector.
[0025] Preferably, the process of mapping the predicted action block to a low-dimensional action flow space based on the action flow matching mechanism includes:
[0026] The layered fully connected network performs dimensionality transformation and dimensionality reduction on the fused feature vector layer by layer, mapping the high-dimensional fused feature constraints to the predefined dimension of the standardized action flow latent space.
[0027] Match the standard action flow distribution constraints corresponding to historical samples within the action flow latent space to correct the latent space feature offset;
[0028] The corrected latent space features are decoded and restored to continuous. Step robotic arm joint command sequence, consisting of continuous Step joint command sequences are combined to form a single predicted action block.
[0029] Preferably, the multidimensional reward function further includes a clinical constraint reward term; the clinical constraint reward term includes an imaging quality reward based on the ultrasound image signal-to-noise ratio, contrast, and artifact level, and a limit reward based on the probe contact force safety threshold.
[0030] Preferably, the primary optimization objective is to minimize the error between the predicted action block and the action block label. Simultaneously, a multi-dimensional reward function is introduced for joint auxiliary supervision. The iterative update of the policy network parameters includes:
[0031] The basic regression loss term is constructed based on the deviation between the predicted action block and the true action block label. The overall composite optimization objective is obtained by weighting the smoothness reward term, the tracking deviation reward term, and the clinical constraint reward term.
[0032] The gradient values of the composite optimization objective with respect to the policy network parameters are obtained by relying on the gradient backpropagation mechanism, and the network parameters are iteratively updated round by round based on a preset step size;
[0033] The iterative process is repeated until the composite optimization objective converges to the set threshold, at which point the parameter optimization process is terminated.
[0034] A second aspect of the present invention also provides a training system for a robot autonomous operation model, characterized in that it includes a context multimodal perception training module, a context dynamics perception training module, and an action generation training module, wherein:
[0035] The scenario multimodal perception training module is used to collect image sequence samples containing multiple frames of ultrasound images and language instruction samples containing clinical operation language, and pair them to form the first training dataset; based on the first training dataset, the pre-trained visual-language model VLM is fine-tuned and trained with instructions, and the labeled scenario description text is used as supervision labels to train and generate the scenario multimodal perception module.
[0036] The scenario dynamics perception training module is used to select historical state sequences as input samples, select multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data as labels, and construct a second training dataset based on the input samples and labels; and fine-tune the video diffusion model (VDM) based on the second training dataset to obtain the scenario dynamics perception module; wherein, the historical state sequence is any one or a combination of two of the following: a continuous ultrasound image frame sequence and a continuous state vector sequence of the robot body;
[0037] The action generation training module is used to combine scenario context representation samples, dynamic prediction representation samples, robot real-time state samples, and historical action sequence samples into input features, and to use continuous action blocks of preset duration as action block labels to build a third training dataset; fix and freeze all network parameters of the scenario multimodal perception module and scenario dynamics perception module, and train the policy network to be trained based on the third training dataset to obtain the action generation module.
[0038] The training steps of the action generation module include: feeding four types of input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a bias reward term representing the organization tracking error.
[0039] A third aspect of the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the training method for the autonomous operation model of a robot as described above.
[0040] According to the above-described solution, the beneficial effects of this invention are as follows:
[0041] 1. This invention constructs a multimodal joint training system to achieve intelligent mapping of language commands, ultrasound images, and robot actions, thus overcoming the traditional reliance on operators' clinical experience in ultrasound operations. It eliminates the subjective variability inherent in traditional manual operations, maintaining stable operational logic and imaging effects in various complex scanning and interventional tasks, significantly improving the standardization of ultrasound diagnostic and treatment processes and image repeatability. Simultaneously, it overcomes the limitation of traditional offline planning schemes being only applicable to static scenarios, possessing real-time scene perception and online reasoning capabilities, adapting to dynamic disturbances during operations, avoiding operational deviations and safety hazards caused by preset path failures, and improving operational safety and scenario adaptability.
[0042] 2. This invention fully utilizes historical temporal state information during model training and introduces a future scene temporal prediction supervision mechanism, changing the short-sighted decision-making mode of traditional VLA models that generate single-step actions based solely on instantaneous observations of a single frame; it can accumulate historical operational experience and predict future scene evolution trends, ensuring the action continuity and logical integrity of multi-step, long-time-series continuous ultrasonic operations, effectively avoiding problems such as action sequence fragmentation and task interruption, and significantly improving the overall execution quality and completion of complex long-time-domain tasks;
[0043] 3. This invention constructs a dedicated multimodal cross-modal feature fusion and mapping mechanism, which performs unified dimensional adaptation and correlation fusion of multiple heterogeneous features, effectively making up for the dimensional differences and information gaps between high-level semantic representation and low-level action representation in the traditional VLA architecture; reducing feature loss and offset in the process of converting semantic understanding into action strategy, achieving accurate matching of clinical operation intention, scene semantics and robot action output, greatly improving the robot operation control accuracy, and ensuring that intelligent execution actions meet the actual clinical operation needs;
[0044] 4. This invention introduces a dynamic scene temporal prediction training and multi-dimensional clinical joint supervision and optimization mechanism, enabling the model to predict and correct deviations in the dynamic changes of the scene. It overcomes the shortcomings of traditional models that cannot adapt to dynamic environmental disturbances and have fixed strategies. It can adapt to the dynamic changes of the scene in real time during the operation, adjust the operation action strategy online, and continuously maintain a stable and standardized autonomous operation state, effectively improving the anti-interference ability and operation stability of the ultrasound robot in dynamic clinical conditions. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram of the training method provided by the present invention. Detailed Implementation
[0047] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0048] It should be noted that when a component is referred to as "fixed," "set," or "connected" to another component, it may be located directly or indirectly on that other component. The terms "upper," "lower," "left," "right," "front," "rear," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or position based on the accompanying drawings, and are for ease of description only, and should not be construed as limiting the technical solution. The terms "first," "second," etc., are used for ease of description only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features. "Many" means two or more, unless otherwise explicitly specified. "Several" means one or more, unless otherwise explicitly specified.
[0049] This invention provides a method for training a robot autonomous operation model, see [link to relevant documentation]. Figure 1 This includes the following steps:
[0050] S1. Collect image sequence samples containing multiple ultrasound images and language instruction samples containing clinical operation scripts, and pair them to form a first training dataset; perform instruction fine-tuning training on the pre-trained visual-language model VLM based on the first training dataset, and use the labeled scenario description text as supervision labels to train and generate a scenario multimodal perception module.
[0051] S2. Select a historical state sequence as an input sample, and select multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data as labels. Construct a second training dataset based on the input sample and the labels. Fine-tune the video diffusion model (VDM) based on the second training dataset to obtain the scenario dynamics perception module. The historical state sequence is any one or a combination of two of the following: a continuous ultrasonic image frame sequence and a continuous state vector sequence of the robot body.
[0052] S3. Combine the context representation samples, dynamic prediction samples, robot real-time state samples, and historical action sequence samples into input features, and use the preset duration continuous action blocks as action block labels to build a third training dataset; fix and freeze all network parameters of the context multimodal perception module and the context dynamics perception module, and train the policy network to be trained based on the third training dataset to obtain the action generation module.
[0053] The training steps of the action generation module include: feeding four types of input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a bias reward term representing the organization tracking error.
[0054] In some embodiments, the process of performing step S1 includes:
[0055] S1.1 Construct the first training dataset;
[0056] S1.2 Training and generating a multimodal perception module for scenarios.
[0057] Specifically, the process of executing step S1.1 includes: screening continuous image sequences under standard clinical ultrasound scanning scenarios and matching them with standardized clinical operation instructions for the corresponding scenarios; preprocessing the ultrasound images to perform noise reduction, size normalization, and frame sequence alignment; cleaning and semantic normalization of the clinical operation instructions and generating accurate scenario description text supervision labels by matching them with the corresponding image scenarios; binding and pairing the preprocessed ultrasound image sequences, clinical language instructions, and scenario description labels one by one, and then summarizing and organizing them to construct the first training dataset.
[0058] Further, in step S1.2, the pre-trained VLM basic weights are loaded and the parameters of the VLM backbone visual coding layer, open text interaction layer and multimodal mapping layer are frozen. Based on transfer learning, the general visual semantic basic capabilities of the model are preserved, while the features of ultrasound clinical scenarios are specifically adapted.
[0059] The ultrasound images and language instructions of a single pair of samples in the first training dataset are input into the model encoder, and the scenario description text bound to the samples is used as the ground truth label to construct a multimodal cross-entropy loss function.
[0060] The open layer network parameters are iteratively optimized using a mini-batch gradient descent algorithm to continuously reduce the deviation between the model's predicted output and the ground truth label. After reaching the preset convergence condition, the optimal model weights are derived, and finally, a scenario-based multimodal perception module adapted to ultrasound clinical scenarios is trained.
[0061] In fact, by locking the core visual encoding parameters and only fine-tuning the training strategy of interaction and mapping layers, the model can be quickly adapted to ultrasound clinical scenarios while retaining its general perception capabilities. This effectively solves the problems of poor adaptability of general visual language models to medical scenarios and the disconnect between semantic understanding and clinical context. It improves the robot's cross-modal perception accuracy of clinical language commands and ultrasound scenarios, and provides a reliable high-level semantic perception foundation for subsequent dynamic scene prediction and precise action generation.
[0062] In some embodiments, the process of performing step S2 includes:
[0063] S2.1 Construct the second training dataset;
[0064] S2.2 Training and generating the scenario dynamics perception module.
[0065] Specifically, the process of executing step S2.1 includes: acquiring a continuous historical ultrasound image sequence and a corresponding continuous state vector sequence of the robot body during ultrasound scanning, combining them to form multiple types of historical state sequence samples; and selecting samples sequentially from the end of the historical state time sequence. frame( >1) The image and robotic arm pose joint sequence constitutes the future scene trajectory of multiple consecutive future image frames and the corresponding robot pose data fusion, which serves as the ground truth label;
[0066] The future image frames and the corresponding robot pose data together constitute a temporal state sequence, wherein the data at each moment includes the ultrasound image and the pose of the robotic arm end effector. This sequence serves as a supervision label during the training of the video diffusion model, constraining the model's predictive output of future states.
[0067] Furthermore, temporal insertion weights are configured for future scene trajectories at different time positions to differentiate the supervision intensity of trajectories at different times; the collected historical state samples and future scene trajectory labels are temporally aligned, data normalized, and sample cleaned to remove abnormal perturbation samples, and standardized sample pairing and summarization are completed to finally construct a second training dataset for temporal dynamic prediction training.
[0068] Specifically, the process of executing step S2.2 includes: loading the pre-trained video diffusion model (VDM) weights, fixing the model's underlying feature extraction parameters, and opening the temporal prediction branch parameters; inputting the historical state sequence consisting of consecutive frames from the second training dataset into the VDM encoder, and sequentially... The future scene trajectory of the frame is used as the ground truth label to construct the diffusion loss function; the batch optimization algorithm is used to iteratively update the temporal prediction branch parameters to continuously optimize the model's temporal prediction accuracy. After the training reaches the preset convergence condition, the model weights are exported, and finally a scene dynamics perception module with dynamic temporal inference capability is formed.
[0069] In fact, by matching long-term historical states with future multi-frame trajectory labels and configuring differentiated temporal weight supervision, the correlation between scene dynamic evolution and robot motion during ultrasound scanning can be accurately discovered. Relying on the underlying feature capabilities of the pre-trained model and optimizing the temporal prediction branch in a targeted manner, it can effectively solve the problems of short-sighted decision-making and inability to predict dynamic disturbances in clinical scenarios in traditional models. It breaks through the limitations of single-frame instantaneous perception, realizes accurate prediction of future scene temporal trajectories, and provides core dynamic prediction support for the robot to output continuous, stable, and scene-adaptive autonomous operation actions, greatly improving the adaptability of operation in dynamic clinical conditions.
[0070] In some embodiments, the process of performing step S3 includes:
[0071] S3.1 Construct the third training dataset;
[0072] S3.2 Training Action Generation Module.
[0073] Specifically, in the process of executing step S3.1, the following steps are included: calling the trained scenario multimodal perception module and scenario dynamics perception module to extract scenario context representation and dynamic prediction representation respectively, synchronously collecting real-time robot state data and historical action sequence data, and integrating four types of features as model input features; at the same time, extracting continuous robotic arm action instructions of a preset duration as action block ground truth labels, standardizing and temporally aligning all input features and action labels, removing abnormal samples, and summarizing and pairing to form the third training dataset.
[0074] Furthermore, the process of executing step S3.2 includes:
[0075] S3.2.1 Multi-source feature fusion;
[0076] S3.2.2, Generate predicted action blocks;
[0077] S3.2.3 Iterative update strategy network parameters.
[0078] Specifically, in the process of executing step S3.2.1, the following steps are included: the policy network to be trained performs dimension alignment mapping on four types of features: context representation, dynamic prediction representation, robot current state, and historical action sequence, to unify the dimensions of each modal feature; and calculates the correlation weight coefficients between different features through the multi-head cross-modal attention mechanism built into the policy network, and performs weighted aggregation of each modal feature based on the weight coefficients to complete the fusion of multi-source features and obtain a global fused feature vector.
[0079] Further, step S3.2.2 is executed, whereby the fused feature vector is subjected to dimensionality transformation and dimensionality reduction layer by layer based on the hierarchically arranged fully connected network, mapping the high-dimensional fused feature constraints to a predefined dimension standardized action flow latent space; within the action flow latent space, the standard action flow distribution constraints corresponding to historical samples are matched to correct the latent space feature offset; and the corrected latent space features are decoded and restored to continuous... Step robotic arm joint command sequence, consisting of continuous Step joint command sequences are combined to form a single predicted action block.
[0080] Specifically, the execution of step S3.2.3 includes: constructing a basic regression loss term based on the deviation between the predicted action block and the true action block label, and weighting it with the smoothness reward term, the tracking deviation reward term and the clinical constraint reward term to obtain the overall composite optimization objective;
[0081] The overall composite optimization objective formula is as follows:
[0082] ;
[0083] In the formula, For the overall composite optimization goal, To predict the basic regression loss term between the action block and the ground truth action block, For smoothness bonus items, To track deviation reward items, This is a clinical constraint reward item. This refers to the preset weighted balance coefficients for each reward item.
[0084] The basic regression loss term uses the mean squared error loss, as shown in the following formula:
[0085] ;
[0086] In the formula, The number of action steps contained within the action block. For the first Predicting actions step by step For the first Step truth value action tag.
[0087] Furthermore, the gradient values of the composite optimization objective with respect to the policy network parameters are obtained by relying on the gradient backpropagation mechanism, and the network parameters are iteratively updated round by round based on a preset step size; the iterative process is repeated until the composite optimization objective converges to the set threshold, and the parameter optimization process is terminated.
[0088] In summary, this invention constructs an end-to-end robotic ultrasound autonomous operation model training system that integrates multimodal perception, temporal dynamic prediction, and intelligent action generation. It completes specialized training for multimodal semantic perception, scene dynamics deduction, and autonomous action decision-making models in a layered and step-by-step manner, with each layer progressing and performing its own function. This effectively solves the technical defects of traditional ultrasound robot autonomous operation, such as scene semantic understanding deviation, lack of dynamic scene prediction capability, stiff action generation, insufficient tissue tracking accuracy, and poor clinical adaptability.
[0089] A second aspect of the present invention also provides a training system for a robot autonomous operation model, characterized in that it includes a context multimodal perception training module, a context dynamics perception training module, and an action generation training module, wherein:
[0090] The scenario multimodal perception training module is used to collect image sequence samples containing multiple frames of ultrasound images and language instruction samples containing clinical operation language, and pair them to form the first training dataset; based on the first training dataset, the pre-trained visual-language model VLM is fine-tuned and trained with instructions, and the labeled scenario description text is used as supervision labels to train and generate the scenario multimodal perception module.
[0091] The scenario dynamics perception training module is used to select historical state sequences as input samples, select multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data as labels, and construct a second training dataset based on the input samples and labels; and fine-tune the video diffusion model (VDM) based on the second training dataset to obtain the scenario dynamics perception module; wherein, the historical state sequence is any one or a combination of two of the following: a continuous ultrasound image frame sequence and a continuous state vector sequence of the robot body;
[0092] The action generation training module is used to combine scenario context representation samples, dynamic prediction representation samples, robot real-time state samples, and historical action sequence samples into input features, and to use continuous action blocks of preset duration as action block labels to build a third training dataset; fix and freeze all network parameters of the scenario multimodal perception module and scenario dynamics perception module, and train the policy network to be trained based on the third training dataset to obtain the action generation module.
[0093] The training steps of the action generation module include: feeding four types of input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a bias reward term representing the organization tracking error.
[0094] A third aspect of the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the training method for the autonomous operation model of a robot as described above.
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for training an autonomous operation model of a robot, characterized in that, Includes the following steps: Image sequence samples containing multiple ultrasound images and language instruction samples containing clinical operation scripts are collected and paired to form a first training dataset. Based on the first training dataset, the pre-trained visual-language model VLM is fine-tuned and trained with instructions, and the annotated scenario description text is used as a supervision label to train and generate a scenario multimodal perception module. Historical state sequences are selected as input samples, and multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data are selected as labels. A second training dataset is constructed based on the input samples and labels. The video diffusion model (VDM) is fine-tuned based on the second training dataset to obtain the scenario dynamics perception module. The historical state sequence is any one or a combination of two of the following: a continuous ultrasonic image frame sequence and a continuous state vector sequence of the robot body. The context representation samples, dynamic prediction samples, robot real-time state samples, and historical action sequence samples are combined into input features, and continuous action blocks of preset duration are used as action block labels to construct a third training dataset; all network parameters of the context multimodal perception module and the context dynamics perception module are fixed and frozen, and the action generation module is obtained by training the policy network to be trained based on the third training dataset. The training steps of the action generation module include: feeding the input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a deviation reward term representing the tissue tracking error, and also includes a clinical constraint reward term; the clinical constraint reward term includes an imaging quality reward based on the signal-to-noise ratio, contrast, and artifact degree of the ultrasound image, and a limit reward based on the probe contact force safety threshold.
2. The training method for a robot autonomous operation model according to claim 1, characterized in that, The training process for the multimodal perception module for generating contexts includes: Load the pre-trained VLM base weights and freeze the parameters of the VLM backbone visual encoding layer, open text interaction layer and multimodal mapping layer; The ultrasound images and language instructions of a single pair of samples in the first training dataset are input into the model encoder, and the scenario description text bound to the samples is used as the ground truth label to construct a multimodal cross-entropy loss function. The parameters of the open layer are iteratively optimized using mini-batch gradient descent. After the preset convergence condition is met, the weights are exported to form a scenario multimodal perception module.
3. The training method for a robot autonomous operation model according to claim 1, characterized in that, The consecutive future image frames and their corresponding robot pose data are sequentially ordered from the end of the historical state time sequence. The image and pose of the frame are jointly predicted; the action block labels are continuous. A set of robotic arm joint motion commands for step length. >
1.
4. The training method for a robot autonomous operation model according to claim 3, characterized in that, The process of fine-tuning the video diffusion model (VDM) based on the second training dataset to obtain the context dynamics perception module includes: Load the weights of the pre-trained video diffusion model and fix the underlying feature extraction parameters and open the temporal prediction branch parameters; The historical state sequence consisting of consecutive frames in the second training dataset is input into the VDM encoder, and the image and pose joint prediction sequence of the temporally delayed τ frames is used as the ground value label to construct the diffusion loss function. The parameters of the time series prediction branch are updated iteratively using a batch optimization algorithm, and the model weights are exported after convergence to form the scenario dynamics perception module.
5. The training method for a robot autonomous operation model according to claim 1, characterized in that, The process of multi-source feature fusion through cross-modal attention mechanisms includes: The four types of features—contextual representation, dynamic prediction representation, robot current state, and historical action sequence—are dimensional alignment mappings. Based on multi-head cross-modal attention, the correlation weight coefficients between different features are calculated. Based on the weight coefficients, the modal features are weighted and aggregated to obtain a fused feature vector.
6. The training method for a robot autonomous operation model according to claim 1, characterized in that, The process of mapping the predicted action block to a low-dimensional action flow space based on the action flow matching mechanism includes: The layered fully connected network performs dimensionality transformation and dimensionality reduction on the fused feature vector layer by layer, mapping the high-dimensional fused feature constraints to the predefined dimension of the standardized action flow latent space. Match the standard action flow distribution constraints corresponding to historical samples within the action flow latent space to correct the latent space feature offset; The corrected latent space features are decoded and restored to continuous. Step robotic arm joint command sequence, consisting of continuous Step joint command sequences are combined to form a single predicted action block.
7. The training method for a robot autonomous operation model according to claim 1, characterized in that, The primary optimization objective is to minimize the error between the predicted action block and its label. A multi-dimensional reward function is introduced for joint auxiliary supervision. The iterative update of the policy network parameters includes: The basic regression loss term is constructed based on the deviation between the predicted action block and the true action block label. The overall composite optimization objective is obtained by weighting the smoothness reward term, the tracking deviation reward term, and the clinical constraint reward term. The gradient values of the composite optimization objective with respect to the policy network parameters are obtained by relying on the gradient backpropagation mechanism, and the network parameters are iteratively updated round by round based on a preset step size; The iterative process is repeated until the composite optimization objective converges to the set threshold, at which point the parameter optimization process is terminated.
8. A training system for an autonomous operation model of a robot, characterized in that, It includes a context-multimodal perception training module, a context-dynamics perception training module, and an action generation training module, among which: The scenario multimodal perception training module is used to collect image sequence samples containing multiple frames of ultrasound images and language instruction samples containing clinical operation language, and pair them to form the first training dataset; based on the first training dataset, the pre-trained visual-language model VLM is fine-tuned and trained with instructions, and the labeled scenario description text is used as supervision labels to train and generate the scenario multimodal perception module. The scenario dynamics perception training module is used to select historical state sequences as input samples, select multiple consecutive future image frames that are temporally continuous with the historical state sequence and the corresponding robot pose data as labels, and construct a second training dataset based on the input samples and labels; and fine-tune the video diffusion model (VDM) based on the second training dataset to obtain the scenario dynamics perception module; wherein, the historical state sequence is any one or a combination of two of the following: a continuous ultrasound image frame sequence and a continuous state vector sequence of the robot body; The action generation training module is used to combine scenario context representation samples, dynamic prediction representation samples, robot real-time state samples, and historical action sequence samples into input features, and to use continuous action blocks of preset duration as action block labels to build a third training dataset; fix and freeze all network parameters of the scenario multimodal perception module and scenario dynamics perception module, and train the policy network to be trained based on the third training dataset to obtain the action generation module. The training steps of the action generation module include: feeding four types of input features into the policy network to be trained, completing multi-source feature fusion through a cross-modal attention mechanism, and mapping the predicted action blocks to a low-dimensional action flow space based on an action flow matching mechanism; taking minimizing the error between the predicted action block and the action block label as the first optimization objective, while introducing a multi-dimensional reward function for joint auxiliary supervision, and iteratively updating the policy network parameters; the multi-dimensional reward function includes at least a smoothness reward term representing the stability of the action and a deviation reward term representing the tissue tracking error, and also includes a clinical constraint reward term; the clinical constraint reward term includes an imaging quality reward based on the signal-to-noise ratio, contrast, and artifact degree of the ultrasound image, and a limit reward based on the probe contact force safety threshold.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the training method for the robot autonomous operation model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Robot control method and device, electronic equipment and medium
CN121374583A
Method for constructing visual language action model based on wash Adaboost
CN121598175A