VLA model generation method and device of surgical robot, equipment and medium

By training the visual language action model and using reinforcement learning to optimize the motion head parameters of the surgical robot, the problem that the existing medical AI system is unable to adjust surgical strategies in real time is solved. The high precision and adaptability of the surgical robot in complex environments are achieved, the deployment cost is reduced, and it is suitable for primary medical institutions.

CN120612733APending Publication Date: 2025-09-09PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510704263.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing AI systems in the medical field are unable to provide real-time feedback and adjust surgical strategies during surgery, resulting in poor surgical results. The model also has weak generalization capabilities, making it difficult to cope with complex and changing clinical situations. It is unable to fully combine multimodal information for comprehensive judgment, and the deployment cost is high, which limits its application in primary medical institutions.

Method used

By acquiring surgical video frames and natural language instructions to train the initial visual-language-action model, the reinforcement learning algorithm is used to optimize the action head parameters, and the model is trained in combination with online datasets to form a closed-loop visual-language-action model, enabling the surgical robot to adjust its strategy in real time.

Benefits of technology

It improves the surgical accuracy and adaptability of surgical robots, enables them to adjust surgical strategies based on real-time feedback, enhances the generalization ability of the model, reduces deployment costs, and is suitable for primary medical institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612733A_ABST
    Figure CN120612733A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence and medical treatment, and discloses a VLA model generation method and device for a surgical robot, equipment and a medium, and the method comprises the steps: inputting a first data subset into a to-be-trained visual language action model for training, and obtaining an initial visual language action model; inputting the first new operation task data into the initial visual language action model, and controlling the operation robot to output and execute an operation action according to the initial visual language action model; utilizing a preset reinforcement learning algorithm to optimize action head parameters of the initial visual language action model to obtain a first visual language action model; storing task trajectory data successfully completed by the surgical robot to an online data set; and inputting the second data subset and the task trajectory data into the first visual language action model for training to obtain a second visual language action model, thereby continuously performing reinforcement learning on the VLA model, enabling the surgical robot to adjust a surgical strategy according to real-time feedback, and further improving the surgical accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and medical technology, and in particular to a VLA model generation method, device, equipment and medium for a surgical robot. Background Art

[0002] In today's medical field, artificial intelligence systems have been widely used in imaging diagnosis, surgical navigation, personalized treatment recommendations and other aspects. However, current AI systems in the medical field mostly rely on static supervised learning, which has many limitations. For example, static supervised learning requires a large amount of labeled data from high-quality experts, but the cost of labeled data is high, and the case diversity is insufficient, resulting in weak generalization ability of the trained model, making it difficult to cope with complex and changing clinical situations. In addition, it is impossible to adjust strategies through real-time interaction (for example, intraoperative feedback, changes in patient physiological signals). When faced with emergencies or new cases, the system finds it difficult to respond promptly and effectively, affecting the surgical effect. At the same time, multimodal information such as medical images, text medical records, and sensor data cannot be fully jointly modeled, limiting the system's comprehensive judgment and decision-making on the condition. In addition, fine-tuning all parameters of large models requires high computing power support, which is difficult to deploy on local medical equipment, limiting the promotion and application of artificial intelligence technology in primary medical institutions. Summary of the Invention

[0003] The embodiments of the present invention provide a VLA model generation method, device, equipment and medium for a surgical robot, aiming to solve the problem in the prior art that the surgical robot cannot provide real-time feedback to adjust the surgical strategy, thereby affecting the surgical effect.

[0004] In a first aspect, an embodiment of the present invention provides a method for generating a VLA model of a surgical robot, the method comprising:

[0005] Obtaining a first data subset, and inputting the first data subset into a visual language action model to be trained for model training to obtain an initial visual language action model; wherein the first data subset includes surgical video frames, natural language instructions, and a first surgical action label;

[0006] Acquiring first new surgical task data, inputting the first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform a surgical action according to an action label output by the initial visual language action model;

[0007] Optimizing the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model;

[0008] storing the task trajectory data of the surgical robot successfully performing the surgical action into an online dataset;

[0009] A second data subset is obtained, and the second data subset and the task trajectory data in the online data set are input into the first visual language action model for model training to obtain a second visual language action model.

[0010] In a second aspect, an embodiment of the present invention further provides a VLA model generation device for a surgical robot, the device comprising:

[0011] A first training unit is configured to obtain a first data subset and input the first data subset into a visual language action model to be trained for model training to obtain an initial visual language action model; wherein the first data subset includes surgical video frames, natural language instructions, and a first surgical action label;

[0012] an execution unit, configured to obtain first new surgical task data, input the first new surgical task data into the initial visual language action model, and control the surgical robot to perform a surgical action according to an action label output by the initial visual language action model;

[0013] a reinforcement learning unit, configured to optimize the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model;

[0014] A storage unit, configured to store the task trajectory data of the surgical robot successfully performing the surgical action into an online data set;

[0015] The second training unit is used to obtain a second data subset, input the second data subset and the task trajectory data in the online data set into the first visual language action model for model training, and obtain a second visual language action model.

[0016] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the VLA model generation method for the surgical robot as described in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the VLA model generation method of the surgical robot described in the first aspect.

[0018] An embodiment of the present invention provides a VLA model generation method, device, equipment and medium for a surgical robot, which obtains an initial visual language action model by inputting a first data subset into a visual language action model to be trained; inputting a first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform surgical actions according to the output of the initial visual language action model; optimizing the action head parameters of the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model; storing task trajectory data successfully completed by the surgical robot into an online data set; inputting a second data subset and task trajectory data into the first visual language action model for training to obtain a second visual language action model, thereby continuously reinforcing the learning of the VLA model, so that the surgical robot can adjust the surgical strategy according to real-time feedback, thereby improving surgical accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A schematic diagram of a first flow chart of a method for generating a VLA model of a surgical robot provided by an embodiment of the present invention;

[0021] Figure 2 A second flow chart of the method for generating a VLA model of a surgical robot provided by an embodiment of the present invention;

[0022] Figure 3 A schematic block diagram of a VLA model generation device for a surgical robot provided in an embodiment of the present invention;

[0023] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0027] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] A VLA model generation method for a surgical robot provided by an embodiment of the present invention is specifically applicable to medical surgical application scenarios, such as cholecystectomy, etc.; and the VLA model generation method for the surgical robot can be applied to a server, and the server obtains a first data subset, inputs the first data subset into the visual language action model to be trained for model training, and obtains an initial visual language action model; wherein the first data subset includes surgical video frames, natural language instructions and first surgical action labels; obtains first new surgical task data, inputs the first new surgical task data into the initial visual language action model, and controls the surgical robot to perform surgical actions according to the action labels output by the initial visual language action model; uses a preset reinforcement learning algorithm to optimize the action head parameters in the initial visual language action model to obtain a first visual language action model; stores the task trajectory data of the surgical robot successfully performing the surgical action in an online dataset; obtains a second data subset, inputs the second data subset and the task trajectory data in the online dataset into the first visual language action model for model training, and obtains a second visual language action model. The VLA model, short for Vision-Language-Action (VLA), is a multimodal AI model that integrates visual perception, language understanding, and action generation. It can process images, text instructions, and operational tasks in complex scenarios, forming a closed loop from perception to decision-making and execution. Furthermore, the server can be a standalone server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and AI platforms.

[0029] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0030] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0031] The present invention is described in detail below through specific examples.

[0032] See also Figure 1 and Figure 2 , Figure 1 1 is a schematic diagram of a first flow chart of a method for generating a VLA model of a surgical robot provided by an embodiment of the present invention; Figure 2 FIG. 2 is a second flow chart of the VLA model generation method for a surgical robot provided by an embodiment of the present invention. Figure 1 and Figure 2 As shown, the method includes the following steps S110-S150.

[0033] S110 , obtaining a first data subset, and inputting the first data subset into a vision-language-action model to be trained for model training to obtain an initial vision-language-action model.

[0034] In this embodiment, to enable the initial visual-language action model to initially understand the first data subset and perform the surgical task in the first data subset, this embodiment specifically inputs the first data subset into a pre-constructed visual-language action model to be trained for model training to obtain an initial visual-language action model. The first data subset is obtained from a preset expert dataset and includes surgical video frames, natural language instructions, and first surgical action labels corresponding to the surgical video frames and natural language instructions. The surgical video frames refer to visual input during the surgical procedure, such as the operation of surgical instruments and changes in the surgical site. The natural language instructions refer to voice instructions that an expert may issue during the procedure, such as "cut the skin." The first surgical action labels are precise action annotations for each surgical video frame and corresponding natural language instruction combination based on standard surgical operating procedures and expert experience. These labels may include rotation data for each control joint of the surgical robot, gripper grip / release status, and other action data during the procedure. The visual-language action model to be trained is thus trained using the first data subset to form an initial visual-language action model with a basic mapping relationship of multimodal information and preliminary decision logic based on expert experience.

[0035] In addition, before obtaining the first data subset, the method also includes a step of responding to a construction instruction and constructing a visual language action model to be trained according to the construction instruction; wherein the visual language action model to be trained specifically includes a visual language model to be trained (VLM) and a lightweight action head; the VLM is responsible for processing surgical video frames and natural language instructions to output hidden representations; the lightweight action head is used to map the hidden representations into specific surgical action labels.

[0036] In one embodiment, inputting the first data subset into a visual-language-action model to be trained for model training to obtain an initial visual-language-action model includes:

[0037] Inputting the surgical video frames and the natural language instructions in the first data subset into the video language model to be trained in the visual language action model to be trained for fusion processing to obtain a hidden representation;

[0038] Inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation;

[0039] Calculating the first surgical action label and the second surgical action label to obtain a mean square error;

[0040] The visual language action model to be trained is iteratively optimized using a preset iterative algorithm and the mean square error until the mean square error in the visual language action model to be trained meets a preset training stop condition, then the iteration is stopped and the initialized visual language action model is obtained.

[0041] In this embodiment, the surgical video frames and corresponding natural language instructions in the first data subset are simultaneously input into the video language model to be trained in the visual language action model to be trained. The video language model first extracts features from the surgical video frames and converts them into high-dimensional feature vectors using a convolutional neural network (CNN). At the same time, the natural language instructions are word-embedded, each word is converted into a vector representation, and then language features are extracted through a recurrent neural network (RNN) or Transformer encoder. The video language model then uses an attention mechanism to fuse image features and language features to obtain a hidden representation; wherein the hidden representation contains visual information of the surgical scene and information about the doctor's operating intentions. The hidden representation is input into the lightweight action head in the visual language action model to be trained. The lightweight action head further performs nonlinear transformation and mapping on the hidden representation and outputs a second surgical action label corresponding to the hidden representation; wherein the second surgical action label is a quantitative representation of the surgical action predicted by the model, including the selection of surgical instruments, specific actions and parameters of the operation, etc. Then, a first surgical action label corresponding to the currently input surgical video frame and natural language instruction is obtained from the first data subset, and the first surgical action label is compared with the second surgical action label output by the model to calculate the mean squared error (MSE) between them. The visual language action model to be trained is iteratively optimized using a preset iterative algorithm and the mean squared error; wherein the preset iterative algorithm can be a gradient descent method; in each iteration, the VLM parameters and action head parameters in the visual language action model to be trained are optimized according to the mean squared error, so that the prediction result of the visual language action model to be trained gradually approaches the true first surgical action label. In addition, during the iterative process, the mean squared error of the visual language action model to be trained is continuously calculated. When the mean squared error meets the preset training stop condition, the iteration is stopped and the initialized visual language action model is obtained; wherein the preset training stop condition can be a preset mean squared error threshold, such as 0.01.

[0042] In one embodiment, inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation includes:

[0043] The lightweight action head includes a learnable module and a multi-layer perceptron module;

[0044] Inputting the hidden representation into a learnable module in the lightweight action head for processing to obtain an adaptive label corresponding to the hidden representation;

[0045] The adaptive label is input into the multi-layer perceptron module in the lightweight action head for processing to obtain the second surgical action label.

[0046] In this embodiment, in order to map the hidden representation to a specific surgical action label, the lightweight action head in the visual language action model to be trained includes a learnable (Token Learner) module and a multi-layer perceptron (MLP) module. Specifically, the hidden representation obtained by fusing the surgical video frame and the natural language instruction of the video language model to be trained is input into the learnable module. The learnable module extracts and compresses the hidden representation to obtain a lower-dimensional adaptive tag, for example, reducing the hidden representation from a dimension of 256 to an adaptive tag of a dimension of 128; wherein the adaptive tag is used to highlight the key information related to the surgical action. Afterwards, the adaptive tag is input into the multi-layer perceptron module for preliminary feature extraction to obtain a first feature vector to capture the more complex relationship in the adaptive tag; then the first feature vector is further abstracted and reduced in dimension to remove some redundant information to obtain a second feature vector, and finally the second feature vector is mapped to a second surgical action label; wherein the second surgical action label contains information such as the type of surgical instrument (e.g., a specific model of suture needle, needle holder, etc.), the operation action (e.g., instrument movement, cutting, suturing, etc.), the operation force and speed, etc. For example, the output second surgical action label may be represented as "pick up a small lancet from a designated area of ​​the operating table at a speed of 10 centimeters per second."

[0047] In one embodiment, obtaining the first data subset includes:

[0048] A preset number of expert data sets are randomly selected from a preset database according to a preset selection strategy as the first data subset.

[0049] In this embodiment, to ensure data diversity, the first data subset is obtained by randomly selecting a predetermined number of data from a pre-stored number of expert datasets using a preset selection strategy. For example, if a preset database contains 2000 expert datasets, a 1000-data subset is randomly selected according to the preset selection strategy during the supervised learning initialization phase as the first data subset for training the visual-language-action model to be trained.

[0050] S120 , obtaining first new surgical task data, inputting the first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform a surgical action according to an action label output by the initial visual language action model.

[0051] In this embodiment, after obtaining the initial visual language action model, the initial visual language action model has the ability to understand basic surgical scenes and natural language instructions; at this time, the first new surgical task data can be obtained specifically by using a camera to continuously capture the surgical scene during the operation, such as the operating table, various surgical tools, etc.; at the same time, the natural language instructions issued by the doctor are captured through a microphone, such as "Please turn on the surgical light", "Please hand the scalpel to the doctor", "Please take the surgical forceps from the doctor", etc., so that the surgical scene and natural language instructions form the first new surgical task data; and the first new surgical task data is input into the initial visual language action model for processing to output an action label corresponding to the first new surgical task data and send it to the surgical robot; after receiving the action label output by the model, the surgical robot parses and converts it, maps the surgical instrument type in the action label to the instrument actually equipped by the surgical robot, and converts the operation action, force and speed and other information into instructions that can be understood by the surgical robot's mechanical arm and end effector, thereby enabling the surgical robot to perform the surgical action corresponding to the action label. For example, if the action label is "pick up a small lancet from the designated area of ​​the operating table at a speed of 10 centimeters per second", the surgical robot will convert "a speed of 10 centimeters per second" into the motion parameters of the robotic arm.

[0052] S130: Optimize the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model.

[0053] In this embodiment, since the initial visual language action model has learned rich visual and language feature representations in the initial training stage, in order to ensure that the initial visual language action model remains stable in the reinforcement learning stage and can adapt to different surgical task requirements, this embodiment freezes the visual language model (VLM) parameters in the initial visual language action model and uses a preset reinforcement learning algorithm to optimize only the action head parameters in the initial visual language action model, thereby obtaining an optimized first visual language action model; wherein, the preset reinforcement learning algorithm is preferably PPO (Proximal Policy Optimization, proximal policy optimization) algorithm; in specific implementation, the surgical robot runs in the first new surgical task data for a period of time and collects a series of task trajectory data; wherein the task trajectory data includes a series of states (visual images, natural language instructions, etc.), actions (action head output), rewards (feedback based on the action effect) and the next state; the task trajectory data is then stored in a preset buffer for subsequent optimization of the initial visual language action model; and the generalized advantage estimation (GAE) is used to calculate the advantage function of each trajectory data in the task trajectory data respectively to measure the quality of the action; then a loss function is constructed to optimize the policy network. After repeated sampling and optimization steps of the task trajectory data, until the loss function calculated by the initial visual language action model is less than a preset threshold, the initial visual language action model is deemed to have converged, the training is terminated and the first visual language action model is obtained.

[0054] S140. Storing the task trajectory data of the surgical robot successfully performing the surgical action into an online data set.

[0055] In this embodiment, in order to enable the initial visual language action model to continuously learn the latest surgical tasks and experiences, this embodiment can specifically judge the task trajectory data generated by the surgical robot performing surgical actions according to preset judgment criteria to determine whether it is successfully completed task trajectory data; if it is successfully completed task trajectory data, it is stored in the online data set, thereby automatically generating high-quality training data, reducing data acquisition costs, and providing training samples for further optimization of the subsequent initial visual language action model.

[0056] For example, when the surgical robot executes the instruction "take a small lancet from the operating table", the data of each frame of the surgical robot (including the various states of the robotic arm joints and the pictures taken by the camera device) is recorded during the execution process to form trajectory data; then the preset judgment criteria are used to judge whether the surgical robot has completed the instruction. If completed, the trajectory data is regarded as the trajectory data of the successfully completed task.

[0057] S150 , obtaining a second data subset, and inputting the second data subset and the task trajectory data in the online data set into the first visual language action model for model training to obtain a second visual language action model.

[0058] In this embodiment, the method of obtaining the second data subset in this embodiment is similar to the method of obtaining the first data subset described above, and both methods randomly select a preset number of data from a preset expert data set as the second data subset, which will not be described in detail here. In addition, in order to obtain a fully optimized visual language action model, this embodiment inputs the second data subset and the successfully completed task trajectory data in the online data set into the first visual language action model for model training, until the loss function in the first visual language action model meets the preset training stop condition, then stops the model training and obtains the second visual language action model, so that the first visual language action model is exposed to more diverse surgical scenarios and task instances, so as to better adapt to new tasks and maintain performance on old tasks, thereby enhancing generalization ability. Among them, the loss function in the first visual language action model is preferably obtained by calculating the mean square error in this embodiment.

[0059] In one embodiment, inputting the second data subset and the task trajectory data in the online dataset into the first visual language action model for model training to obtain the second visual language action model includes:

[0060] performing data preprocessing on the second data subset and the task trajectory data in the online data set to obtain a preprocessed second data subset and preprocessed task trajectory data;

[0061] Merging the preprocessed second data subset and the preprocessed task trajectory data according to a preset merging strategy to obtain merged data;

[0062] The combined data is input into the first visual language action model for model training to obtain the second visual language action model.

[0063] In this embodiment, the successfully completed task trajectory data in the second data subset and the online dataset are first preprocessed to remove noise and outliers, unify the data format, and process missing values, thereby improving the data quality. Then, to help the first visual language action model learn more comprehensive knowledge and skills, this embodiment also merges the preprocessed second data subset and the preprocessed task trajectory data to increase the diversity of the training data. The expert dataset may contain the operation data of experienced doctors or experts in specific surgical scenarios, while the online dataset contains various situations in the actual surgical task environment. Therefore, the merged data covers a wider range of surgical scenarios and operation methods. Therefore, the merged data is input into the first visual language action model for model training, thereby obtaining a fully optimized second visual language action model.

[0064] In one embodiment, after step S150, the method further includes:

[0065] S160: Fine-tune the parameters of the visual language model in the second visual language action model using a preset fine-tuning algorithm to obtain a fine-tuned visual language action model.

[0066] In this embodiment, after obtaining a fully optimized second visual language action model, although the overall model performance is good, in order to further improve the accuracy of surgical action generation, this embodiment also fine-tunes the parameters of the visual language model in the second visual language action model using a preset fine-tuning algorithm. Specifically, the LORA (Low-Rank Adaptation) method can be used to fine-tune some parameters of the visual language model (VLM) in the second visual language action model to obtain a fine-tuned visual language action model, thereby reducing the computational burden while optimizing the model performance of the second visual language action model.

[0067] In one embodiment, after step S160, the method further includes:

[0068] S170. Acquire the second new surgical task data, update the second new surgical task data into the first data subset, update the fine-tuning visual language action model into the visual language action model to be trained, and return to execute the step of inputting the first data subset into the visual language action model to be trained for model training to obtain the initial visual language action model.

[0069] In this embodiment, in order to enable the visual language action model to adapt to more different surgical task scenarios, this embodiment forms a feedback loop by continuously repeating the training steps to continuously adjust and optimize the parameters and performance of the model itself, and gradually improve the generalization ability of the visual language action (VLA) model. Specifically, first, by obtaining the second new surgical task data; wherein, the method of obtaining the second new surgical task data is similar to the method of obtaining the first new surgical task data mentioned above, and will not be repeated here. Then, the second new surgical task data is updated to the first data subset, the fine-tuned visual language action model is updated to the visual language action model to be trained, and the step of inputting the first data subset into the visual language action model to be trained for model training to obtain the initial visual language action model is returned to execute, and the execution is continuously repeated to improve the visual language action model to adapt to more complex surgical task scenarios, so that the auxiliary surgical robot can adjust the surgical strategy according to real-time feedback, thereby improving the targeting and success rate of complex operations.

[0070] It can be seen from the above technical solution that the present invention obtains an initial visual language action model by inputting a first data subset into the visual language action model to be trained; inputting the first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform surgical actions according to the output of the initial visual language action model; optimizing the action head parameters of the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model; storing the task trajectory data successfully completed by the surgical robot into an online data set; inputting the second data subset and the task trajectory data into the first visual language action model for training to obtain a second visual language action model, thereby continuously reinforcing the learning of the VLA model, so that the surgical robot can adjust the surgical strategy according to real-time feedback, thereby improving surgical accuracy.

[0071] Figure 3 Schematic block diagram of a VLA model generation device for a surgical robot provided by an embodiment of the present invention. Figure 3 As shown, the present invention also provides a VLA model generation device for a surgical robot, the VLA model generation device for the surgical robot includes a unit for executing the VLA model generation method of the surgical robot. Specifically, see Figure 3 The VLA model generation device 100 of the surgical robot includes a first training unit 110, an execution unit 120, a reinforcement learning unit 130, a storage unit 140, a second training unit 150, a fine-tuning unit 160 and an updating unit 170.

[0072] The first training unit 110 is configured to obtain a first data subset, and input the first data subset into a visual language action model to be trained for model training to obtain an initial visual language action model.

[0073] In this embodiment, to enable the initial visual-language action model to initially understand the first data subset and perform the surgical task in the first data subset, this embodiment specifically inputs the first data subset into a pre-constructed visual-language action model to be trained for model training to obtain an initial visual-language action model. The first data subset is obtained from a preset expert dataset and includes surgical video frames, natural language instructions, and first surgical action labels corresponding to the surgical video frames and natural language instructions. The surgical video frames refer to visual input during the surgical procedure, such as the operation of surgical instruments and changes in the surgical site. The natural language instructions refer to voice instructions that an expert may issue during the procedure, such as "cut the skin." The first surgical action labels are precise action annotations for each surgical video frame and corresponding natural language instruction combination based on standard surgical operating procedures and expert experience. These labels may include rotation data for each control joint of the surgical robot, gripper grip / release status, and other action data during the procedure. The visual-language action model to be trained is thus trained using the first data subset to form an initial visual-language action model with a basic mapping relationship of multimodal information and preliminary decision logic based on expert experience.

[0074] In addition, before obtaining the first data subset, the method also includes a step of responding to a construction instruction and constructing a visual language action model to be trained according to the construction instruction; wherein the visual language action model to be trained specifically includes a visual language model to be trained (VLM) and a lightweight action head; the VLM is responsible for processing surgical video frames and natural language instructions to output hidden representations; the lightweight action head is used to map the hidden representations into specific surgical action labels.

[0075] In one embodiment, inputting the first data subset into a visual-language-action model to be trained for model training to obtain an initial visual-language-action model includes:

[0076] Inputting the surgical video frames and the natural language instructions in the first data subset into the video language model to be trained in the visual language action model to be trained for fusion processing to obtain a hidden representation;

[0077] Inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation;

[0078] Calculating the first surgical action label and the second surgical action label to obtain a mean square error;

[0079] The visual language action model to be trained is iteratively optimized using a preset iterative algorithm and the mean square error until the mean square error in the visual language action model to be trained meets a preset training stop condition, then the iteration is stopped and the initialized visual language action model is obtained.

[0080] In this embodiment, the surgical video frames and corresponding natural language instructions in the first data subset are simultaneously input into the video language model to be trained in the visual language action model to be trained. The video language model first extracts features from the surgical video frames and converts them into high-dimensional feature vectors using a convolutional neural network (CNN). At the same time, the natural language instructions are word-embedded, each word is converted into a vector representation, and then language features are extracted through a recurrent neural network (RNN) or Transformer encoder. The video language model then uses an attention mechanism to fuse image features and language features to obtain a hidden representation; wherein the hidden representation contains visual information of the surgical scene and information about the doctor's operating intentions. The hidden representation is input into the lightweight action head in the visual language action model to be trained. The lightweight action head further performs nonlinear transformation and mapping on the hidden representation and outputs a second surgical action label corresponding to the hidden representation; wherein the second surgical action label is a quantitative representation of the surgical action predicted by the model, including the selection of surgical instruments, specific actions and parameters of the operation, etc. Then, a first surgical action label corresponding to the currently input surgical video frame and natural language instruction is obtained from the first data subset, and the first surgical action label is compared with the second surgical action label output by the model to calculate the mean squared error (MSE) between them. The visual language action model to be trained is iteratively optimized using a preset iterative algorithm and the mean squared error; wherein the preset iterative algorithm can be a gradient descent method; in each iteration, the VLM parameters and action head parameters in the visual language action model to be trained are optimized according to the mean squared error, so that the prediction result of the visual language action model to be trained gradually approaches the true first surgical action label. In addition, during the iterative process, the mean squared error of the visual language action model to be trained is continuously calculated. When the mean squared error meets the preset training stop condition, the iteration is stopped and the initialized visual language action model is obtained; wherein the preset training stop condition can be a preset mean squared error threshold, such as 0.01.

[0081] In one embodiment, inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation includes:

[0082] The lightweight action head includes a learnable module and a multi-layer perceptron module;

[0083] Inputting the hidden representation into a learnable module in the lightweight action head for processing to obtain an adaptive label corresponding to the hidden representation;

[0084] The adaptive label is input into the multi-layer perceptron module in the lightweight action head for processing to obtain the second surgical action label.

[0085] In this embodiment, in order to map the hidden representation to a specific surgical action label, the lightweight action head in the visual language action model to be trained includes a learnable (Token Learner) module and a multi-layer perceptron (MLP) module. Specifically, the hidden representation obtained by fusing the surgical video frame and the natural language instruction of the video language model to be trained is input into the learnable module. The learnable module extracts and compresses the hidden representation to obtain a lower-dimensional adaptive tag, for example, reducing the hidden representation from a dimension of 256 to an adaptive tag of a dimension of 128; wherein the adaptive tag is used to highlight the key information related to the surgical action. Afterwards, the adaptive tag is input into the multi-layer perceptron module for preliminary feature extraction to obtain a first feature vector to capture the more complex relationship in the adaptive tag; then the first feature vector is further abstracted and reduced in dimension to remove some redundant information to obtain a second feature vector, and finally the second feature vector is mapped to a second surgical action label; wherein the second surgical action label contains information such as the type of surgical instrument (e.g., a specific model of suture needle, needle holder, etc.), the operation action (e.g., instrument movement, cutting, suturing, etc.), the operation force and speed, etc. For example, the output second surgical action label may be represented as "pick up a small lancet from a designated area of ​​the operating table at a speed of 10 centimeters per second."

[0086] In one embodiment, obtaining the first data subset includes:

[0087] A preset number of expert data sets are randomly selected from a preset database according to a preset selection strategy as the first data subset.

[0088] In this embodiment, to ensure data diversity, the first data subset is obtained by randomly selecting a predetermined number of data from a pre-stored number of expert datasets using a preset selection strategy. For example, if a preset database contains 2000 expert datasets, a 1000-data subset is randomly selected according to the preset selection strategy during the supervised learning initialization phase as the first data subset for training the visual-language-action model to be trained.

[0089] The execution unit 120 is used to obtain first new surgical task data, input the first new surgical task data into the initial visual language action model, and control the surgical robot to perform the surgical action according to the action label output by the initial visual language action model.

[0090] In this embodiment, after obtaining the initial visual language action model, the initial visual language action model has the ability to understand basic surgical scenes and natural language instructions; at this time, the first new surgical task data can be obtained specifically by using a camera to continuously capture the surgical scene during the operation, such as the operating table, various surgical tools, etc.; at the same time, the natural language instructions issued by the doctor are captured through a microphone, such as "Please turn on the surgical light", "Please hand the scalpel to the doctor", "Please take the surgical forceps from the doctor", etc., so that the surgical scene and natural language instructions form the first new surgical task data; and the first new surgical task data is input into the initial visual language action model for processing to output an action label corresponding to the first new surgical task data and send it to the surgical robot; after receiving the action label output by the model, the surgical robot parses and converts it, maps the surgical instrument type in the action label to the instrument actually equipped by the surgical robot, and converts the operation action, force and speed and other information into instructions that can be understood by the surgical robot's mechanical arm and end effector, thereby enabling the surgical robot to perform the surgical action corresponding to the action label. For example, if the action label is "pick up a small lancet from the designated area of ​​the operating table at a speed of 10 centimeters per second", the surgical robot will convert "a speed of 10 centimeters per second" into the motion parameters of the robotic arm.

[0091] The reinforcement learning unit 130 is configured to optimize the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model.

[0092] In this embodiment, since the initial visual language action model has learned rich visual and language feature representations in the initial training stage, in order to ensure that the initial visual language action model remains stable in the reinforcement learning stage and can adapt to different surgical task requirements, this embodiment freezes the visual language model (VLM) parameters in the initial visual language action model and uses a preset reinforcement learning algorithm to optimize only the action head parameters in the initial visual language action model, thereby obtaining an optimized first visual language action model; wherein, the preset reinforcement learning algorithm is preferably PPO (Proximal Policy Optimization, proximal policy optimization) algorithm; in specific implementation, the surgical robot runs in the first new surgical task data for a period of time and collects a series of task trajectory data; wherein the task trajectory data includes a series of states (visual images, natural language instructions, etc.), actions (action head output), rewards (feedback based on the action effect) and the next state; the task trajectory data is then stored in a preset buffer for subsequent optimization of the initial visual language action model; and the generalized advantage estimation (GAE) is used to calculate the advantage function of each trajectory data in the task trajectory data respectively to measure the quality of the action; then a loss function is constructed to optimize the policy network. After repeated sampling and optimization steps of the task trajectory data, until the loss function calculated by the initial visual language action model is less than a preset threshold, the initial visual language action model is deemed to have converged, the training is terminated and the first visual language action model is obtained.

[0093] The storage unit 140 is used to store the task trajectory data of the surgical robot successfully performing the surgical action into the online data set.

[0094] In this embodiment, in order to enable the initial visual language action model to continuously learn the latest surgical tasks and experiences, this embodiment can specifically judge the task trajectory data generated by the surgical robot performing surgical actions according to preset judgment criteria to determine whether it is successfully completed task trajectory data; if it is successfully completed task trajectory data, it is stored in the online data set, thereby automatically generating high-quality training data, reducing data acquisition costs, and providing training samples for further optimization of the subsequent initial visual language action model.

[0095] For example, when the surgical robot executes the instruction "take a small lancet from the operating table", the data of each frame of the surgical robot (including the various states of the robotic arm joints and the pictures taken by the camera device) is recorded during the execution process to form trajectory data; then the preset judgment criteria are used to judge whether the surgical robot has completed the instruction. If completed, the trajectory data is regarded as the trajectory data of the successfully completed task.

[0096] The second training unit 150 is configured to obtain a second data subset, input the second data subset and the task trajectory data in the online data set into the first visual language action model for model training, and obtain a second visual language action model.

[0097] In this embodiment, the method of obtaining the second data subset in this embodiment is similar to the method of obtaining the first data subset described above, and both methods randomly select a preset number of data from a preset expert data set as the second data subset, which will not be described in detail here. In addition, in order to obtain a fully optimized visual language action model, this embodiment inputs the second data subset and the successfully completed task trajectory data in the online data set into the first visual language action model for model training, until the loss function in the first visual language action model meets the preset training stop condition, then stops the model training and obtains the second visual language action model, so that the first visual language action model is exposed to more diverse surgical scenarios and task instances, so as to better adapt to new tasks and maintain performance on old tasks, thereby enhancing generalization ability. Among them, the loss function in the first visual language action model is preferably obtained by calculating the mean square error in this embodiment.

[0098] In one embodiment, inputting the second data subset and the task trajectory data in the online dataset into the first visual language action model for model training to obtain the second visual language action model includes:

[0099] performing data preprocessing on the second data subset and the task trajectory data in the online data set to obtain a preprocessed second data subset and preprocessed task trajectory data;

[0100] Merging the preprocessed second data subset and the preprocessed task trajectory data according to a preset merging strategy to obtain merged data;

[0101] The combined data is input into the first visual language action model for model training to obtain the second visual language action model.

[0102] In this embodiment, the successfully completed task trajectory data in the second data subset and the online dataset are first preprocessed to remove noise and outliers, unify the data format, and process missing values, thereby improving the data quality. Then, to help the first visual language action model learn more comprehensive knowledge and skills, this embodiment also merges the preprocessed second data subset and the preprocessed task trajectory data to increase the diversity of the training data. The expert dataset may contain the operation data of experienced doctors or experts in specific surgical scenarios, while the online dataset contains various situations in the actual surgical task environment. Therefore, the merged data covers a wider range of surgical scenarios and operation methods. Therefore, the merged data is input into the first visual language action model for model training, thereby obtaining a fully optimized second visual language action model.

[0103] In one embodiment, after the second training unit 150, the following further steps are included:

[0104] The fine-tuning unit 160 is configured to fine-tune the parameters of the visual language model in the second visual language action model using a preset fine-tuning algorithm to obtain a fine-tuned visual language action model.

[0105] In this embodiment, after obtaining a fully optimized second visual language action model, although the overall model performance is good, in order to further improve the accuracy of surgical action generation, this embodiment also fine-tunes the parameters of the visual language model in the second visual language action model using a preset fine-tuning algorithm. Specifically, the LORA (Low-Rank Adaptation) method can be used to fine-tune some parameters of the visual language model (VLM) in the second visual language action model to obtain a fine-tuned visual language action model, thereby reducing the computational burden while optimizing the model performance of the second visual language action model.

[0106] In one embodiment, after the fine-tuning unit 160, the following further comprises:

[0107] The updating unit 170 is used to obtain the second new surgical task data, update the second new surgical task data into the first data subset, update the fine-tuning visual language action model into the visual language action model to be trained, and return to execute the step of inputting the first data subset into the visual language action model to be trained for model training to obtain the initial visual language action model.

[0108] In this embodiment, in order to enable the visual language action model to adapt to more different surgical task scenarios, this embodiment forms a feedback loop by continuously repeating the training steps to continuously adjust and optimize the parameters and performance of the model itself, and gradually improve the generalization ability of the visual language action (VLA) model. Specifically, first, by obtaining the second new surgical task data; wherein, the method of obtaining the second new surgical task data is similar to the method of obtaining the first new surgical task data mentioned above, and will not be repeated here. Then, the second new surgical task data is updated to the first data subset, the fine-tuned visual language action model is updated to the visual language action model to be trained, and the step of inputting the first data subset into the visual language action model to be trained for model training to obtain the initial visual language action model is returned to execute, and the execution is continuously repeated to improve the visual language action model to adapt to more complex surgical task scenarios, so that the auxiliary surgical robot can adjust the surgical strategy according to real-time feedback, thereby improving the targeting and success rate of complex operations.

[0109] It can be seen from the above technical solution that the present invention obtains an initial visual language action model by inputting a first data subset into the visual language action model to be trained; inputting the first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform surgical actions according to the output of the initial visual language action model; optimizing the action head parameters of the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model; storing the task trajectory data successfully completed by the surgical robot into an online data set; inputting the second data subset and the task trajectory data into the first visual language action model for training to obtain a second visual language action model, thereby continuously reinforcing the learning of the VLA model, so that the surgical robot can adjust the surgical strategy according to real-time feedback, thereby improving surgical accuracy.

[0110] The VLA model generation device of the surgical robot can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer equipment shown.

[0111] See also Figure 4 , Figure 4 is a schematic block diagram of a computer device provided in an embodiment of the present invention. Computer device 500 is a server, which can be a standalone server or a server cluster consisting of multiple servers. Computer device 500 can also be a transmitter, which can be a communication-capable electronic device such as a smartphone, tablet computer, laptop computer, desktop computer, personal digital assistant, or wearable device.

[0112] See Figure 4The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a storage medium 503 and an internal memory 504 .

[0113] The storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute a VLA model generation method for a surgical robot.

[0114] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0115] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute the VLA model generation method of the surgical robot.

[0116] The network interface 505 is used for network communication, such as providing data information transmission. Those skilled in the art will understand that Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0117] The processor 502 is configured to run a computer program 5032 stored in a memory to implement the VLA model generation method for a surgical robot disclosed in an embodiment of the present invention.

[0118] Those skilled in the art will understand that Figure 4 The embodiment of the computer device shown in the figure does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structure and function of the memory and processor are the same as those in the figure. Figure 4 The embodiments shown are consistent and will not be described again here.

[0119] It should be understood that in the embodiment of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0120] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the method for generating a VLA model for a surgical robot disclosed in an embodiment of the present invention is implemented.

[0121] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the devices, systems and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0122] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, systems or units, or may be an electrical, mechanical or other form of connection.

[0123] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.

[0124] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0126] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A VLA model generation method for a surgical robot, characterized in that: The method comprises: Obtaining a first data subset, and inputting the first data subset into a visual language action model to be trained for model training to obtain an initial visual language action model; wherein the first data subset includes surgical video frames, natural language instructions, and a first surgical action label; Acquiring first new surgical task data, inputting the first new surgical task data into the initial visual language action model, and controlling the surgical robot to perform a surgical action according to an action label output by the initial visual language action model; Optimizing the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model; storing the task trajectory data of the surgical robot successfully performing the surgical action into an online dataset; A second data subset is obtained, and the second data subset and the task trajectory data in the online data set are input into the first visual language action model for model training to obtain a second visual language action model.

2. The VLA model generation method for a surgical robot according to claim 1, characterized in that: After obtaining the second visual language action model, the method further includes: The parameters of the visual language model in the second visual language action model are fine-tuned using a preset fine-tuning algorithm to obtain a fine-tuned visual language action model.

3. The VLA model generation method for a surgical robot according to claim 2, characterized in that: After obtaining the fine-tuned vision-language-action model, the method further includes: Acquire the second new surgical task data, update the second new surgical task data into the first data subset, update the fine-tuned visual language action model into the visual language action model to be trained, and return to execute the step of inputting the first data subset into the visual language action model to be trained for model training to obtain the initial visual language action model.

4. The VLA model generation method for a surgical robot according to claim 1, characterized in that: The step of inputting the first data subset into the visual language action model to be trained to perform model training to obtain an initial visual language action model includes: Inputting the surgical video frames and the natural language instructions in the first data subset into the video language model to be trained in the visual language action model to be trained for fusion processing to obtain a hidden representation; Inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation; Calculating the first surgical action label and the second surgical action label to obtain a mean square error; The visual language action model to be trained is iteratively optimized using a preset iterative algorithm and the mean square error until the mean square error in the visual language action model to be trained meets a preset training stop condition, then the iteration is stopped and the initialized visual language action model is obtained.

5. The VLA model generation method for a surgical robot according to claim 4, characterized in that: Inputting the hidden representation into the lightweight action head in the visual language action model to be trained for processing to obtain a second surgical action label corresponding to the hidden representation includes: The lightweight action head includes a learnable module and a multi-layer perceptron module; Inputting the hidden representation into a learnable module in the lightweight action head for processing to obtain an adaptive label corresponding to the hidden representation; The adaptive label is input into the multi-layer perceptron module in the lightweight action head for processing to obtain the second surgical action label.

6. The VLA model generation method for a surgical robot according to claim 1, characterized in that: The obtaining of the first data subset includes: A preset number of expert data sets are randomly selected from a preset database according to a preset selection strategy as the first data subset.

7. The VLA model generation method for a surgical robot according to claim 1, characterized in that: The step of inputting the second data subset and the task trajectory data in the online data set into the first visual language action model for model training to obtain a second visual language action model includes: performing data preprocessing on the second data subset and the task trajectory data in the online data set to obtain a preprocessed second data subset and preprocessed task trajectory data; Merging the preprocessed second data subset and the preprocessed task trajectory data according to a preset merging strategy to obtain merged data; The combined data is input into the first visual language action model for model training to obtain the second visual language action model.

8. A VLA model generation device for a surgical robot, characterized in that: The device comprises: A first training unit is configured to obtain a first data subset and input the first data subset into a visual language action model to be trained for model training to obtain an initial visual language action model; wherein the first data subset includes surgical video frames, natural language instructions, and a first surgical action label; an execution unit, configured to obtain first new surgical task data, input the first new surgical task data into the initial visual language action model, and control the surgical robot to perform a surgical action according to an action label output by the initial visual language action model; a reinforcement learning unit, configured to optimize the action head parameters in the initial visual language action model using a preset reinforcement learning algorithm to obtain a first visual language action model; A storage unit, configured to store the task trajectory data of the surgical robot successfully performing the surgical action into an online data set; The second training unit is used to obtain a second data subset, input the second data subset and the task trajectory data in the online data set into the first visual language action model for model training, and obtain a second visual language action model.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the VLA model generation method of the surgical robot according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the VLA model generation method for the surgical robot according to any one of claims 1 to 7 can be implemented.

Citation Information

Cited By

  • Training method of humanoid robot task planning model and task planning method

    CN121340271A

  • Robot control method, computing device and readable storage medium

    CN121403420A

  • Robust vision-language-action model post-training method based on reinforcement learning

    CN122200228A