Ultrasonic scanning robot control method and system based on enhanced large language model
Through the ultrasonic scanning robot control method based on the enhanced large language model, the direct interaction between the robot and the patient is achieved, the problem of inability to understand the patient's intentions in the prior art is solved, the autonomy and perception ability of the robot are enhanced, and it is suitable for clinical applications.
Patent Information
- Application Number
- CN202510411312.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing ultrasound scanning robot lacks the ability to interact directly with patients, is unable to effectively understand the patient's intentions and adjust the scanning strategy in real time, which limits its autonomy and its promotion in actual clinical applications.
The ultrasonic scanning robot control method based on the enhanced large language model is used to receive patient instruction information, use prompt word engineering to encapsulate instructions, call the large language model for chain reasoning and imitation learning, combine multimodal feature generation scanning strategy, dynamically adjust contact force and position pose, and realize autonomous ultrasonic scanning.
It realizes flexible human-computer interaction between ultrasonic scanning robots and patients, can accurately analyze patient needs, adjust scanning strategies in real time, enhances the robot's autonomy and perception and adaptability in complex environments, and is easy to promote in clinical applications.
Smart Images

Figure CN120280112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasonic scanning robots, and in particular to an ultrasonic scanning robot control method and system based on an enhanced large language model. Background Art
[0002] Ultrasound scanning robots are considered to be an effective means to solve the problems of traditional ultrasound scanning, which is over-dependent on the operator's experience and skills, susceptible to human factors, resulting in unstable imaging quality and diagnostic bias, and a shortage of high-level ultrasound physicians in remote areas.
[0003] However, existing ultrasound scanning robots still have limitations in terms of operational flexibility and intelligence. Although they can achieve a certain degree of automated scanning, these robots lack the ability to interact directly with patients and cannot effectively understand patients' intentions and adjust scanning strategies in real time, including online adjustment of parameters such as scanning force, trajectory, and scanning area. These limitations not only restrict the autonomy of ultrasound scanning robots, but also limit their promotion in actual clinical applications. Summary of the invention
[0004] The purpose of the present invention is to provide an ultrasonic scanning robot control method and system based on an enhanced large language model, which is implemented by calling the ChatGPT API or other large language model API interfaces, mapping high-dimensional, complex language and environmental information into specific scanning subtasks, accurately analyzing patient needs and clinical scenarios, effectively understanding the patient's intentions and adjusting the scanning strategy in real time, aiming to solve the problems in the prior art.
[0005] The present invention is implemented in this way: an ultrasonic scanning robot control method based on an enhanced large language model is applied to an ultrasonic scanning robot, specifically comprising the following steps:
[0006] S101: receiving real-time instruction information of a patient, converting and cleaning the received real-time instruction information, and encapsulating the patient's instruction by using a prompt word project;
[0007] S102: By calling the API of the large language model, the prompt words that encapsulate the patient's instructions are passed to the model. Based on the chain reasoning technology, the model performs multiple and repeated large model calls according to the prompt word requirements to complete the iterative decomposition of the prompt word;
[0008] S103: Under the deep learning framework, receiving the imitation learning structure instruction generated by the large language model, and defining the specific structure of the imitation learning model according to the structure instruction, and inferring the next moment position instruction of the probe by the imitation learning neural network by combining multimodal features;
[0009] S104: Based on the impedance control strategy, perform contact force control, establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, dynamically correct the contact force, and calculate the target angles of each joint of the robotic arm through inverse kinematics to achieve the movement of the probe from the current state to the target pose.
[0010] S105: Receive the patient's next voice command in real time, and dynamically adjust the progress, contact force, and scanning position of the ultrasonic scan according to the command to make the operating state of the robotic arm consistent with the patient's needs, so as to achieve autonomous ultrasonic scanning with flexible human-machine interaction capabilities.
[0011] Further, in S101, receive the patient's real-time instruction information, convert and clean the received real-time instruction information, and complete the generation of the prompt words, including:
[0012] Receive the patient's real-time voice command. Based on automatic speech recognition technology, convert the patient's voice input into a transcribed text.
[0013] Clean and standardize the transcribed text to eliminate meaningless characters and duplicate characters in the transcribed text. The cleaning includes denoising and grammar adjustment operations.
[0014] Further, use prompt engineering to encapsulate the patient's instructions, including:
[0015] Define the core target features of the prompt words according to the clinical task characteristics and patient needs of the patient, and clearly define the core target features in the prompt words.
[0016] Embed domain-related background knowledge and specific terms in the prompt words to make the prompt words conform to the context learning ability of the model.
[0017] Specify the output format in the prompt words, and clearly require the model to return the results in a standardized form.
[0018] Further, in S102, the model is based on the chain-of-thought reasoning technique and performs multiple and repeated large model calls according to the requirements of the prompt words to complete the iterative decomposition of the prompt words, including:
[0019] Decompose the requirements of the prompt words into multiple low-level subtasks with clear logic and distinct levels through multiple and repeated calls to the large model.
[0020] Use chain-of-thought reasoning to perform reasoning on the low-level subtasks, and in each step of reasoning, combine the current context subtask with the previous result to predict and generate the next step.
[0021] The model uses the context consistency check mechanism of the task to verify the order and logical relationship of the generated subtasks. After initially generating the subtasks, the model further optimizes the order of the subtasks.
[0022] Further, in S103, under the deep learning framework, receive the imitation learning structure instructions generated by the large language model, define the specific structure of the imitation learning model according to the structure instructions, and through combining multi-modal features, the imitation learning neural network infers the pose instruction of the probe at the next moment, including:
[0023] Use the depth camera to collect environmental data in real time, and obtain the RGB image and the corresponding point cloud data;
[0024] Use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, extract the features of RGB and point cloud, and obtain Tokens;
[0025] Input the extracted RGB and point cloud Tokens into the imitation learning neural network, and infer the pose instruction of the ultrasonic probe at the next moment by combining multi-modal features.
[0026] Further, before using ResNet and PointNet as the backbone networks for processing RGB and point cloud data, it is necessary to preprocess the point cloud data, clean and normalize the data, including image normalization, noise filtering, and point cloud downsampling, to improve the quality and consistency of the data, provide more accurate input for feature extraction, and reflect the spatial environment and target area characteristics around the ultrasonic probe.
[0027] Further, use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, extract the features of RGB and point cloud, and obtain Tokens, including:
[0028] Set the network structures of ResNet and PointNet respectively as the backbone networks for processing RGB images and point cloud data, and adapt to the input data;
[0029] Load the pre-trained parameters of ResNet and PointNet;
[0030] Extract the high-dimensional semantic features of the RGB image through ResNet, and extract the spatial structure features of the point cloud data through PointNet, and convert them into semantic Tokens for the imitation learning network to use.
[0031] Further, in S104, perform contact force control based on the impedance control strategy, establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, and dynamically correct the contact force, including:
[0032] Adopt the impedance control equation: Establish the dynamic relationship between the contact force and the probe pose, where K p , K v , K fSet according to prior experience and repeatedly optimized in experiments;
[0033] The contact force of the probe in the normal direction is monitored in real time by a force sensor, and the error between the actual force and the target force is used as feedback to input into the controller. The controller adjusts the displacement of the probe in the normal direction based on this feedback and dynamically corrects the contact force.
[0034] Further, in S104, the target angles of each joint of the robotic arm are calculated through inverse kinematics to enable the probe to move from the current state to the target pose, including:
[0035] Based on the DH parameters, define the link length a, link twist angle α i , joint offset d i and joint angle θ i , and construct the homogeneous transformation matrix A for each joint i ;
[0036] Obtain the homogeneous transformation matrix T from the end effector to the base through successive matrix multiplications T = A1·A2...A n and decouple the target position P and orientation R of the end effector;
[0037] According to the dynamics model, calculate the homogeneous transformation matrix A i , homogeneous transformation matrix T and target position P and orientation R through inverse kinematics, and obtain the target angles of each joint of the robotic arm to achieve precise control of the ultrasonic probe from the current state to the target pose.
[0038] Compared with the prior art, the ultrasonic scanning robot control method and system based on an enhanced large language model provided by the present invention have the following beneficial effects:
[0039] 1. This technical solution is based on scene understanding, task planning, and realizes an autonomous ultrasonic scanning strategy for atomic operations through an imitation learning architecture. Through prompt engineering, the system can input the patient's verbal instructions and environmental information into the large language model to generate low-level tasks. The reasoning ability of the large language model can be achieved by calling the ChatGPT API or other large language model API interfaces, mapping high-dimensional and complex language and environmental information into specific scanning subtasks, accurately parsing the patient's needs and clinical scenarios, effectively understanding the patient's intentions, and adjusting the scanning strategy in real time.
[0040] 2. Based on this architecture, the system can also receive and understand patient feedback in real time during the scanning process, dynamically adjust the scanning strategy and progress, learn from expert operation demonstrations through imitation learning methods, and achieve highly robust execution of subtasks. Combining the reasoning ability of the large language model with the execution ability of imitation learning, the ultrasound robot's perception and adaptability in complex environments are significantly enhanced. In addition, the specific implementation of each link mainly includes: encapsulating patient instructions through prompt word engineering, low-level task generation and decomposition based on the large language model, generating posture instructions based on imitation learning methods, precise instruction execution by the controller, real-time perception and online adjustment, and achieving direct interaction with patients, including online adjustment of parameters such as scanning force, trajectory and scanning position, so that the ultrasound scanning robot has autonomy and is easy to promote in actual clinical applications.
[0041] An ultrasonic scanning robot control system based on an enhanced large language model is used to execute the above-mentioned ultrasonic scanning robot control method, and the control system includes:
[0042] An acquisition module, used for receiving real-time instruction information of patients;
[0043] The encapsulation module is used to convert and clean the received real-time instruction information, generate prompt words, and use prompt word engineering to encapsulate patient instructions;
[0044] The chained reasoning module is used to call the API of the large language model and perform multiple and repeated large model calls according to the prompt word requirements to complete the iterative decomposition of the prompt word;
[0045] The feature reasoning module is used to combine multimodal features and imitate the learning neural network to infer the next moment position and posture instructions of the scanning probe;
[0046] A correction module is used to establish a dynamic relationship between the contact force and the ultrasonic probe posture and to dynamically correct the contact force;
[0047] The ultrasound execution module is used to dynamically adjust the progress, contact force and scanning position of the ultrasound scan according to the instructions, ensure that the operation target posture is consistent with the patient's required posture, and realize autonomous ultrasound scanning with flexible human-computer interaction capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic flow chart of the ultrasonic scanning robot control method based on the enhanced large language model proposed in the present invention;
[0049] Figure 2 This is a schematic diagram of the operation logic of the ultrasonic scanning robot control method based on the enhanced large language model proposed by the present invention;
[0050] Figure 3Schematic diagram of the control system of the ultrasonic scanning robot based on the enhanced large language model proposed by the present invention. Specific embodiments
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] The implementation of the present invention will be described in detail below with reference to specific embodiments.
[0053] In the accompanying drawings of this embodiment, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be understood as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0054] Refer to Figure 1-2 As shown, the ultrasonic scanning robot control method based on the enhanced large language model is applied to an ultrasonic scanning robot and specifically includes the following steps:
[0055] S101: Receive the real-time instruction information of the patient, convert and clean the received real-time instruction information to generate a prompt word, and use prompt engineering to encapsulate the patient's instruction;
[0056] Among them, receiving the real-time instruction information of the patient and converting and cleaning the received real-time instruction information to complete the generation of the prompt word includes:
[0057] Receive the real-time voice instruction of the patient. The real-time voice instruction is based on automatic speech recognition technology to convert the patient's voice input into a transcribed text;
[0058] Clean and standardize the transcribed text to eliminate meaningless characters and duplicate characters in the transcribed text. The cleaning includes denoising and grammar adjustment operations;
[0059] Specifically, cleaning and standardization include operations such as denoising and grammar adjustment. Denoising includes using regular expressions to remove redundant punctuation, special characters, whitespace, etc., and deleting meaningless words such as common transcription fillers (e.g., "um", "this", "you know"). Grammar adjustment uses NLP tools (such as SpaCy or LanguageTool, etc.) to analyze syntactic structures and adjust grammar.
[0060] S102: By calling the API of the large language model (such as ChatGPT API or similar services), the prompt words encapsulating the patient instructions are passed to the model. The model, based on the Chain-of-Thought (CoT) technique, performs multiple and repeated large model calls according to the requirements of the prompt words to complete the iterative decomposition of the prompt words;
[0061] Among them, the model, based on the Chain-of-Thought technique, performs multiple and repeated large model calls according to the requirements of the prompt words to complete the iterative decomposition of the prompt words, including:
[0062] The requirements of the prompt words are disassembled into multiple low-level subtasks with clear logic and distinct levels through multiple and repeated calls to the large model;
[0063] Use Chain-of-Thought for the reasoning of low-level subtasks, and in each step of reasoning, combine the current context subtask with the previous result to predict and generate the next step;
[0064] The model uses the context consistency check mechanism of the task to verify the order and logical relationship of the generated subtasks. After initially generating the subtasks, the model further optimizes the order of the subtasks;
[0065] S103: Under the deep learning framework (such as PyTorch or TensorFlow), receive the imitation learning structure instructions generated by the large language model, define the specific structure of the imitation learning model according to the structure instructions, and through combining multi-modal features, the imitation learning neural network infers the pose instruction of the probe at the next moment;
[0066] Among them, through combining multi-modal features, the imitation learning neural network infers the pose instruction of the probe at the next moment, including:
[0067] Use a depth camera to collect environmental data in real time to obtain RGB images and corresponding point cloud data;
[0068] Use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, extract the features of RGB and point cloud, and obtain Tokens;
[0069] Input the extracted RGB and point cloud tokens into the imitation learning neural network, and infer the pose command of the ultrasonic probe at the next moment by combining multi-modal features;
[0070] S104: Based on the impedance control strategy, perform contact force control, establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, dynamically correct the contact force, and calculate the target angles of the joints of the robotic arm through inverse kinematics to achieve the probe moving from the current state to the target pose;
[0071] Among them, based on the impedance control strategy, perform contact force control, establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, and dynamically correct the contact force, including:
[0072] Adopt the impedance control equation: Establish the dynamic relationship between the contact force and the probe pose, where K p , K v , K f Are set according to prior experience and repeatedly optimized in experiments;
[0073] Use a force sensor to monitor the contact force of the probe in the normal direction in real time, and use the error between the actual force and the target force as feedback to input into the controller. The controller adjusts the displacement of the probe in the normal direction based on this feedback to dynamically correct the contact force;
[0074] S105: Receive the patient's next voice command in real time, and dynamically adjust the progress, contact force, and scanning position of the ultrasonic scan according to the command to ensure that the operation target pose is consistent with the patient's required pose, so as to achieve autonomous ultrasonic scanning with flexible human-computer interaction capabilities. This technical solution is based on scene understanding, task planning, and realizes the autonomous ultrasonic scanning strategy of atomic operations through an imitation learning architecture. Through prompt engineering, the system can input the patient's language commands and environmental information into the large language model to generate low-level tasks. The reasoning ability of the large language model can be realized by calling the ChatGPT API or other large language model API interfaces, mapping high-dimensional and complex language and environmental information into specific scanning subtasks, accurately parsing the patient's needs and clinical scenarios, effectively understanding the patient's intentions, and adjusting the scanning strategy in real time.
[0075] In S101 of this embodiment, use prompt engineering to encapsulate patient instructions, including:
[0076] Clarify the core goals of the prompts according to the characteristics of clinical tasks and patient needs, such as the adjustment range of the contact force and the format of the results to be returned. Clearly define the goals in the prompts, such as "Set the contact force size according to the patient's needs" and "Set the key area of ultrasonic scanning according to the patient's needs", etc.;
[0077] Encapsulate domain knowledge, embed domain-related background knowledge and specific terms in the prompt, so that the prompt conforms to the model's in-context learning ability;
[0078] Explicitly require the model to return the results in a standardized form in the prompt, such as mathematical formulas, parameter lists, or code snippets. Specifying the output format can help the model generate structured results for subsequent direct application.
[0079] In S103 of this embodiment, before using ResNet and PointNet as the backbone networks for processing RGB and point cloud data, the point cloud data needs to be preprocessed, including data cleaning and normalization, such as image normalization, noise filtering, and point cloud downsampling, to improve the quality and consistency of the data and provide more accurate inputs for feature extraction, reflecting the spatial environment and target area characteristics around the ultrasonic probe.
[0080] In S103 of this embodiment, use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, and extract the features of RGB and point cloud to obtain Tokens, including:
[0081] Respectively set the network structures of ResNet and PointNet as the backbone networks for processing RGB images and point cloud data to adapt to the input data;
[0082] Load the pre-trained parameters of ResNet and PointNet;
[0083] Extract the high-dimensional semantic features of the RGB image through ResNet, and extract the spatial structure features of the point cloud data through PointNet, and convert them into semantic Tokens for the imitation learning network to use.
[0084] In S104 of this embodiment, calculate the target angles of each joint of the robotic arm through inverse kinematics to achieve the probe from the current state to the target pose, including:
[0085] Based on the DH parameters, define the link length a, link twist angle α i , joint offset d i and joint angle θ i , and construct the homogeneous transformation matrix A of each joint i ;
[0086] Obtain the homogeneous transformation matrix T from the end effector to the base through successive matrix multiplications T = A1·A2...A n to decouple the target position P and orientation R of the end effector;
[0087] According to the dynamic model, calculate the homogeneous transformation matrix A through inverse kinematics i, the homogeneous transformation matrix T and the target position P and posture R are used to obtain the target angles of each joint of the robotic arm to achieve precise control of the ultrasound probe from the current state to the target posture.
[0088] The system of this technical solution can input the patient's language instructions and environmental information into a large language model to generate low-level tasks. The reasoning ability of the large language model can be realized by calling ChatGPTAPI or other large language model API interfaces, mapping high-dimensional, complex language and environmental information into specific scanning subtasks, and accurately analyzing patient needs and clinical scenarios. Based on this architecture, the system can also receive and understand patient feedback in real time during the scanning process, dynamically adjust the scanning strategy and progress, learn from expert operation demonstrations through imitation learning methods, and achieve highly robust execution of subtasks. Combining the reasoning ability of the large language model with the execution ability of imitation learning, the ultrasound robot's perception and adaptability in complex environments are significantly enhanced.
[0089] In this embodiment, the specific implementation of each link of the technical solution mainly includes: encapsulating patient instructions through prompt word engineering, generating and decomposing low-level tasks based on a large language model, generating posture instructions based on imitation learning methods, precise instruction execution by the controller, real-time perception and online adjustment, and realizing direct interaction capabilities with patients, including online adjustment of parameters such as scanning force, trajectory and scanning position, so that the ultrasound scanning robot has autonomy and is easy to promote in actual clinical applications.
[0090] Reference Figure 3As shown, the ultrasonic scanning robot control system based on an enhanced large language model is used to execute the above ultrasonic scanning robot control method. The control system includes: an acquisition module for receiving real-time instruction information of a patient; an encapsulation module for converting and cleaning the received real-time instruction information to generate a prompt, and using prompt engineering to encapsulate the patient's instructions; a chain reasoning module for calling the API of the large language model to perform multiple and repeated large model calls for tasks according to the requirements of the prompt to complete the iterative decomposition of the prompt; a feature reasoning module for imitating and learning the neural network to infer the pose instruction of the scanning probe at the next moment by combining multi-modal features; a correction module for establishing the dynamic relationship between the contact force and the pose of the ultrasonic probe and dynamically correcting the contact force; an ultrasonic execution module for dynamically adjusting the progress, contact force, and scanning position of the ultrasonic scan according to the instruction, ensuring that the operation target pose is consistent with the patient's required pose, and realizing autonomous ultrasonic scanning with flexible human-computer interaction capabilities. Based on this architecture system, it can also receive and understand the patient's feedback in real time during the scanning process, dynamically adjust the scanning strategy and progress, learn from the expert operation demonstration through the imitation learning method, and achieve highly robust execution of subtasks. By combining the reasoning ability of the large language model and the execution ability of imitation learning, the perception and adaptive ability of the ultrasonic robot in a complex environment are significantly enhanced. In addition, the specific implementation of each link mainly includes: encapsulating the patient's instructions through prompt engineering, generating and decomposing low-level tasks based on the large language model, generating pose instructions based on the imitation learning method, precise instruction execution by the controller, real-time perception and online adjustment, to achieve the direct interaction ability with the patient, including online adjustment of parameters such as scanning force, trajectory, and scanning part, enabling the ultrasonic scanning robot to have autonomy and facilitating its popularization in actual clinical applications.
[0091] In this embodiment, the entire operation process can be controlled by a computer. Signal feedback can be achieved by setting sensors to enable sequential execution of the steps. These are all common knowledge in current automation control and will not be elaborated one by one in this embodiment.
[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for controlling an ultrasonic scanning robot based on an enhanced large language model, characterized in that, Applied to an ultrasound scanning robot, it specifically includes the following steps: S101: Receive the real-time instruction information of the patient, convert and clean the received real-time instruction information, and use prompt engineering to encapsulate the patient instructions; S102: By calling the API of the large language model, transfer the prompt with the encapsulated patient instructions to the model. The model is based on the chain reasoning technology and performs multiple and repeated large model calls according to the requirements of the prompt to complete the iterative decomposition of the prompt; S103: Under the deep learning framework, receive the imitation learning structure instructions generated by the large language model, define the specific structure of the imitation learning model according to the structure instructions, and through combining multi-modal features, the imitation learning neural network infers the pose instruction of the probe at the next moment; S104: Perform contact force control based on the impedance control strategy, establish the dynamic relationship between the contact force and the pose of the ultrasound probe, dynamically correct the contact force, and calculate the target angles of each joint of the robotic arm through inverse kinematics to achieve the transition of the probe from the current state to the target pose; S105: Receive the patient's next voice instruction in real time, and dynamically adjust the progress, contact force, and scanning position of the ultrasound scan according to the instruction to make the operating state of the robotic arm consistent with the patient's needs, so as to achieve autonomous ultrasound scanning with flexible human-computer interaction capabilities.
2. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 1, characterized in that, In S101, receive the real-time instruction information of the patient, convert and clean the received real-time instruction information, and complete the generation of the prompt, including: Receive the real-time voice instruction of the patient. The real-time voice instruction is based on the automatic speech recognition technology to convert the patient's voice input into a transcribed text; Clean and standardize the transcribed text, eliminate meaningless characters and repeated characters in the transcribed text. The cleaning includes denoising and grammar adjustment operations.
3. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 2, wherein, Use prompt engineering to encapsulate the patient instructions, including: Clarify the core target features of the prompt according to the clinical task characteristics and patient needs of the patient, and clearly define the core target features in the prompt; Embed domain-related background knowledge and specific terms in the prompt to make the prompt conform to the context learning ability of the model; Specify the output format in the prompt and clearly require the model to return the result in a standardized form.
4. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 3, wherein, In S102, the model is based on the chain reasoning technology and performs multiple and repeated large model calls according to the requirements of the prompt to complete the iterative decomposition of the prompt, including: Decompose the requirements of the prompt into multiple low-level subtasks with clear logic and distinct levels through multiple and repeated calls of the large model; Use chain reasoning to perform the reasoning of low-level subtasks, and in each step of reasoning, combine the current context subtask with the previous result to predict and generate the next step; The model uses the context consistency check mechanism of the task to verify the order and logical relationship of the generated subtasks. After initially generating the subtasks, the model further optimizes the order of the subtasks.
5. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 4, wherein In S103, under the deep learning framework, receive the imitation learning structure instructions generated by the large language model, define the specific structure of the imitation learning model according to the structure instructions, and through combining multi-modal features, the imitation learning neural network infers the pose instruction of the probe at the next moment, including: Collect environmental data in real time using a depth camera to obtain RGB images and corresponding point cloud data; Use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, extract the features of RGB and point cloud, and obtain Tokens; Input the extracted RGB and point cloud Tokens into the imitation learning neural network, and infer the pose command of the ultrasonic probe at the next moment by combining multi-modal features.
6. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 5, wherein, Before using ResNet and PointNet as the backbone networks for processing RGB and point cloud data, it is necessary to preprocess the point cloud data, clean and normalize the data, including image normalization, noise filtering, and point cloud downsampling, to improve the quality and consistency of the data, provide more accurate input for feature extraction, and reflect the spatial environment and target area characteristics around the ultrasonic probe.
7. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 6, wherein Use ResNet and PointNet as the backbone networks for processing RGB and point cloud data, extract the features of RGB and point cloud, and obtain Tokens, including: Set the network structures of ResNet and PointNet respectively as the backbone networks for processing RGB images and point cloud data to adapt to the input data; Load the pre-trained parameters of ResNet and PointNet; Extract the high-dimensional semantic features of the RGB image through ResNet, and extract the spatial structure features of the point cloud data through PointNet, and convert them into semantic Tokens for use by the imitation learning network.
8. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 7, wherein, In S104, based on the impedance control strategy, perform contact force control, establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, and dynamically correct the contact force, including: Adopt the impedance control equation: Establish the dynamic relationship between the contact force and the probe pose, where K p , K v , K f Is set according to prior experience and obtained through repeated optimization in experiments; Real-time monitor the contact force of the probe in the normal direction through a force sensor, and use the error between the actual force and the target force as the feedback input to the controller. The controller adjusts the displacement of the probe in the normal direction based on this feedback to dynamically correct the contact force.
9. The method for controlling an ultrasonic scanning robot based on an enhanced large language model according to claim 8, wherein, In S104, calculate the target angles of the joints of the robotic arm through inverse kinematics to achieve the probe moving from the current state to the target pose, including: Define the link length a, link twist angle α of the ultrasonic probe manipulator based on the DH parameters i , joint offset d i and joint angle θ i , and construct the homogeneous transformation matrix A for each joint i ; Obtain the homogeneous transformation matrix T from the end to the base through the progressive matrix product T = A1·A2...A n to decouple the target position P and attitude R of the end According to the kinetic model, the homogeneous transformation matrix A is calculated through inverse solution i , the homogeneous transformation matrix T, and the target position P and attitude R to obtain the target angles of each joint of the robotic arm, so as to achieve precise control of the ultrasonic probe from the current state to the target pose.
10. The ultrasonic scanning robot control system based on an enhanced large language model is characterized in that For implementing the ultrasonic scanning robot control method according to any one of claims 1-9, the control system includes: An acquisition module, used to receive the real-time instruction information of the patient; An encapsulation module, used to convert and clean the received real-time instruction information, generate a prompt word, and encapsulate the patient's instruction using prompt engineering; A chain reasoning module, used to call the API of the large language model, and perform multiple and repeated large model calls according to the requirements of the prompt word to complete the iterative decomposition of the prompt word; A feature reasoning module, used to infer the pose command of the scanning probe at the next moment through the imitation learning neural network by combining multi-modal features; A correction module, used to establish the dynamic relationship between the contact force and the pose of the ultrasonic probe, and dynamically correct the contact force; An ultrasonic execution module, used to dynamically adjust the progress, contact force, and scanning position of the ultrasonic scan according to the instruction, ensure that the operation target pose is consistent with the patient's required pose, and achieve autonomous ultrasonic scanning with flexible human-computer interaction capabilities.