Robot thyroid ultrasonic probe automatic adjusting method and system based on AI intensive training

By using a robot based on AI-intensive training to automatically adjust the position and posture of the ultrasound probe in thyroid ultrasound examination, the complex and inefficient problems of traditional manual operations are solved, and efficient and accurate ultrasound examination is achieved.

CN120108679APending Publication Date: 2025-06-06SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267652.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In traditional thyroid ultrasound examination, the operator needs to manually adjust the ultrasound probe, which leads to complex operation, low efficiency, unstable image quality, and long-term operation leads to operator fatigue.

Method used

The automatic adjustment method of robot thyroid ultrasound probe based on AI-intensive training is adopted, and the position and posture of the ultrasound probe are automatically adjusted by establishing a reinforcement learning model and a manual scanning teaching pool.

Benefits of technology

It improves the efficiency and quality of thyroid ultrasound examination, reduces the work burden of operators, and provides more accurate and efficient medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108679A_ABST
    Figure CN120108679A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of thyroid ultrasound probe AI automatic adjustment, and discloses a robot thyroid ultrasound probe automatic adjustment method and system based on AI intensive training. The automatic adjusting method for the robot thyroid ultrasound probe is applied to a medical AI robot and specifically comprises the following steps that S101, a request for establishing an MDP model by a user terminal is received, and the MDP model request carries a state space, an action space and reward function data defined by the robot thyroid ultrasound probe; and S102, acquiring the defined reward function data, inputting the reward function data into a preset content recognition model, and acquiring a content recognition result output by the content recognition model based on the reward function data. The thyroid ultrasound examination efficiency and quality can be improved, the workload of operators can be relieved, and more accurate and efficient medical services are provided for patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AI automatic adjustment of a thyroid ultrasound probe, and in particular to a method and system for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training. Background Art

[0002] With the continuous advancement of medical technology, thyroid ultrasound examination has become one of the important means of diagnosing thyroid diseases. However, the quality of thyroid ultrasound examination is closely related to the operating skills of the ultrasound probe.

[0003] In traditional ultrasound examinations, operators usually need to rely on experience to manually adjust the ultrasound probe in order to obtain high-quality images. This operation method not only requires high operator experience, but also easily causes operator fatigue after long-term operation, reducing examination efficiency and image quality. Therefore, how to accurately adjust the ultrasound probe in an automated way is an important challenge in medical imaging technology. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training. By establishing a reinforcement learning model and a manual scanning teaching pool, the position and posture of the ultrasound probe are automatically adjusted, which solves the problems of traditional manual operation. It can not only improve the efficiency and quality of thyroid ultrasound examination, but also reduce the operator's workload, aiming to solve the problems in the prior art.

[0005] The present invention is implemented in this way: a robot thyroid ultrasound probe automatic adjustment method based on AI intensive training is applied to a medical AI robot, specifically comprising the following steps:

[0006] S101: receiving a request from a user terminal to establish an MDP model, wherein the MDP model request carries state space, action space, and reward function data defined by a robotic thyroid ultrasound probe;

[0007] S102: Obtain defined reward function data, input the reward function data into a preset content recognition model, obtain a content recognition result output by the content recognition model based on the reward function data, and obtain, from the content recognition result, an instant reward value obtained by the reward function from environmental factors after performing a certain action;

[0008] S103: Establish a manual scanning teaching pool to collect high-quality initial experience data for the reinforcement learning algorithm to learn in the early stage and complete the establishment of the initialization model;

[0009] S104: Obtain the initialization model, select different adjustment actions through interactive optimization strategy reinforcement learning training with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy, and complete the cumulative rewards for strategy update;

[0010] S105: The model is continuously optimized in each training cycle, and the position and posture of the ultrasound probe are automatically adjusted based on the real-time feedback of the image quality and lesion recognition of the real-time ultrasound scanning.

[0011] Further, in S101, receiving a request from a user terminal to establish an MDP model includes:

[0012] Receive a connection request sent by a user terminal through a preset network, wherein the connection request is used to request to establish a connection with the medical AI robot;

[0013] Detecting whether the current account of the user terminal is the target account;

[0014] If the current account of the user terminal is the target account, a connection is established with the user terminal according to the connection request, and after the connection is completed, a request for establishing an MDP model is received from the user terminal.

[0015] Furthermore, the MDP model request carries the state space, action space and reward function data defined by the robotic thyroid ultrasound probe, including:

[0016] The defined state space includes: the position and posture of the robotic thyroid ultrasound probe, the quality of the ultrasound image, the location of the lesion area, and time information.

[0017] The defined action space includes: all actions performed by the robotic thyroid ultrasound probe, and state adjustment operations performed in the state space of all actions performed;

[0018] The defined reward function data include: the quality of ultrasound images and the improvement of lesion identification accuracy, which are used to evaluate the impact of each robotic thyroid ultrasound probe action on the system.

[0019] Further, in S102, in the content recognition result, an immediate reward value obtained by a reward function from environmental factors after performing a certain action is obtained, and the environmental factors based on the reward function include:

[0020] The quality of ultrasound images, used to assess the similarity or quality between the generated or processed images and the original or desired images;

[0021] Lesion classification accuracy, used to evaluate the accuracy of the medical AI robot in classifying lesions;

[0022] Lesion segmentation accuracy, used to evaluate the accuracy of the medical AI robot in segmenting the lesion area;

[0023] Action cost is used to evaluate the cost or price when a medical AI robot performs an action.

[0024] Furthermore, the reward function obtains the weighted sum of the immediate reward values ​​in the environmental factors as follows:

[0025] R = w1*image quality + w2*lesion classification accuracy + w3*lesion segmentation accuracy - w4*action cost;

[0026] Among them, w1, w2, w3, and w4 are the weights of each part, which are used to adjust the relative importance of each part in the total reward. w1, w2, w3, and w4 can be set according to the needs and importance of the actual task.

[0027] Furthermore, in S103, a manual scanning teaching pool is established to collect high-quality initial experience data, including:

[0028] The medical AI robot sends a request to the medical system to obtain the stored thyroid ultrasound examination image data, wherein the request includes a generation path based on the convergence of the medical AI robot acceleration model;

[0029] The medical system receives the request sent by the medical AI robot, identifies the generation path of the convergence of the medical AI robot acceleration model, and performs pathological classification on the stored thyroid ultrasound examination image data according to the generation path;

[0030] Obtain thyroid disease items after pathological classification, classify and store them in sequence, and transmit them to the medical AI robot to complete the establishment of the initialization model.

[0031] Further, in S104, an initialization model is obtained, and different adjustment actions are selected through interactive optimization strategy reinforcement learning training with the environment to obtain feedback about the environment to bring higher rewards, including:

[0032] Initialize the obtained initialization model, including the state space, action space, and reward function, and then use the data provided by the constructed manual scanning teaching pool for pre-training;

[0033] By choosing different adjustment actions, we get feedback about the environment and obtain the reward set by the reward function.

[0034] Furthermore, the Q-learning algorithm is used to optimize the strategy and complete the cumulative reward for strategy update. The Q-learning algorithm process includes:

[0035] Initialize the Q table, set the Q values ​​of all state-action pairs to the initial values, and select actions according to the ε-greedy strategy in each state, randomly select actions with a probability of ε and select the action with the largest current Q value with a probability of 1-ε;

[0036] The robotic thyroid ultrasound probe performs the selected action and observes the next state, the immediate reward obtained, and whether the termination condition is reached. If the termination condition is reached, the Q value is updated. The Q value update uses the Bellman equation formula:

[0037]

[0038] Where α is the learning rate, γ is the discount factor, R(s,a) is the immediate reward for performing action a in state s, s' is the next state, and a' is the possible action for the next state;

[0039] Repeat the above steps until the maximum number of iterations is reached or the Q value converges.

[0040] Compared with the prior art, the method and system for automatic adjustment of robot thyroid ultrasound probe based on AI intensive training provided by the present invention have the following beneficial effects:

[0041] 1. By establishing a reinforcement learning model and a manual scanning teaching pool, the position and posture of the ultrasound probe are automatically adjusted, solving the problem of traditional manual operation. This can not only improve the efficiency and quality of thyroid ultrasound examinations, but also reduce the workload of operators, providing patients with more accurate and efficient medical services;

[0042] 2. The main logic of this solution is to establish a Markov decision process (MDP) model, reward function design, construction of a manual scanning teaching pool, reinforcement learning training process, real-time feedback and adjustment of ultrasound scanning, and optimize the strategy with the Q-learning algorithm. The Q-learning algorithm is a model-free prediction algorithm. It evaluates the expected utility of taking an action under a given state by learning an action value Q function, thereby determining the value of the model for subsequent probe adjustments during learning. The reinforcement learning algorithm learns in the early stages, and the data in the teaching pool will be used to initialize the model, and will gradually increase as the training progresses. By continuously optimizing the model, it will learn how to automatically adjust the position and posture of the ultrasound probe according to the real-time image quality and lesion recognition.

[0043] The robot thyroid ultrasound probe automatic adjustment system based on AI intensive training performs the above-mentioned robot thyroid ultrasound probe automatic adjustment method, and the robot thyroid ultrasound probe automatic adjustment system includes:

[0044] A request module, used to receive a request from a user terminal to establish an MDP model;

[0045] A reward function definition module, used to obtain a content recognition result output by the content recognition model based on the reward function data, and obtain, in the content recognition result, an instant reward value obtained by the reward function from environmental factors after performing a certain action;

[0046] The model building module is used to establish a manual scanning teaching pool, collect high-quality initial experience data for the reinforcement learning algorithm to learn in the early stage, and complete the establishment of the initialization model;

[0047] The reinforcement training module is used to initialize the model, select different adjustment actions through interaction with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy;

[0048] The probe adjustment module is used to automatically adjust the position and posture of the ultrasound probe based on real-time feedback of the image quality and lesion recognition of the real-time ultrasound scan.

[0049] Specifically, the intensive training module includes:

[0050] An action acquisition unit is used to set the Q values ​​of all state-action pairs to initial values. In each state, the robotic thyroid ultrasound probe selects an action according to an ε-greedy strategy, randomly selects an action with a probability of ε, and selects the action with the largest current Q value with a probability of 1-ε;

[0051] The Q-value updating unit is used to determine whether the robot thyroid ultrasound probe performs the selected action, and observe the next state, the immediate reward obtained, and whether the termination condition is met. If the termination condition is met, the Bellman equation formula is used to complete the Q-value update. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a schematic flow chart of the method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training proposed in the present invention;

[0053] Figure 2 A schematic block diagram of the process of receiving a request from a user terminal to establish an MDP model in the method for automatically adjusting a robot thyroid ultrasound probe based on AI intensive training proposed by the present invention;

[0054] Figure 3 This is a schematic diagram of the structure of the robot thyroid ultrasound probe automatic adjustment system based on AI intensive training proposed by the present invention;

[0055] Figure 4This is a structural schematic diagram of the medium-term training module of the robot thyroid ultrasound probe automatic adjustment system based on AI enhanced training proposed in the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0057] The implementation of the present invention is described in detail below in conjunction with specific embodiments.

[0058] The same or similar numbers in the drawings of this embodiment correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limitations on the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0059] Reference Figure 1-2 As shown, the automatic adjustment method of the robot thyroid ultrasound probe based on AI intensive training is applied to the medical AI robot, and specifically includes the following steps:

[0060] S101: receiving a request from a user terminal to establish an MDP model, where the MDP model request carries state space, action space, and reward function data defined by a robotic thyroid ultrasound probe;

[0061] The step of receiving a request from a user terminal to establish an MDP model includes:

[0062] Receive a connection request sent by a user terminal through a preset network, where the connection request is used to request to establish a connection with the medical AI robot;

[0063] Check whether the current account of the user terminal is the target account;

[0064] If the current account of the user terminal is the target account, connect with the user terminal according to the connection request, and after the connection is completed, receive the user terminal's request to establish the MDP model;

[0065] S102: Obtain defined reward function data, input the reward function data into a preset content recognition model, obtain a content recognition result output by the content recognition model based on the reward function data, and obtain, from the content recognition result, an instant reward value obtained by the reward function from environmental factors after performing a certain action;

[0066] Among them, in the content recognition result, the instant reward value obtained by the reward function from the environmental factors after performing a certain action is obtained. The environmental factors based on the reward function include:

[0067] The quality of ultrasound images, used to assess the similarity or quality between the generated or processed images and the original or desired images;

[0068] Lesion classification accuracy, used to evaluate the accuracy of the medical AI robot in classifying lesions;

[0069] Lesion segmentation accuracy, used to evaluate the accuracy of the medical AI robot in segmenting the lesion area;

[0070] Action cost, used to evaluate the cost or price of the medical AI robot when performing an action;

[0071] Specifically, the reward function obtains the weighted sum of the immediate reward values ​​in the environmental factors as follows:

[0072] R = w1*image quality + w2*lesion classification accuracy + w3*lesion segmentation accuracy - w4*action cost;

[0073] Among them, w1, w2, w3, and w4 are the weights of each part, which are used to adjust the relative importance of each part in the total reward. w1, w2, w3, and w4 can be set according to the needs and importance of the actual task;

[0074] S103: Establish a manual scanning teaching pool to collect high-quality initial experience data for the reinforcement learning algorithm to learn in the early stage and complete the establishment of the initialization model;

[0075] S104: Obtain the initialization model, select different adjustment actions through interactive optimization strategy reinforcement learning training with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy, and complete the cumulative rewards for strategy update;

[0076] Among them, the initialization model is obtained, and different adjustment actions are selected through interactive optimization strategy reinforcement learning training with the environment to obtain feedback about the environment to bring higher rewards, including:

[0077] Initialize the obtained initialization model, including the state space, action space, and reward function, and then use the data provided by the constructed manual scanning teaching pool for pre-training;

[0078] By selecting different adjustment actions, we can obtain feedback about the environment and bring rewards set by the reward function.

[0079] S105: The model is continuously optimized in each training cycle, and the position and posture of the ultrasound probe are automatically adjusted based on real-time feedback on the image quality and lesion recognition of real-time ultrasound scanning. By establishing a reinforcement learning model and a manual scanning teaching pool, the position and posture of the ultrasound probe are automatically adjusted, which solves the problems of traditional manual operation. It can not only improve the efficiency and quality of thyroid ultrasound examinations, but also reduce the workload of operators, providing patients with more accurate and efficient medical services.

[0080] In this embodiment, the MDP model request carries the state space, action space, and reward function data defined by the robotic thyroid ultrasound probe, including:

[0081] The defined state space includes: the position and posture of the robotic thyroid ultrasound probe, the quality of the ultrasound image, the location of the lesion area, and time information.

[0082] The defined action space includes: all actions performed by the robotic thyroid ultrasound probe and the state adjustment operations performed in the state space of all actions performed;

[0083] The defined reward function data include: the quality of ultrasound images and the improvement of lesion identification accuracy, which are used to evaluate the impact of each robotic thyroid ultrasound probe action on the system.

[0084] In this embodiment, a manual scanning teaching pool is established to collect high-quality initial experience data, including:

[0085] The medical AI robot sends a request to the medical system to obtain the stored thyroid ultrasound examination image data, where the request includes a generation path based on the convergence of the medical AI robot acceleration model;

[0086] The medical system receives the request sent by the medical AI robot, identifies the generation path of the convergence of the medical AI robot acceleration model, and performs pathological classification on the stored thyroid ultrasound examination image data according to the generation path;

[0087] Obtain thyroid disease items after pathological classification, classify and store them in sequence, and transmit them to the medical AI robot to complete the establishment of the initialization model.

[0088] In this embodiment, the Q-learning algorithm is used to optimize the strategy and complete the cumulative reward for strategy update. The Q-learning algorithm process includes:

[0089] Initialize the Q table, set the Q values ​​of all state-action pairs to the initial values, and select actions according to the ε-greedy strategy in each state, randomly select actions with a probability of ε and select the action with the largest current Q value with a probability of 1-ε;

[0090] The robotic thyroid ultrasound probe performs the selected action and observes the next state, the immediate reward obtained, and whether the termination condition is reached. If the termination condition is reached, the Q value is updated. The Q value update uses the Bellman equation formula:

[0091]

[0092] Where α is the learning rate, γ is the discount factor, R(s,a) is the immediate reward for performing action a in state s, s' is the next state, and a' is the possible action for the next state;

[0093] Repeat the above steps until the maximum number of iterations is reached or the Q value converges.

[0094] The main logic of this solution is to establish a Markov decision process (MDP) model, reward function design, construction of a manual scanning teaching pool, reinforcement learning training process, real-time feedback and adjustment of ultrasound scanning, and use the Q-learning algorithm to optimize the strategy. The Q-learning algorithm is a model-free prediction algorithm. It evaluates the expected utility of taking an action under a given state by learning an action value Q function, thereby determining the value of the model for subsequent probe adjustments during learning. The reinforcement learning algorithm learns in the early stages, and the data in the teaching pool will be used to initialize the model. As the training progresses, it will gradually increase. By continuously optimizing the model, it will learn how to automatically adjust the position and posture of the ultrasound probe according to the real-time image quality and lesion recognition.

[0095] For example, the following is a simple MDP model implementation example:

[0096]

[0097]

[0098] Reference Figure 3-4As shown, a robot thyroid ultrasound probe automatic adjustment system based on AI reinforcement training executes the above-mentioned robot thyroid ultrasound probe automatic adjustment method, and the robot thyroid ultrasound probe automatic adjustment system includes: a request module, which is used to receive a request from a user terminal to establish an MDP model; a reward function definition module, which is used to obtain a content recognition result output by a content recognition model based on reward function data, and in the content recognition result, obtain an instant reward value obtained by the reward function from environmental factors after performing a certain action; a model building module, which is used to establish an artificial scanning teaching pool, collect high-quality initial experience data, and provide reinforcement learning algorithms for learning in the early stages to complete the establishment of an initialization model; a reinforcement learning algorithm ... The training module is used to initialize the model, select different adjustment actions through interactive optimization strategy reinforcement learning training with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy; the probe adjustment module is used to automatically adjust the position and posture of the ultrasound probe according to the real-time ultrasound scanning image quality and lesion recognition situation. By establishing a reinforcement learning model and a manual scanning teaching pool, the position and posture of the ultrasound probe are automatically adjusted, which solves the problem of traditional manual operation. It can not only improve the efficiency and quality of thyroid ultrasound examination, but also reduce the workload of operators, providing patients with more accurate and efficient medical services;

[0099] In this embodiment, the reinforcement training module includes: an action acquisition unit, which is used to set the Q value of all state-action pairs to an initial value. In each state, the robotic thyroid ultrasound probe selects an action according to the ε-greedy strategy, randomly selects an action with a probability of ε, and selects the action with the largest current Q value with a probability of 1-ε; a Q value update unit, which is used to determine whether the robotic thyroid ultrasound probe executes the selected action, and observes the next state, the immediate reward obtained, and whether the termination condition is met. If the termination condition is met, the Bellman equation formula is used to complete the Q value update. The Q-learning algorithm is a model-free prediction algorithm, which evaluates the expected utility of taking an action in a given state by learning an action value Q function, thereby determining the value of the model to subsequent probe adjustments during learning, and the reinforcement learning algorithm is learned in the early stage, and the data in the teaching pool will be used to initialize the model.

[0100] In this embodiment, the entire operation process can be controlled by a computer to realize automatic operation control, perform signal feedback, and implement the steps in sequence. These are all common knowledge of current automatic control and will not be described in detail in this embodiment.

[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training, characterized in that: Applied to medical AI robots, the specific steps include: S101: receiving a request from a user terminal to establish an MDP model, wherein the MDP model request carries state space, action space, and reward function data defined by a robotic thyroid ultrasound probe; S102: Obtain defined reward function data, input the reward function data into a preset content recognition model, obtain a content recognition result output by the content recognition model based on the reward function data, and obtain, from the content recognition result, an instant reward value obtained by the reward function from environmental factors after performing a certain action; S103: Establish a manual scanning teaching pool to collect high-quality initial experience data for the reinforcement learning algorithm to learn in the early stage and complete the establishment of the initialization model; S104: Obtain the initialization model, select different adjustment actions through interactive optimization strategy reinforcement learning training with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy, and complete the cumulative rewards for strategy update; S105: The model is continuously optimized in each training cycle, and the position and posture of the ultrasound probe are automatically adjusted based on the real-time feedback of the image quality and lesion recognition of the real-time ultrasound scanning.

2. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 1, characterized in that: In S101, receiving a request from a user terminal to establish an MDP model includes: Receive a connection request sent by a user terminal through a preset network, wherein the connection request is used to request to establish a connection with the medical AI robot; Detecting whether the current account of the user terminal is the target account; If the current account of the user terminal is the target account, a connection is established with the user terminal according to the connection request, and after the connection is completed, a request for establishing an MDP model is received from the user terminal.

3. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 2, characterized in that: The MDP model request carries the state space, action space and reward function data defined by the robotic thyroid ultrasound probe, including: The defined state space includes: the position and posture of the robotic thyroid ultrasound probe, the quality of the ultrasound image, the location of the lesion area, and time information. The defined action space includes: all actions performed by the robotic thyroid ultrasound probe, and state adjustment operations performed in the state space of all actions performed; The defined reward function data include: the quality of ultrasound images and the improvement of lesion identification accuracy, which are used to evaluate the impact of each robotic thyroid ultrasound probe action on the system.

4. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 3, characterized in that: In S102, in the content recognition result, an immediate reward value obtained by a reward function from environmental factors after performing a certain action is obtained, and the environmental factors based on which the reward function is based include: The quality of ultrasound images, used to assess the similarity or quality between the generated or processed images and the original or desired images; Lesion classification accuracy, used to evaluate the accuracy of the medical AI robot in classifying lesions; Lesion segmentation accuracy, used to evaluate the accuracy of the medical AI robot in segmenting the lesion area; Action cost is used to evaluate the cost or price when a medical AI robot performs an action.

5. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 4, characterized in that: The reward function obtains the weighted sum of the immediate reward values ​​in the environmental factors as follows: R = w1*image quality + w2*lesion classification accuracy + w3*lesion segmentation accuracy - w4*action cost; Among them, w1, w2, w3, and w4 are the weights of each part, which are used to adjust the relative importance of each part in the total reward. w1, w2, w3, and w4 can be set according to the needs and importance of the actual task.

6. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 5, characterized in that: In S103, a manual scanning teaching pool is established to collect high-quality initial experience data, including: The medical AI robot sends a request to the medical system to obtain the stored thyroid ultrasound examination image data, wherein the request includes a generation path based on the convergence of the medical AI robot acceleration model; The medical system receives the request sent by the medical AI robot, identifies the generation path of the convergence of the medical AI robot acceleration model, and performs pathological classification on the stored thyroid ultrasound examination image data according to the generation path; Obtain thyroid disease items after pathological classification, classify and store them in sequence, and transmit them to the medical AI robot to complete the establishment of the initialization model.

7. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 6, characterized in that: In S104, an initialization model is obtained, and different adjustment actions are selected through interactive optimization strategy reinforcement learning training with the environment to obtain feedback about the environment to bring higher rewards, including: Initialize the obtained initialization model, including the state space, action space, and reward function, and then use the data provided by the constructed manual scanning teaching pool for pre-training; By choosing different adjustment actions, we get feedback about the environment and obtain the reward set by the reward function.

8. The method for automatic adjustment of a robot thyroid ultrasound probe based on AI intensive training as claimed in claim 7, characterized in that: The Q-learning algorithm is used to optimize the strategy and complete the cumulative rewards for strategy update. The Q-learning algorithm process includes: Initialize the Q table, set the Q values ​​of all state-action pairs to the initial values, and select actions according to the ε-greedy strategy in each state, randomly select actions with a probability of ε and select the action with the largest current Q value with a probability of 1-ε; The robotic thyroid ultrasound probe performs the selected action and observes the next state, the immediate reward obtained, and whether the termination condition is reached. If the termination condition is reached, the Q value is updated. The Q value update uses the Bellman equation formula: Where α is the learning rate, γ is the discount factor, R(s,a) is the immediate reward for performing action a in state s, s' is the next state, and a' is the possible action for the next state; Repeat the above steps until the maximum number of iterations is reached or the Q value converges.

9. The robot thyroid ultrasound probe automatic adjustment system based on AI intensive training is characterized by: The method for automatically adjusting a robot thyroid ultrasound probe according to any one of claims 1 to 8 is implemented, wherein the robot thyroid ultrasound probe automatic adjustment system comprises: A request module, used to receive a request from a user terminal to establish an MDP model; A reward function definition module, used to obtain a content recognition result output by the content recognition model based on the reward function data, and in the content recognition result, obtain an instant reward value obtained by the reward function from environmental factors after performing a certain action; The model building module is used to establish a manual scanning teaching pool, collect high-quality initial experience data for the reinforcement learning algorithm to learn in the early stage, and complete the establishment of the initialization model; The reinforcement training module is used to initialize the model, select different adjustment actions through interaction with the environment, obtain feedback about the environment to bring higher rewards, and use the Q-learning algorithm to optimize the strategy; The probe adjustment module is used to automatically adjust the position and posture of the ultrasound probe based on real-time feedback of the image quality and lesion recognition of the real-time ultrasound scan.

10. The robot thyroid ultrasound probe automatic adjustment system based on AI intensive training as claimed in claim 9, characterized in that: The intensive training modules include: An action acquisition unit is used to set the Q values ​​of all state-action pairs to initial values. In each state, the robotic thyroid ultrasound probe selects an action according to an ε-greedy strategy, randomly selects an action with a probability of ε, and selects the action with the largest current Q value with a probability of 1-ε; The Q value update unit is used to determine whether the robot thyroid ultrasound probe performs the selected action, and observe the next state, the immediate reward obtained, and whether the termination condition is met. If the termination condition is met, the Bellman equation formula is used to complete the Q value update.