Ultrasonic view quality control method based on reinforcement learning algorithm
By combining a unified reference clock, anatomical key point atlas, and graph Transformer, the problems of unstable multi-source data fusion and insufficient closed-loop control in ultrasound imaging are solved, achieving efficient and accurate standard view acquisition and quality control.
Patent Information
- Application Number
- CN202511731950.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-06
AI Technical Summary
Existing ultrasound imaging technology suffers from several shortcomings in acquiring standard views, including insufficient modeling of observable states, unstable fusion of multi-source data, and a lack of closed-loop control and planning capabilities. This makes it difficult to stably express anatomical visibility and posture deviations, and also lacks real-time optimization capabilities.
A unified reference clock is used for time synchronization and drift correction. An atlas of anatomical key points is constructed and embedded with guide priors and physical constraints. A graph Transformer is used for time-series filtering to output belief states. Combined with TD-MPC2 short field-of-view rolling optimization, probe micro-movements and equipment parameters are planned to achieve closed-loop quality control.
It improves the accuracy and efficiency of quality control, enhances the ability to resist occlusion and drift, reduces manual adjustment and repeated scanning, and achieves real-time operable optimization under safety constraints.
Smart Images

Figure FT_1
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging, and more particularly to an ultrasound image quality control method based on reinforcement learning algorithms. Background Technology
[0002] Ultrasound imaging is widely used in cardiovascular, obstetric, and abdominal organ examinations due to its high real-time performance and low cost. In clinical practice, the acquisition and quality control of standard views directly affect diagnostic conclusions. Existing research mainly focuses on standard view recognition and quality assessment, including view classification and segmentation based on convolutional neural networks, anatomical localization based on key points, and short sequence feature fusion based on temporal networks or Transformers. To improve guidance capabilities, some studies have introduced probe inertial measurement unit data or external tracking and attempted to provide operational prompts, while others have utilized automatic gain and depth adaptation of device parameters. However, overall, these methods are mostly offline assessments or weak closed-loop guidance.
[0003] The existing technology still has the following shortcomings:
[0004] 1. Insufficient modeling of some observable states: Affected by occlusion, speckle noise, breathing / heartbeat and operator jitter, single-frame or short-segment methods are difficult to stably express time-varying anatomical visibility and posture deviation, and lack belief state estimation and robust temporal fusion for POMDP.
[0005] 2. Insufficient fusion of multi-source data and utilization of structured priors: Video, inertial measurement unit and equipment parameters are often fused by simple splicing or post-calibration, and timestamp inconsistencies and drifts are difficult to be fully corrected; there is a lack of unified and computable structured representation of the topological relationships between anatomy, the multi-structure joint constraints in clinical guidelines and acoustic-physical constraints (incident angle, near / far field coverage);
[0006] 3. Insufficient closed-loop control and planning capabilities: Existing solutions mostly rely on post-event judgment or experience-based prompts, lacking real-time rolling optimization with a clear quality control cost function as the goal; existing reinforcement learning is mostly model-independent and has low sample efficiency, making it difficult to jointly optimize probe micro-movements and equipment parameters under safety constraints, resulting in slow convergence and poor operability.
[0007] Therefore, a method for quality control of ultrasound images that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose an ultrasound image quality control method based on reinforcement learning algorithms. Addressing the problems in existing technologies regarding standard view acquisition, such as difficulty in modeling some observable states, unstable clock alignment and fusion of multi-source data (video, inertial measurement unit, equipment parameters), and lack of closed-loop optimization involving clinical guidelines and acoustic physical constraints, this invention proposes a unified reference clock for time synchronization and drift correction, constructing an anatomical keypoint atlas carrying pose encoding and equipment context, embedding guideline priors with hyperedges and introducing physical constraint edges, using graph Transformer temporal filtering to output belief states, and combining TD-MPC2 short-field-of-view rolling optimization to jointly plan probe micro-movements and equipment parameters for closed-loop judgment. This invention achieves the technical effects of improving quality control accuracy and standard view acquisition efficiency, enhancing anti-occlusion and anti-drift capabilities, reducing manual intervention and repeated scanning, and achieving real-time operable optimization under safety constraints.
[0009] An ultrasound image quality control method based on a reinforcement learning algorithm according to an embodiment of the present invention is characterized by comprising:
[0010] S1. Acquire ultrasonic video frame sequences, inertial measurement unit data, and equipment parameter data; perform time synchronization and drift correction, and output the data.
[0011] S2. Input the aligned ultrasound video frame sequence, inertial measurement unit data, and equipment parameter data into the anatomical key point detection and atlas construction module to construct an anatomical key point atlas;
[0012] S3. Construct an anatomical hypermap by embedding guide priors and physical constraints into the anatomical key point atlas;
[0013] S4. Perform temporal attention fusion and state update on the Transformer filter of the anatomical hypergraph input graph, and output the belief state vector;
[0014] S5. Input the belief state vector into the TD-MPC2 short-field planner, construct the quality control cost function and perform short-field rolling optimization, and output action suggestions;
[0015] S6. Perform the suggested actions on the ultrasonic equipment and collect observations of the execution process, outputting the ultrasonic video frame sequence, inertial measurement unit data, and equipment parameter data after execution;
[0016] S7. The executed ultrasonic video frame sequence, inertial measurement unit data, and equipment parameter data are processed sequentially through steps S1 to S4 to obtain the updated belief state vector, which is then input into the quality control judgment module. The process terminates when the quality control judgment result is qualified; otherwise, steps S5 to S7 are repeated.
[0017] Optionally, step S1 specifically includes:
[0018] A timestamp is generated for each ultrasonic video frame, each set of inertial measurement unit data, and each sampling of equipment parameters under a unified reference clock; ultrasonic video frame sequences are continuously acquired and output under the timestamp markings;
[0019] Inertial measurement unit (IMU) data is continuously acquired and output under the timestamp marker, and the IMU data includes at least triaxial acceleration and triaxial angular velocity.
[0020] Under the timestamp mark, device parameter data is read and output in real time through the device interface. The device parameter data includes at least gain, depth, focus position and transmission frequency.
[0021] Using the timestamps of the ultrasonic video frame sequence as a reference time axis, time offset estimation and time-varying drift compensation are performed on the inertial measurement unit data and equipment parameter data based on the reference time axis. The time offset estimation and time-varying drift compensation are obtained by minimizing the cross-modal alignment error on the reference time axis to obtain the corrected timestamps.
[0022] Interpolation and resampling are performed on the inertial measurement unit data and equipment parameter data under the corrected timestamp to make the sampling time consistent with the reference time axis, and timestamp correction and necessary interpolation processing are performed on the dropped and repeated frames of the ultrasonic video frame sequence.
[0023] After completing time synchronization and drift correction, the system outputs aligned ultrasonic video frame sequences, aligned inertial measurement unit data, and aligned device parameter data.
[0024] Optionally, step S2 specifically includes:
[0025] Using the aligned ultrasound video frame sequence as image input, anatomical key points are detected in each frame, and coordinates and visibility confidence are generated for each anatomical key point.
[0026] The relative pose code of the probe is calculated using the aligned inertial measurement unit data, and the relative pose code is used as the node attribute and edge attribute of the anatomical key point map;
[0027] Based on the spatial relationship between anatomical key points, edges are generated for the anatomical key point atlas, and the edges are used to represent the spatial adjacency relationship between adjacent key points;
[0028] The aligned device parameter data is mapped to a spectrum context attribute, which includes at least the current values of gain, depth, focal position, and transmission frequency.
[0029] Output an anatomical key point atlas, which includes nodes, edges, and atlas context attributes. Nodes contain the coordinates and visibility confidence of the anatomical key points, edges contain the spatial adjacency relationship of adjacent key points, and nodes and edges carry probe relative pose encoding.
[0030] Definitions:
[0031] The anatomical key points are predefined image markers located on the target anatomical structure, used to characterize the geometric relationship of the standard view, and their positions are taken in the two-dimensional image coordinate system of the current frame;
[0032] The coordinates are the pixel coordinates of the anatomical key points in the current frame's two-dimensional image coordinate system, preferably represented by the top left corner of the image as the origin, the horizontal axis to the right as the x-axis, and the vertical axis downwards as the y-axis;
[0033] The visibility confidence score is a numerical index indicating the visibility of the corresponding anatomical key point in the current frame, with a value range of [0,1], and is estimated by the key point detection algorithm;
[0034] The anatomical key point detection and atlas construction module is a functional unit that receives aligned multimodal input and outputs an anatomical key point atlas, including at least key point localization, attribute calculation and graph structure generation subunits;
[0035] The probe relative pose encoding is a vector representation of the rotation and translation increments of the ultrasonic probe relative to a preset reference pose, obtained from aligned inertial measurement unit data and imaging geometry calculations, and includes at least rotation and translation parameters.
[0036] The node is a vertex in the anatomical key point atlas, corresponding to a single anatomical key point instance;
[0037] The edge is the line connecting two nodes in the anatomical key point atlas, used to express the relationship between the two anatomical key points;
[0038] The node attributes are a set of features associated with each node, including at least the node's coordinates, visibility confidence, and probe relative pose encoding;
[0039] The edge attribute is a set of features associated with each edge, including at least the probe relative pose code associated with that edge;
[0040] The spatial adjacency relationship is a relationship rule used to determine whether two anatomical key points are connected. It can be determined based on at least the distance threshold between key points, a predefined anatomical topology, or a proximity criterion based on triangulation.
[0041] The adjacent key points are two anatomical key points that have a spatial adjacency relationship.
[0042] The context attributes of the map are context information that is globally associated with the entire map, including at least the current values of gain, depth, focal position and transmission frequency in the aligned device parameter data;
[0043] The anatomical key point atlas is a graph data structure consisting of nodes, edges and their corresponding attributes, and atlas context attributes, used to uniformly express information such as anatomical geometry, visibility, and probe pose.
[0044] Optionally, step S3 specifically includes:
[0045] Using the anatomical key point atlas as input, the multi-structure joint constraints of the standard view are written into the anatomical key point atlas in the form of hyperedges according to clinical guidelines. The hyperedges are used to represent at least one set of geometric and visibility conditions that are simultaneously satisfied by the anatomical structures. The geometric and visibility conditions include at least one of symmetry, collinearity, relative angle range, relative distance range, or coverage threshold.
[0046] Weights are assigned to each hyperedge based on thresholds from clinical guidelines;
[0047] Physical constraint attributes are constructed on the edges of the map. The physical constraint attributes include at least one of the following: incident angle deviation, near-field coverage, far-field coverage, or expected reflection intensity. The physical constraint attributes are calculated by the probe relative pose encoding, map context attributes, and anatomical key point coordinates.
[0048] The relative pose of the probe is preserved as node attributes and edge attributes;
[0049] After completing the guide prior embedding and physical constraint construction, the anatomical hypergraph is output.
[0050] Definition of noun:
[0051] The aforementioned guideline-based prior embedding is the process of incorporating standard view-related geometric and visibility constraints into an anatomical keypoint atlas in a structured form, based on clinical guidelines, including expressing joint constraints with hyperedges and assigning weights.
[0052] The standard view is a standardized ultrasound imaging view for diagnosis as specified in clinical guidelines, which meets the display requirements of specific anatomical structures and geometric constraints.
[0053] The multi-structure joint constraint is a set of conditions that are proposed simultaneously for at least one set of anatomical structures and must be satisfied together, used to characterize the comprehensive judgment requirements of the standard view;
[0054] The hyperedge is a high-order relational unit in the graph structure that can simultaneously associate two or more nodes, and is used to represent multi-structure joint constraints;
[0055] The hyperedge weight is a numerical weight that measures the relative importance of the corresponding hyperedge in the determination and optimization process, and can be set according to the threshold or priority in clinical guidelines;
[0056] The anatomical structures are a collection of organs or their components that can be identified and used for diagnosis in ultrasound images, and are the objects of joint constraint;
[0057] The geometric and visibility conditions are the categories of conditions used to determine whether the joint constraint is valid, and at least include symmetry, collinearity, relative angle range, relative distance range, and coverage threshold.
[0058] The symmetry refers to the degree to which the geometric relationship between the target anatomical structure or key point and the reference axis or center is satisfied, which is a mirror image or approximately a mirror image.
[0059] Collinearity refers to the degree to which multiple key points or structural boundary points are located on the same straight line or approximately a straight line.
[0060] The relative angle range is a constraint that the angle between two anatomical axes or boundary tangents must fall within a preset range.
[0061] The relative distance range is a constraint that the distance between key points or structures must fall within a preset range.
[0062] The coverage threshold is the lower limit of the proportion or area ratio that the target structure should achieve in the current imaging field of view.
[0063] The physical constraint properties are a set of computable properties related to imaging acoustics and geometry and defined on the edge of the map, including at least incident angle deviation, near-field coverage, far-field coverage and expected reflection intensity;
[0064] The incident angle deviation is the degree to which the angle between the sound beam direction and the normal of the target structure interface deviates from the ideal incident angle, and can be expressed by the angle or its normalized value.
[0065] The near-field coverage rate is the proportion of the target structure covered by the sound beam within the near-field depth range before the focal point.
[0066] The far-field coverage rate is the proportion of the target structure covered by the sound beam within the far-field depth range after the focal point.
[0067] The expected reflection intensity is a predictive index of the echo signal intensity under the current equipment parameters and tissue model assumptions, used to reflect the physical accessibility of imaging quality;
[0068] The physical constraint construction is a process of calculating and associating the above physical constraint attributes with edges based on the probe relative pose encoding, map context attributes and anatomical key point coordinates.
[0069] The anatomical hypergraph is an extended graph structure formed by adding hyperedges and their weights to the anatomical key point atlas and attaching physical constraint attributes to the edges. It is used to uniformly express the prior constraints of the guide and the physical features of the imaging.
[0070] Optionally, step S4 specifically includes:
[0071] The anatomical hypergraph is used as the input at the current time step, and cross-time fusion is performed through the temporal attention of the graph Transformer filter under the sequence of anatomical hypergraphs at adjacent time steps. The cross-time fusion takes the node attributes, edge attributes and hyperedge weights of the anatomical hypergraph as attention inputs, and uses physical constraint attributes as attention biases to control weight allocation.
[0072] Update node embeddings and edge embeddings after completing cross-time fusion;
[0073] Based on the updated node embedding and edge embedding, a belief state vector is calculated and output through a readout function. The belief state vector includes at least anatomical coverage, angular deviation, symmetry index, and confidence in the visibility of key structures.
[0074] Definition of noun:
[0075] The graph Transformer filter is a temporal model that uses an attention mechanism for feature weighting and information propagation on a graph structure, and is used for noise suppression and state estimation of anatomical hypergraphs.
[0076] The temporal attention is a weighting mechanism that calculates the correlation between the historical and current graph representations in the time dimension, and is used to guide information aggregation;
[0077] The cross-temporal fusion is a process of weighted combination of information between anatomical hypergraphs at adjacent time points based on temporal attention;
[0078] The attention input is a set of features used to calculate attention weights, including at least the node attributes, edge attributes, and hyperedge weights of the anatomical hypergraph.
[0079] The attention bias is a correction term added to the attention weight calculation, which is used to adjust the weight allocation according to the physical constraint attributes to reflect the imaging physics and guide priority.
[0080] The node embedding is a vectorized representation of the node after processing by the graph Transformer, which carries the temporal and topological fusion features of the node;
[0081] The edge embedding is a vectorized representation of the edge after processing by the graph Transformer, which carries the temporal and topological fusion features of the edge;
[0082] The state update is a process of iteratively refreshing the node embedding and edge embedding after cross-time fusion to reflect the information at the current moment.
[0083] The readout function is a function that aggregates the updated node embeddings and edge embeddings (and necessary hyperedge information) and maps them to a global belief state vector;
[0084] The belief state vector is a compact numerical description of the current imaging quality and geometric relationship under partially observable conditions, and is used as the basis for subsequent planning and quality control judgment.
[0085] The anatomical coverage rate is a measure of the coverage ratio of the target structure in the current imaging field of view, preferably expressed as a normalized value of the area of the target region as a percentage of the field of view area;
[0086] The angular deviation is a measure of the difference between the currently estimated anatomical axis or structural boundary and the guide target angle, preferably expressed as an absolute angular difference or its normalized value;
[0087] The symmetry index is a measure of the consistency of the left-right or top-bottom mirror image of the structure around a reference axis or center, and is preferably expressed as the reciprocal of the normalized similarity or difference.
[0088] The visibility confidence score of the key structure is a comprehensive confidence score of the visibility of the key anatomical structure to be determined in the current frame, which is derived from the confidence fusion of structure detection, segmentation or key point aggregation.
[0089] Optionally, step S5 specifically includes:
[0090] Using the belief state vector as input, a quality control cost function is constructed based on the anatomical coverage, angle deviation, symmetry index, and key structure visibility confidence in the belief state vector. The quality control cost function is composed of a weighted sum of the hyperedge satisfaction, the deviation term corresponding to the physical constraint edge attribute, the anatomical coverage improvement term, the angle deviation and symmetry reduction term, and the penalty term for insufficient visibility confidence.
[0091] Candidate action sequences are generated under equipment safety constraints and step size range constraints. The candidate action sequences consist of translation step size, tilt step size, and rotation step size of probe micro-movements, as well as gain step size, depth step size, focal position step size, and transmission frequency step size of equipment parameter adjustments.
[0092] The TD-MPC2 short-field planner performs rolling optimization on the candidate action sequences within a preset short field of view, evaluates the expected value of each candidate action sequence for the quality control cost function, selects the action that minimizes the quality control cost function, and outputs action suggestions.
[0093] Definition of noun:
[0094] The quality control cost function is a scalar objective function used to measure the difference between the current state and the standard view quality control target and to drive optimization. It is composed of multiple quality and physical indicators combined according to weights.
[0095] The hyperedge satisfaction is a quantitative indicator of the degree of satisfaction of the joint constraints of clinical guidelines in the current state, preferably normalized to [0,1] and participating in cost weighting;
[0096] The physical constraint deviation term is a penalty term that aggregates the deviations of the edge-level physical constraint attributes relative to the target or threshold, and is used to reflect the physical consistency of the incident angle, near / far field coverage and reflection intensity.
[0097] The anatomical coverage enhancement term is a term in the cost function used to encourage increased coverage of the target structure, preferably added as a negative mapping of coverage or its monotonic function form to reduce the cost;
[0098] The angle deviation and symmetry reduction terms are penalty combinations used in the cost function to reduce the angle difference of the anatomical axis and improve symmetry, and are preferably represented as a weighted sum of the two.
[0099] The penalty for insufficient visibility confidence is the amount of penalty applied when the visibility confidence of a critical structure is lower than a set threshold, preferably in the form of hinge loss or threshold truncation.
[0100] The equipment safety constraints are a set of constraints that limit candidate actions and parameter adjustments to ensure the safety of the equipment and the patient, including at least mechanical travel, output power / thermal load and manufacturer's safety limit.
[0101] The step size range constraint is a constraint that limits the upper and lower bounds of the probe pose increment and the step size of the device parameters, and is used to determine the feasible region of candidate actions.
[0102] The candidate action sequence is a set of combined actions arranged in chronological order within a preset short field of view, including the probe micro-movements and device parameter adjustments at each step;
[0103] The probe micro-motion is a small control increment relative to the current probe pose, including at least translation, tilt and rotation components;
[0104] The translation step size is the incremental position of the probe in the three axes of the imaging reference coordinate system;
[0105] The tilt step is the increment of the tilt angle of the probe relative to the reference plane in the pitch and roll directions.
[0106] The rotation step size is the increment of the probe's rotation angle around the acoustic beam axis;
[0107] The device parameter adjustment is an incremental setting of the imaging parameters of the ultrasound device, which affects the imaging quality and coverage;
[0108] The gain step size is the single adjustment range of the gain parameter;
[0109] The depth step size is the single adjustment range of the maximum imaging depth parameter;
[0110] The focal position step size is the single adjustment range of the electronic focusing focal depth position;
[0111] The transmission frequency step size is the single adjustment range of the transmission center frequency parameter;
[0112] The TD-MPC2 short-field planner is a planner that combines temporal difference and model predictive control to evaluate candidate actions and select the optimal action within a limited field of view.
[0113] The short field of view is a finite time step window used for prediction and optimization, with a length of a preset number of steps;
[0114] The rolling optimization is a recursive optimization strategy that uses a short field of view as a window, repeatedly predicts and optimizes at each time step, and only executes the optimal action of the current step.
[0115] The expected value is the statistical expectation or average evaluation value of the quality control cost function under the prediction model and uncertainty assumptions.
[0116] The action suggestion is a combined action to be performed at the current moment, selected through short field-of-view rolling optimization, including the specific values of probe micro-movements and device parameter step sizes.
[0117] Optionally, step S6 specifically includes:
[0118] Taking the motion suggestion as input, the motion suggestion is parsed into incremental targets of the translation step, tilt step and rotation step of the probe micro-motion relative to the current probe pose, and incremental targets of the gain step, depth step, focus position step and transmission frequency step of the device parameter adjustment relative to the current device parameters.
[0119] Under the constraints of equipment safety, mechanical travel and legal parameter range, the pose increment is executed through the probe pose control interface and the parameter increment is applied through the equipment control interface, or the operator is provided with direction and amplitude prompts through the human-machine interface and implemented by the operator. Unreachable increments are truncated according to the constraints and the actual execution after truncation shall prevail.
[0120] Under a unified reference clock, from the start of the action to the completion of the action, the sequence of ultrasonic video frames, the data of the inertial measurement unit, and the data of the equipment parameters are continuously acquired and output. Each frame of ultrasonic video, each set of data of the inertial measurement unit, and each sampling of equipment parameters are timestamped to characterize the actual observation after the action.
[0121] Definition of noun:
[0122] The probe pose is a set of parameters describing the spatial position and orientation of the probe in the imaging reference coordinate system, including at least three-dimensional translation and three-dimensional rotation;
[0123] The pose increment is a small change that needs to be performed relative to the current probe pose, including incremental components of translation, tilting and rotation around the acoustic beam axis;
[0124] The parameter increment is a single adjustment relative to the current device parameters, including increments in gain, depth, focus position, and transmission frequency;
[0125] The incremental target is the target value of the pose increment and parameter increment to be applied, obtained from the action suggestion parsing.
[0126] The mechanical travel constraint is a pose boundary constraint defined by the movable range of the probe's mechanical structure and mounting device, to prevent exceeding the physical travel or collision.
[0127] The legal range constraint of the parameters is the upper and lower limit of the imaging parameters specified by the manufacturer's specifications or safety regulations, to ensure that the parameters are set within the permissible range;
[0128] The probe pose control interface is a control and communication interface that issues pose increment commands to the probe or mechanical actuator and obtains feedback.
[0129] The device control interface is a control and communication interface that interacts with the ultrasound host or the whole machine software to apply parameter increments and read current parameters.
[0130] The human-computer interaction interface is an interface that presents and confirms the pose and parameter adjustment suggestions to the operator, and supports graphical / text / sound interaction;
[0131] The direction and amplitude prompts are clear guidance information on the direction of movement and the magnitude of adjustment on the human-computer interaction interface;
[0132] The unreachable increment is the pose or parameter increment that cannot be fully executed according to the target value due to equipment safety, mechanical travel, or legal parameter range constraints.
[0133] The truncation is a process of clipping unreachable increments according to the upper and lower bounds of the constraints and mapping them to the feasible region.
[0134] The actual execution refers to the pose and parameter adjustment results that are actually applied to the probe and device after constraint truncation.
[0135] The ultrasound video frame sequence after execution is a sequence of image frames acquired and time-stamped during the execution of the action, used to record the imaging results after execution;
[0136] The inertial measurement unit data after execution is IMU data collected and timestamped during the execution of the action, used to record the probe motion state during execution;
[0137] The executed device parameter data is a sequence of device parameter values collected and timestamped during the execution of the action, used to record the actual parameter adjustment results;
[0138] The actual observations are a collection of timestamped images, inertial data, and equipment parameter data collected during the execution of actions under a unified reference clock, used to objectively reflect the state after execution.
[0139] Optionally, step S7 specifically includes:
[0140] Using the ultrasound video frame sequence after execution, the inertial measurement unit data after execution, and the device parameter data after execution as inputs, repeat steps S1 to S4 to obtain and output the updated belief state vector;
[0141] The updated belief state vector is input into the quality control judgment module, which makes a judgment based on the preset threshold for the number of consecutive qualified frames and outputs the quality control judgment result.
[0142] The method terminates when the quality control judgment result meets the preset continuous qualified frame count threshold. When the quality control judgment result does not meet the preset continuous qualified frame count threshold, the updated belief state vector and the quality control judgment result are used as input constraints for short field of view scrolling optimization in step S5, and steps S5 to S7 are repeated until the preset continuous qualified frame count threshold is met.
[0143] The beneficial effects of this invention are:
[0144] 1. Improve robustness and accuracy under partially observable conditions: By using time synchronization and time-varying drift compensation of a unified reference clock, anatomical key point map and guide prior hyperedge, and physical constraint edge as attention bias graph Transformer temporal fusion, a stable belief state is output, which significantly reduces cross-modal alignment error and improves the reliability of anatomical coverage, angle deviation and key structure visibility determination.
[0145] 2. Improve the efficiency of standard view acquisition and achieve operable closed-loop quality control: Construct a quality control cost function based on the super-edge satisfaction and physical constraint deviation, and use TD-MPC2 short field of view rolling optimization to jointly plan probe micro-movements and equipment parameters under safety and step size constraints, shorten the time to reach the continuous qualified frame number threshold, and reduce repeated scanning and invalid adjustments;
[0146] 3. Balancing safety and human-machine collaboration: Action suggestions are constrained by the legal range of mechanical stroke, acoustics, and parameters. They can be executed automatically or implemented through human-machine collaboration with directional / amplitude prompts. Actions that cannot be achieved are automatically truncated and the actual execution is the standard, reducing risks and improving clinical consistency and reproducibility. Attached Figure Description
[0147] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0148] Figure 1 This is a flowchart of an ultrasound image quality control method based on a reinforcement learning algorithm proposed in this invention. Detailed Implementation
[0149] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0150] refer to Figure 1 An ultrasonic image quality control method based on reinforcement learning algorithm, characterized by comprising:
[0151] S1. Acquire ultrasonic video frame sequences, inertial measurement unit data, and equipment parameter data; perform time synchronization and drift correction, and output the data.
[0152] S2. Input the aligned ultrasound video frame sequence, inertial measurement unit data, and equipment parameter data into the anatomical key point detection and atlas construction module to construct an anatomical key point atlas;
[0153] S3. Construct an anatomical hypermap by embedding guide priors and physical constraints into the anatomical key point atlas;
[0154] S4. Perform temporal attention fusion and state update on the Transformer filter of the anatomical hypergraph input graph, and output the belief state vector;
[0155] S5. Input the belief state vector into the TD-MPC2 short-field planner, construct the quality control cost function and perform short-field rolling optimization, and output action suggestions;
[0156] S6. Perform the suggested actions on the ultrasonic equipment and collect observations of the execution process, outputting the ultrasonic video frame sequence, inertial measurement unit data, and equipment parameter data after execution;
[0157] S7. The executed ultrasonic video frame sequence, inertial measurement unit data, and equipment parameter data are processed sequentially through steps S1 to S4 to obtain the updated belief state vector, which is then input into the quality control judgment module. The process terminates when the quality control judgment result is qualified; otherwise, steps S5 to S7 are repeated.
[0158] In this specific embodiment, S1 specifically refers to:
[0159] To perform time synchronization and drift correction on multi-source data under a unified reference clock, firstly, using ultrasound video as the reference time axis, a timestamp is generated for each frame and recorded as follows. ,in Indicates the video number Frame timestamp, The video frame number;
[0160] Next, data from the inertial measurement unit and equipment parameters are collected, and timestamps are generated for each. and ,in Indicating the inertial measurement unit (IMU) timestamps of group samples Indicates the device parameter number The timestamp of the next sample For non-video modal samples, the inertial measurement unit must contain at least three-axis acceleration and three-axis angular velocity.
[0161] The equipment parameters should include at least the gain, depth, focal position, and transmission frequency;
[0162] To eliminate cross-modal time offset and time-varying drift, the non-video modal is referenced to the video timeline. Correction is performed, and the corrected timestamp is written. ,in For modality No. The corrected timestamp of each sample For the timestamp before correction, For modality constant time offset relative to the video timeline For modality The time-varying drift function, and The joint estimation is obtained by minimizing the cross-modal alignment error and combining it with the drift smoothing prior. Specifically, a joint optimization strategy based on the initial value search of cross-correlation peaks and piecewise linear drift fitting can be adopted. After obtaining the corrected timestamp, the non-video modalities are resampled onto the video reference time axis to achieve consistent sampling times. The resampling result is expressed as follows:
[0163] ;
[0164] in For video timestamps The obtained mode Resampled values For modality Interpolation and resampling operators (optional zero-order preserved, linear, or spline) For modality The original time series (corresponding to) The time series includes triaxial acceleration and triaxial angular velocity, corresponding to (Time series including gain, depth, focal position and transmission frequency) For modality Corrected timestamp sequence For video reference timestamps;
[0165] Lost and duplicate frames on the video side are corrected by timestamp consistency detection and local time axis fine-tuning to ensure the monotonicity of the reference time axis and sampling stability. After the above processing is completed, the video frame sequence with corrected timestamps, the aligned inertial measurement unit data, and the aligned device parameter data are output for subsequent steps.
[0166] In this specific embodiment, S2 specifically refers to:
[0167] Using aligned ultrasound video frames, inertial measurement unit, and device parameters as input, first in the... Keypoint detection is performed on a frame image using the following formula:
[0168] ;
[0169] in Indicates the first Frame image, For frame number, For key point detection functions, For the first The first frame Pixel coordinates of key points and The x and y coordinates are respectively. The visibility confidence level of this key point, For key point indexing, The total number of key points;
[0170] Subsequently, the relative pose code of the probe is obtained based on the aligned inertial data, written as:
[0171] ;
[0172] in For the first Pose encoding vector at frame reference time, For relative attitude quaternions, It is a relative translation vector;
[0173] Edges are generated based on spatial proximity, and the adjacency indicator is as follows:
[0174] ;
[0175] in For the first Frame key points and Adjacency indication, For key point indexing, For indicator functions, For Euclidean second norm, This is the spatial proximity threshold;
[0176] Mapping device parameters to map context attributes is represented as follows:
[0177] ;
[0178] in For the first Frame reference time context vector, For gain, For depth, For the focal position, For transmission frequency;
[0179] Based on this, an atlas of key anatomical points was constructed and denoted as follows:
[0180] ;
[0181] in For the first Frame graphs, For a set of nodes, For the set of edges, For node attribute collection, For node feature vectors, For the context of the graph;
[0182] And the pose encoding As for each edge The edge attributes are used to reflect the imaging geometry related to the probe pose. The final output is an anatomical key point map containing node, edge and context attributes, with each node and edge carrying the probe's relative pose code for subsequent steps.
[0183] In this specific embodiment, S3 specifically refers to:
[0184] Using the anatomical key point atlas as input, for the first While keeping the node and edge sets unchanged, the multi-structure joint constraints from clinical guidelines are written as hyperedges and an anatomical hypergraph is formed. This is done by using a subset of candidate key points. The above represents each super edge The satisfaction level is calculated and weighted according to the guideline threshold, where the hyperedge satisfaction level is obtained by a normalized weighted sum of geometric and visibility conditions after compression mapping, written as:
[0185] ;
[0186] in Indicates the first Frame-time superedge satisfaction To map real numbers to Interval compression functions (such as Sigmoid) Non-negative weights set by the importance of the guidelines For the compliance of the relative angle range (by (The angle deviation between the inner paired anatomical axes is obtained by normalization) For the compliance of relative distance range (by (obtained by normalizing the spacing deviation of internal key points) The compliance of the coverage threshold (obtained by normalizing the proportion of the target area within the imaging field of view);
[0187] The weight of each hyperedge is obtained by mapping the guide threshold, denoted as:
[0188] ;
[0189] in For super-edge weights, For monotonic mapping functions, To and The corresponding threshold vector (including thresholds for angle, distance, coverage, etc.);
[0190] When constructing physical constraint properties, for each edge Calculate the attribute vector related to imaging geometry, denoted as:
[0191] ;
[0192] in For the first Frame edge Physical property vectors For incident angle deviation, and Near-field and far-field coverage, respectively The expected reflection intensity;
[0193] The incident angle deviation is calculated from the angle between the local structure normal and the sound beam direction, and is written as:
[0194] ;
[0195] in For the key points Estimated unit normal vector The unit acoustic beam direction vector is obtained from the pose and imaging geometry;
[0196] Near-field and far-field coverage are obtained by piecewise integration of device context and depth window. Expected reflection intensity is calculated based on empirical mapping of device parameters and tissue incidence model. To maintain temporal consistency and imaging geometric interpretability, the probe relative pose encoding continues to be a common attribute of nodes and edges, where the pose encoding is denoted as... Device context is denoted as , respectively representing the first The pose quaternion of the frame, relative translation, as well as gain, depth, focal position, and emission frequency, are ultimately combined with the hyperedge set, hyperedge weights, and physical constraint properties to form an anatomical hypergraph, which is denoted as:
[0197] ;
[0198] in For the first Frame anatomical hypermap For a set of nodes, For the set of edges, For hyperedge sets, For the set of superedge weights, It is a set of physical constraint attributes.
[0199] In this specific embodiment, S4 specifically refers to:
[0200] Using the anatomical hypergraph sequence as input, at the current frame index and length is Within the temporal window, a graph Transformer with physical constraint bias is used for cross-temporal fusion and state update. Specifically, the historical embeddings of nodes and edges are first aggregated temporally within the window to obtain the initial representation of the current frame, and attention weights on the adjacency topology are calculated accordingly to fuse multi-source cues. The core attention includes a physical bias term and is written as:
[0201] ;
[0202] in Indicates the current frame Time from node To the node Attention weights and For node indexing, For nodes Neighborhood set For query vectors, For key vectors, and For the embedding of nodes after time-series aggregation, and For learnable linear projection matrices, For hidden space dimension, The Softmax function is normalized over the neighbor set. This refers to attention bias.
[0203] The bias term is obtained by mapping the physical constraint properties to the hyperedge weights of the guide prior and is written as:
[0204] ;
[0205] in For learnable parameter vectors, For the edge At any moment Biased eigenvectors For incident angle deviation, and Near-field and far-field coverage, respectively For the expected reflection intensity, This represents the average weight of the super-edges connected to this edge;
[0206] After performing cross-time and cross-topology information fusion based on the above attention, the nodes and edges are updated, and the updated node embeddings are denoted as follows. To enhance temporal consistency, the readout function ultimately performs global aggregation of nodes, edges, and hyperedges and outputs a belief state vector. The readout is written as follows:
[0207] ;
[0208] in For belief state vector, For anatomical coverage, For angle deviation, For symmetry indicators, For the visibility confidence of key structures, For learnable readout functions, For global aggregation operators, and They are time points The set of nodes and edges For a moment side Physical constraint attribute vector, For a moment hyperedge set, This is the set of superedge weights.
[0209] In this specific embodiment, S5 specifically includes:
[0210] Using belief state as input and employing TD-MPC2 for rolling optimization under a preset short field of view to jointly plan probe micro-movements and device parameter adjustments, the current time is denoted as... The short field of vision length is denoted as In time index Define candidate action vectors above:
[0211] ;
[0212] in For the first Combined movements of steps The three-axis increment of the probe translation step size, The increment of the probe tilt step on the two vertical axes, For the rotational step increment around the sound beam axis, For gain step increment, For depth step increment, For the step size increment at the focal position, This represents the step size increment of the transmission frequency;
[0213] To meet equipment safety and step size range constraints, the feasible region of motion is denoted as:
[0214] ;
[0215] in For actionable assembly, For pose increment subvectors, For parameter increment subvectors, For infinite norm, The maximum allowable range of pose increment, This represents the maximum allowable increment of the parameter.
[0216] In constructing quality control objectives, the key components of belief states are used as the core of the cost, and aggregation terms of hyperedges and physical constraints, as well as action regularization terms, are introduced. The stage cost is defined as follows:
[0217] ;
[0218] in For the first The stage cost function of the step For the first The belief state vector of the step, The weighted average of the satisfaction of the superedges. The aggregate value of physical constraint deviation, For anatomical coverage, For angle deviation, For symmetry indicators, For the visibility confidence of key structures, For visibility qualification threshold, For hinge functions, For each non-negative weight, For weighted quadratic action regularization, It is a symmetric positive semidefinite weight matrix;
[0219] To achieve rolling prediction, the short-field evolution of belief states is denoted as:
[0220] ;
[0221] in To develop a predictive model that characterizes the impact of actions on belief states within a short field of view;
[0222] Based on the above definitions and using the TD-MPC2 optimization criteria to select the action sequence, it can be written as:
[0223] ;
[0224] in For the optimal short visual field action sequence, In the prediction model Expectations The terminal cost function is used to emphasize the endpoint objective of achieving a quality control qualified state;
[0225] The first action of the sequence will be finalized. The action suggestions are output and parsed as step increments for probe translation, tilting, and rotation, as well as parameter step increments for gain, depth, focus position, and transmission frequency, for the next step of execution.
[0226] In this specific embodiment, S6 specifically refers to:
[0227] The motion suggestions are executed on the ultrasound equipment, and real-time observations of the execution process are collected. First, at the current reference time, the motion suggestions are resolved into incremental targets relative to the existing probe pose and equipment parameters. The joint suggestion vector is denoted as:
[0228] ;
[0229] in Recommendations for joint actions to be implemented; The translation step size increment of the probe in three axes, The tilt step increment of the probe on the two tilt axes, For the rotational step increment around the sound beam axis, For gain step increment, For depth step increment, For the step size increment at the focal position, This represents the step size increment of the transmission frequency;
[0230] To satisfy equipment safety constraints, mechanical travel constraints, and parameter legal range constraints, the pose and parameter increments are projected onto the movable domain, as follows:
[0231] ;
[0232] in For the constrained executable pose increment vector, The parameter increment vector that is executable after constraints For pose increment subvectors, For parameter increment subvectors, and For component-wise projection operator, The pose feasible region defined by the upper / lower limits of safety and mechanical travel. For the parameter feasible region defined by the parameter valid range, These are the lower / upper bound vectors of the pose increment, respectively. These are the lower and upper bound vectors for the parameter increment, respectively;
[0233] Unreachable increments are automatically truncated by the above projection and... and For accurate execution, the execution method can be automatically applied through the probe pose control interface and the equipment parameter control interface, or the operator can be provided with direction and amplitude prompts through the human-machine interface. Under a unified reference clock, real observations are continuously collected from the start to the completion of the action, and each observation is timestamped. The collection sequence is recorded as follows:
[0234] and ;
[0235] in For the set of reference timestamps during the execution phase, For the first Reference timestamp of the second sampling Number of samples during the execution phase For the first time after execution Frame ultrasound images, For the timestamp of this frame, For the first time after execution Group inertial measurement unit data, Its timestamp, For the first time after execution Secondary equipment parameter sampling, Its timestamp;
[0236] The above-mentioned timestamped post-execution video, inertial, and parameter data are used as real observation outputs for time alignment and state updates in subsequent steps.
[0237] In this specific embodiment, S7 specifically refers to:
[0238] Using post-execution observations as input, the process follows steps S1 to S4 to sequentially complete time synchronization and drift correction, anatomical keypoint map construction, guide prior hyperedge and physical constraint injection, and graph Transformer temporal filtering, thereby obtaining the updated belief state vector, denoted as:
[0239] ;
[0240] in Indicates the current frame index is Belief status after the update For anatomical coverage, For angle deviation, For symmetry indicators, For the visibility confidence of key structures, To map post-execution observations sequentially through steps S1 to S4 into belief states using a composite operator, For the ultrasound video frame sequence after execution, For the inertial measurement unit data sequence after execution, This is the sequence of device parameter data after execution;
[0241] The belief state is then fed into the quality control judgment module to obtain the single-frame quality score and whether it meets the standard, written as:
[0242] ;
[0243] in The quality score of the current frame. To map belief states to a decision function that normalizes quality scores, As an indicator variable for whether or not it is qualified, For indicator functions, The threshold for a single frame to pass;
[0244] To characterize the number of consecutive qualified frames, a counter is defined for recursion:
[0245] And at the initial time let ;
[0246] in For the continuous valid frame count up to the current frame, This is a count from the previous frame;
[0247] when The time counter naturally resets to 0 when Reaching the preset threshold The method is terminated if the continuous qualified frame count requirement is met; otherwise, it will... and As an input constraint for step S5 short-field-of-view scrolling optimization to strengthen guidance on terminal qualification, the constraint can be written as... ,in To achieve short field of view The following is a prediction model The terminal belief state obtained through evolution For prediction models The uncertainty expectation is addressed by repeating steps S5 to S7 until the target is not met, through the aforementioned closed loop. until.
[0248] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0249] This invention addresses three key technical challenges by linking time synchronization and drift correction under a unified reference clock, anatomical keypoint atlases carrying pose encoding and device context, anatomical hypergraphs embedding clinical guideline priors, graph Transformer temporal fusion with physical constraints as attention biases, and TD-MPC2 short-field-of-view scrolling optimization and quality control judgment centered on belief states into a closed loop. Each submodule specifically alleviates these challenges: difficulty in modeling some observable states, instability in multi-source data fusion, and lack of operable closed-loop control. Specifically, cross-modal time alignment and time-varying drift compensation significantly reduce video, inertial measurement, and device drift. Alignment errors between parameters, structured keypoint maps and hyperedges transform the geometric and visibility conditions in the guidelines into computable constraints. Combined with the physical properties of incident angle, near / far field coverage and expected reflection intensity, temporal attention can still stably output belief state indicators such as coverage, angle deviation, symmetry and visibility under occlusion, jitter and tissue movement. Finally, short field-of-view rolling optimization driven by quality control cost function jointly plans probe micro-movements and parameter adjustments within safety and step size constraints, shortens the time to reach the continuous qualified frame rate threshold, reduces repeated scanning and manual parameter adjustment, and improves the stability and reproducibility of standard view acquisition.
[0250] In terms of algorithm structure, this case has made multiple improvements to address the technical issues:
[0251] On the one hand, it encodes the multi-structure joint constraints of clinical guidelines in the form of hyperedges and introduces learnable weights to close the link of "guideline prior - graph structure - attention bias - belief state", making up for the inadequacy of traditional feature splicing in expressing complex geometry and visibility relationships.
[0252] On the other hand, acoustic properties such as incident angle deviation and field coverage are explicitly injected into the edge-level physical constraints as bias terms for attention, thereby improving the interpretability and robustness of cross-time and cross-topology fusion.
[0253] Simultaneously, the coverage, angle deviation, symmetry, and visibility in the belief state are directly mapped to the quality control cost function. TD-MPC2 is used to jointly optimize pose and parameters within a short field of view. Combined with the constraint projection of mechanical travel, equipment safety, and the legal range of parameters, as well as the human-machine collaborative execution mechanism, an end-to-end closed-loop structure of "belief-cost-action" is formed. This makes it more resistant to occlusion, drift, and operator tremors in real clinical scenarios and enables the quality control target to be achieved faster and more stably within an operable and auditable safety boundary.
Claims
1. An ultrasound view quality control method based on a reinforcement learning algorithm, characterized in that, Comprise: S1, collect ultrasound video frame sequence, inertial measurement unit data and device parameter data, time synchronization and drift correction and output; S2, input the aligned ultrasound video frame sequence, inertial measurement unit data, device parameter data into the anatomical key point detection and atlas construction module to construct the anatomical key point atlas; S3, the anatomical key point atlas is embedded with guide priori and constructed with physical constraint, and the anatomical supergraph is output; S4, input the anatomical supergraph into the graph Transformer filter for time attention fusion and state update, and output the belief state vector; S5, input the belief state vector into TD-MPC2 short field planner, construct quality control cost function and perform short field rolling optimization, and output action suggestion; S6, execute the action suggestion on the ultrasound device and collect the observation of the execution process, output the ultrasound video frame sequence, inertial measurement unit data and device parameter data after execution; S7, the executed ultrasound video frame sequence, inertial measurement unit data and device parameter data are sequentially processed by steps S1 to S4 to obtain the updated belief state vector and input the quality control judgment module, when the quality control judgment result is qualified, terminate, otherwise, repeat steps S5 to S7.
2. The method of claim 1, wherein the method is based on a reinforcement learning algorithm. S1 is specifically: Under the unified reference clock, time stamps are generated for each frame of ultrasound video frame, each group of inertial measurement unit data and each device parameter sampling; under the time stamp mark, the ultrasound video frame sequence is continuously collected and output; Under the time stamp mark, the inertial measurement unit data is continuously collected and output, and the inertial measurement unit data at least includes three-axis acceleration and three-axis angular velocity; Under the time stamp mark, the device parameter data is read and output in real time through the device interface, and the device parameter data at least includes gain, depth, focus position and transmission frequency; Based on the reference time axis, time offset estimation and time-varying drift compensation are performed on the inertial measurement unit data and the device parameter data, the time offset estimation and the time-varying drift compensation are corrected through minimizing the cross-modal alignment error on the reference time axis to obtain the corrected time stamp; Under the corrected time stamp, the inertial measurement unit data and the device parameter data are interpolated and resampled to make their sampling time consistent with the reference time axis, and the time stamp of the ultrasound video frame sequence is corrected and necessary interpolation processing is performed on the lost frame and the repeated frame; After completing the time synchronization and drift correction, the aligned ultrasound video frame sequence, the aligned inertial measurement unit data and the aligned device parameter data are output.
3. The method of claim 1, wherein the method further comprises: S2 is specifically: Take the aligned ultrasound video frame sequence as the image input, detect the anatomical key points in each frame and generate the coordinates and visibility confidence of each anatomical key point; The probe relative pose encoding is calculated based on the aligned inertial measurement unit data, and the relative pose encoding is taken as the node attribute and edge attribute of the anatomical key point atlas; According to the spatial relationship between the anatomical key points, edges of the anatomical key point atlas are generated, which are used to represent the spatial adjacency relationship of adjacent key points; Map the aligned device parameter data as atlas context attributes, the atlas context attributes at least including current values of gain, depth, focal position and emission frequency; Output an anatomical keypoint atlas, the anatomical keypoint atlas including nodes, edges and atlas context attributes, wherein the nodes contain coordinates and visibility confidence of anatomical keypoints, and the edges contain spatial adjacency relationship of adjacent keypoints, and the nodes and edges carry probe relative pose encoding.
4. The method of claim 1, wherein the method further comprises: S3 specifically is: Take the anatomical keypoint atlas as input, and write the multi-structure joint constraints of standard views into the anatomical keypoint atlas in the form of hyperedges according to clinical guidelines, the hyperedges are used to represent the geometric and visibility conditions that are simultaneously satisfied by at least one group of anatomical structures, and the geometric and visibility conditions include at least one of symmetry, collinearity, relative angle range, relative distance range or coverage threshold; Determine the weight of each hyperedge according to the threshold of the clinical guidelines; Construct physical constraint attributes on the edges of the atlas, the physical constraint attributes include at least one of incident angle deviation, near-field coverage, far-field coverage or expected reflection intensity, and the physical constraint attributes are calculated from the probe relative pose encoding, atlas context attributes and anatomical keypoint coordinates; Keep the probe relative pose encoding as node attributes and edge attributes; After completing the prior embedding of the guidelines and the construction of the physical constraints, output the anatomical hypergraph.
5. The method of claim 1, wherein the method further comprises: S4 specifically is: Take the anatomical hypergraph as the current time input, and perform cross-time fusion through the temporal attention of the graph Transformer filter under the sequence of anatomical hypergraphs at adjacent times, the cross-time fusion takes the node attributes, edge attributes and hyperedge weights of the anatomical hypergraph as attention input, and takes the physical constraint attributes as attention bias to control the weight distribution; Update the node embedding and edge embedding after completing the cross-time fusion; Calculate and output the belief state vector based on the updated node embedding and edge embedding through the readout function, the belief state vector at least includes anatomical coverage, angle deviation, symmetry index and key structure visibility confidence.
6. The method of claim 1, wherein the method further comprises: S5 specifically is: Take the belief state vector as input, and construct a quality control cost function according to the anatomical coverage, angle deviation, symmetry index and key structure visibility confidence in the belief state vector, the quality control cost function is composed of hyperedge satisfaction degree, deviation term corresponding to physical constraint edge attribute, anatomical coverage improvement term, angle deviation and symmetry reduction term, and penalty term for insufficient visibility confidence; Generate a candidate action sequence under the constraints of device safety and step range, the candidate action sequence is composed of translation step, tilt step and rotation step of probe micro-motion, and gain step, depth step, focal position step and emission frequency step of device parameter adjustment; Rolling optimize the candidate action sequence in the preset short view field through the TD-MPC2 short view field planner, evaluate the expected value of each candidate action sequence on the quality control cost function, and select the action that minimizes the quality control cost function, and output the action suggestion.
7. The method of claim 1, wherein the method further comprises: S6 specifically is: taking the action suggestion as input, parsing the action suggestion into an incremental target of a translation step, a tilt step and a rotation step of a probe micro-motion relative to a current probe pose, and an incremental target of a gain step, a depth step, a focal position step and a transmit frequency step of a device parameter relative to a current device parameter; under the constraints of device safety, mechanical travel and parameter legal range, executing the pose increment through a probe pose control interface and applying the parameter increment through a device control interface, or providing a direction and amplitude prompt to an operator through a human-machine interaction interface and implementing by the operator, the unattainable increment being truncated according to the constraints and the actual execution after truncation being used as a reference; under a unified reference clock, continuously collecting and outputting an executed ultrasound video frame sequence, executed inertial measurement unit data and executed device parameter data during the execution of the action from the beginning to the end, wherein each ultrasound video frame, each set of inertial measurement unit data and each device parameter sampling is time-stamped to represent a real observation after execution.
8. The method of claim 1, wherein the method is based on a reinforcement learning algorithm. S7 specifically comprises: taking the executed ultrasound video frame sequence, the executed inertial measurement unit data and the executed device parameter data as input, repeating the steps of S1 to S4 to obtain and output an updated belief state vector; taking the updated belief state vector as input into a quality control judgment module, judging according to a preset continuous qualified frame number threshold and outputting a quality control judgment result; when the quality control judgment result meets the preset continuous qualified frame number threshold, terminating the method, and when the quality control judgment result does not meet the preset continuous qualified frame number threshold, taking the updated belief state vector and the quality control judgment result as input constraints of the short field of view rolling optimization of step S5, repeating steps S5 to S7 until the preset continuous qualified frame number threshold is met.