Thyroid ultrasonic robot automatic scanning method, device and equipment based on RGB image and depth information and medium
By combining RGB images and depth information encoding processing, the scanning probe posture and path are adjusted in real time, the problem of unstable thyroid scanning imaging quality in the prior art is solved, and high-quality and standardized automatic ultrasound imaging effect is achieved.
Patent Information
- Application Number
- CN202510679551.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing robot ultrasound scanning system lacks the ability to dynamically adjust the posture and path of the scanning probe in thyroid scans, resulting in unstable imaging quality and low scanning standardization, making it difficult to adapt to the characteristics of hidden thyroid position and subtle structure, and lacks the utilization of semantic relationships in different tissues, which reduces the universality and practicality of the system.
By obtaining RGB image and depth information of the target area, the encoding process is performed and the target scanning area is identified. The scanning probe is controlled in real time to analyze the scan image to recognize the preset target and artifact area, dynamically adjust the scanning attitude and path, and terminate the scanning when the preset target is not recognized.
It realizes accurate positioning, stable image quality, and intelligent optimization of scanning paths in thyroid scans, improves the effect and standardization of automatic ultrasound imaging, and enhances the adaptability and robustness of the system.
Smart Images

Figure CN120580274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a thyroid ultrasound robot automatic scanning method, device, equipment and storage medium based on RGB images and depth information. Background Art
[0002] Thyroid nodules are a relatively common disease in clinical practice. According to epidemiological surveys, the prevalence of thyroid nodules is high in the general population, and some studies have shown that it can reach nearly half. Early and accurate identification of malignant nodules is of great significance for disease management and treatment effectiveness. Ultrasound examination, as a non-invasive, radiation-free, economical and convenient imaging method, has become an important tool for clinical detection of thyroid nodules, especially in detecting tiny cancerous changes. Despite this, the current process of acquiring ultrasound images still heavily relies on the operator's manual control of the probe, resulting in the scanning process being limited by the operator's experience level and differences in technique, resulting in problems such as unstable image quality, low standardization, and poor repeatability. In addition, long-term and continuous manual operation of the ultrasound probe can easily cause musculoskeletal strain and increase occupational health risks.
[0003] With the development of robotic technology, robotic ultrasound scanning systems (RUS) have been proposed in an effort to reduce the burden on operators and improve scanning standardization. Based on the degree of human-computer interaction, existing technologies categorize robotic ultrasound systems into manual, semi-automatic, and fully automatic types. The manual method is primarily used for remote probe operation and is suitable for telemedicine scenarios, but still requires real-time human intervention. The semi-automatic method typically relies on the operator to pre-mark the region of interest or design the scanning path, with the robot performing the scanning action. The fully automatic method aims to autonomously complete probe positioning, path planning, and image acquisition through the robot. However, existing robotic ultrasound scanning systems have significant technical limitations when it comes to the specific anatomical structure of the thyroid gland.
[0004] Some existing automatic path planning methods based on RGB-D imagery have shown problems in application, such as large path coverage, low scanning efficiency, and poor local imaging accuracy. These methods are difficult to adapt to the hidden location and subtle structure of the thyroid gland. Existing studies often estimate paths based on surface brightness changes or contours, lacking accurate modeling of deep anatomical locations. This leads to large deviations in scanning positioning and an inability to effectively guarantee the acquisition of high-quality images. In addition, although deep learning methods have been introduced in recent years for tissue recognition and segmentation in images, most technologies still rely on the operator to provide the initial position or require manual preprocessing, making it difficult to truly achieve full-process autonomous scanning.
[0005] During the scanning process, existing systems often lack the ability to dynamically adjust the probe posture and scanning path based on real-time image recognition results. This results in an inability to promptly correct the scanning trajectory or posture even when artifact areas or scan offsets are identified, leading to a significant amount of invalid information or occlusion in the acquired images. This is particularly true for imaging-intensive areas like the thyroid gland, where the lack of an effective mechanism for real-time optimization of scan control based on image content severely impacts the effectiveness of robotic ultrasound systems.
[0006] Furthermore, existing technologies for the thyroid gland and its surrounding anatomical structures generally lack the ability to fully exploit the spatial semantic relationships between different tissues. This failure to fully incorporate existing spatial structural knowledge into target positioning, path adjustment, and control strategies results in insufficient scanning intelligence and adaptability in complex scenarios. This issue often prevents existing methods from flexibly adjusting to varying patient body shapes and neck postures, reducing the system's universality and practicality. Summary of the Invention
[0007] The main purpose of the present invention is to provide a thyroid ultrasound robot automatic scanning method, device, equipment and storage medium based on RGB images and depth information, aiming to solve the technical problem that the existing technology cannot dynamically adjust the scanning posture and scanning path of the scanning probe based on real-time image recognition results during the scanning process, resulting in unstable imaging quality and low scanning standardization.
[0008] To achieve the above objectives, the present invention provides a method for automatic thyroid ultrasound robot scanning based on RGB images and depth information, comprising:
[0009] Obtain image information and depth information of the target area;
[0010] Encoding the depth information to obtain encoded depth information;
[0011] identifying a target scanning area within the target area based on the image information and the encoded depth information;
[0012] Determining an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0013] Controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process;
[0014] Analyzing the real-time scanning image to identify preset targets and artifact areas in the real-time scanning image;
[0015] Adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0016] Monitoring the continuous presence of a preset target in the real-time scanning image;
[0017] If the preset target is not recognized in the continuous multiple frames of images, the scanning is terminated.
[0018] Furthermore, to achieve the above-mentioned purpose, the present invention provides a scanning control device based on image analysis, comprising:
[0019] Image and depth information acquisition module, used to obtain image information and depth information of the target area;
[0020] A depth information encoding module, configured to encode the depth information to obtain encoded depth information;
[0021] A target scanning area identification module, configured to identify a target scanning area within the target area based on the image information and the encoded depth information;
[0022] An initial scanning planning module, configured to determine an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0023] A scanning execution and image acquisition module, configured to control the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and to acquire a real-time scanning image during the scanning process;
[0024] A real-time image analysis module, configured to analyze the real-time scanned image and identify preset targets and artifact areas in the real-time scanned image;
[0025] A scanning posture and path adjustment module, configured to adjust the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0026] A preset target monitoring module, used to monitor the continuous existence state of the preset target in the real-time scanning image;
[0027] The scanning termination control module is used to terminate the scanning if the preset target is not recognized in the continuous multiple frames of images.
[0028] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image analysis-based scanning control program stored in the memory and runnable on the processor. When the image analysis-based scanning control program is executed by the processor, the steps of the thyroid ultrasound robot automatic scanning method based on RGB images and depth information as described above are implemented.
[0029] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a scanning control program based on image analysis is stored. When the scanning control program based on image analysis is executed by a processor, the steps of the automatic scanning method of the thyroid ultrasound robot based on RGB image and depth information as described above are implemented.
[0030] Beneficial effects: The present invention relates to the field of computer vision technology, and discloses a thyroid ultrasound robot automatic scanning method, device, equipment and medium based on RGB images and depth information, including: acquiring image information and depth information of a target area, encoding the depth information to generate encoded depth information, identifying a target scanning area within the target area based on the image information and the encoded depth information, determining an initial scanning point and an initial scanning direction according to the position of the target scanning area, controlling a scanning probe to scan and collect real-time scanning images based on the initial scanning point and the initial scanning direction, analyzing the real-time scanning image to identify preset targets and artifact areas, adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset targets and artifact areas, monitoring the continuous existence status of preset targets in the real-time scanning image, and terminating the scanning when the preset target is not recognized in multiple consecutive frames of images. The present invention accurately identifies the target scanning area by fusing image information and coded depth information, controls the scanning probe to carry out standardized scanning in combination with the initial scanning point and scanning direction, and dynamically adjusts the scanning posture and scanning path based on the real-time image recognition results during the scanning process. At the same time, the scanning termination timing is determined according to the continuous detection status of the preset target, thereby achieving accurate positioning, stable image quality, and intelligent optimization of the scanning path during the ultrasonic scanning process, thereby improving the automatic ultrasonic imaging effect and standardization of the thyroid gland and similar superficial organs. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0032] Figure 1 A schematic diagram of an application environment of a method for automatic thyroid ultrasound robot scanning based on RGB images and depth information in one embodiment of the present invention;
[0033] Figure 2 This is a flow chart of an embodiment of a method for automatic thyroid ultrasound robot scanning based on RGB images and depth information according to the present invention;
[0034] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a scanning control device based on image analysis of the present invention;
[0035] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0036] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0038] The thyroid ultrasound robot automatic scanning method based on RGB image and depth information provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain image information and depth information of the target area through the user terminal, encode the depth information to generate encoded depth information, identify the target scanning area within the target area based on the image information and the encoded depth information, determine the initial scanning point and the initial scanning direction according to the position of the target scanning area, control the scanning probe to scan according to the initial scanning point and the initial scanning direction and collect real-time scanning images, analyze the real-time scanning images to identify preset targets and artifact areas, adjust the scanning posture and scanning path of the scanning probe according to the recognition results of the preset targets and artifact areas, monitor the continuous presence of the preset targets in the real-time scanning images, and terminate the scanning when the preset targets are not recognized in multiple consecutive frames of images. The present invention accurately identifies the target scanning area by fusing image information and encoded depth information, controls the scanning probe to carry out standardized scanning in combination with the initial scanning point and scanning direction, and dynamically adjusts the scanning posture and scanning path based on the real-time image recognition results during the scanning process. At the same time, the scanning termination time is determined according to the continuous detection status of the preset targets, thereby achieving accurate positioning, stable image quality, and intelligent optimization of the scanning path during the ultrasonic scanning process, thereby improving the automatic ultrasonic imaging effect and standardization of the thyroid gland and similar superficial organs. The user end may be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server end may be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below using specific embodiments.
[0039] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of a method for automated thyroid ultrasound scanning based on RGB images and depth information provided by the present invention. It should be noted that although a logical sequence is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0040] like Figure 2 As shown, the thyroid ultrasound robot automatic scanning method based on RGB image and depth information proposed in the present invention includes the following steps:
[0041] S10, acquiring image information and depth information of the target area;
[0042] In this embodiment, to achieve automated ultrasonic scanning control, it is first necessary to accurately obtain image and depth information of the target area. Image information generally refers to data reflecting the target area's surface texture, color distribution, structural outline, and other content. This information can be acquired using a standard two-dimensional imaging sensor, such as an RGB camera. To achieve higher spatial resolution, a high-pixel-density imaging sensor, such as a CMOS or CCD sensor, can be selected, while ensuring that the acquisition frame rate meets the real-time requirements of subsequent processing.
[0043] Depth information, typically expressed as point-to-surface distances, is the distance between each point on the target surface and the sensor's reference coordinate system. It can be acquired using a structured light depth camera, a time-of-flight (ToF) camera, or a binocular vision system. Depth information acquisition requires robust interference resistance under varying ambient light conditions. Therefore, when selecting hardware, a depth sensor with active infrared projection is recommended to mitigate the effects of external light interference on measurement accuracy.
[0044] To ensure the spatial correspondence between image and depth information, the imaging system is typically pre-calibrated. This involves establishing a one-to-one mapping between image coordinates and depth coordinates using an orthogonal calibration plate or spatial geometric calibration structure. Calibration data includes both internal parameters (such as focal length, principal point position, and distortion coefficients) and external parameters (such as rotation matrices and translation vectors), enabling precise alignment of the two types of data in the spatial domain.
[0045] During the acquisition process, the target area must remain within the sensor's effective imaging field of view. This can be achieved by adjusting the sensor's mounting angle, moving the robotic arm's initial position, or resizing the imaging window to ensure that the acquisition process covers the entire target area while avoiding data distortion caused by sensor edge effects. Holes in the depth information caused by varying material reflectivity can be repaired during acquisition through multi-angle fusion or by applying a neighborhood interpolation algorithm in post-processing to improve data integrity.
[0046] Through the above means, a set of spatially corresponding, content-complementary, and fully covered image information and depth information can be obtained, providing the necessary data basis for subsequent target scanning area identification, scanning path planning, and dynamic control adjustment.
[0047] In practical implementation, simultaneous image acquisition can be performed using a color depth camera that integrates an RGB imaging module and a depth measurement module. Imaging devices equipped with structured light or time-of-flight ranging capabilities are preferred, such as those using active infrared projection and high dynamic range processing technology to improve imaging consistency in complex lighting environments. The device is mounted on the end of a robotic arm or on a standalone bracket. During acquisition, the camera's posture and acquisition angle must be adjusted based on the size and position of the preset target area to ensure that the target subject is contained in the center of the image.
[0048] During the acquisition process, image and depth information are synchronized using a unified timestamp. Holes in the depth information are interpolated and filled using a local neighborhood least-squares fit method to form a complete and continuous depth map. To improve overall system responsiveness, image information can be spatially downsampled using a pyramid, preserving edge detail at all scales.
[0049] To accommodate the varying surface contours caused by varying patient body shapes, the depth sensor can be configured with multiple ranging modes, automatically switching between precise close-range mode and coarse long-range mode. For highly reflective skin surfaces or hair interference, an anti-scattering compensation module can be introduced at the hardware level, or a filtering strategy based on depth map gradient sparsity constraints can be employed during software processing to further optimize data quality.
[0050] The above acquisition methods can be adjusted according to the actual application environment. For example, in a strong indoor light environment, infrared active imaging devices with ambient light filtering capabilities are preferred; in scenarios requiring high spatial resolution, the frame rate can be appropriately reduced in exchange for higher pixel accuracy; in scenarios with moving targets, the acquisition module can add short-time exposure control to reduce image blur.
[0051] Example description: In the application of automatic thyroid ultrasound scanning, in order to ensure complete coverage of the target area, a camera equipped with an active infrared depth measurement module and a high-resolution RGB imaging module can be used to synchronously capture images and depth of the patient's neck. Before acquisition, the imaging angle is set so that the thyroid area is located in the center of the image. The internal parameters of the imaging device are calibrated based on a checkerboard calibration plate, and the integrity of the depth map data is monitored in real time. The detected void areas are repaired using neighborhood least squares smoothing interpolation. In this way, even under different patient body shapes, different neck curvatures, and different skin reflectivity conditions, stable and highly consistent image information and depth information can still be obtained, providing reliable support for subsequent thyroid identification, scanning path planning, and real-time adjustment, thereby effectively improving the standardization of thyroid ultrasound examinations and image acquisition quality.
[0052] This embodiment can provide a complete and accurate spatial data foundation before the start of ultrasound scanning by synchronously acquiring image information and depth information of the target area, and performing calibration alignment and data repair processing. Image information provides rich texture and structural information for subsequent target feature recognition, and depth information provides a spatial geometric basis for path planning and probe posture adjustment. The organic combination of the two effectively improves the accuracy and stability of subsequent automatic ultrasound scanning. The use of high-precision acquisition equipment and adaptive processing strategies can stably obtain high-quality input data under different lighting, different surface conditions, and different scene requirements, thereby enhancing the adaptability and robustness of the entire system and improving the consistency and imaging quality of automatic ultrasound inspections.
[0053] S20, encoding the depth information to obtain encoded depth information;
[0054] In this embodiment, to enhance depth information features and extract structural details, the acquired depth information is encoded to generate encoded depth information with multi-dimensional spatial representation capabilities. Depth information encoding involves converting the original single-channel depth data into a composite representation that reflects the surface morphology, spatial relationships, and orientation attributes of the target area through spatial geometric feature extraction and multi-channel mapping.
[0055] The first step in the encoding process is to calculate the horizontal disparity channel based on the disparity model of the depth sensor. Horizontal disparity is a measure of the displacement of targets at different depths of field in the camera baseline direction. It can enhance the lateral change characteristics of the target due to depth differences on the image plane, and help distinguish the structural contours at different depth levels.
[0056] The second step of the encoding process is to calculate the ground height channel based on the gravity direction reference axis. The ground height reflects the vertical height information of each point relative to the preset horizontal plane. By introducing the gravity direction vector or the attitude reference provided by the inertial measurement unit (IMU), vertical height features that conform to the physical properties of the real world can be stably extracted. This is especially suitable for determining the height changes of superficial organs such as the thyroid gland and surrounding structures.
[0057] The third step in the encoding process is to generate a normal angle channel based on the surface normal estimation algorithm. The normal angle describes the degree of inclination of the local surface relative to the reference direction. The surface normal can be derived through plane fitting based on the local point set or triangular mesh reconstruction. The angle is then calculated with the set reference direction (such as gravity upwards) to form a feature channel with direction perception capabilities. The multi-channel merging of the horizontal disparity channel, the ground height channel, and the normal angle channel can fuse three types of key spatial information: lateral displacement, vertical height, and local orientation within a single data structure to form the encoded depth information, providing a richer and more structured feature foundation for subsequent target area recognition.
[0058] In a specific implementation, the horizontal disparity value can be first inferred based on the raw pixel values of the depth map and sensor parameters. For example, a back-projection model is used to jointly map the pixel coordinates and depth values into a three-dimensional spatial coordinate system. The projection difference between adjacent points in the baseline direction is then calculated to obtain the horizontal disparity channel. The ground height can be calculated by reprojecting each point in the depth map, calculating the height coordinate of the point in the direction of gravity, and normalizing it relative to a uniformly set reference horizontal plane to form a ground height channel. The surface normal vector can be estimated by selecting a local fixed neighborhood window, such as a three-by-three or five-by-five pixel neighborhood, performing least squares plane fitting to obtain the normal vector, and then calculating the angle between the normal vector and the gravity direction to form a normal angle channel. During the multi-channel merging process, the horizontal disparity, ground height, and normal angle can be sequentially spliced as three independent channels into a three-dimensional tensor. At the same time, normalization is performed to standardize the value range to a set range, such as zero to one, to ensure consistency in subsequent neural network processing. In order to adapt to the depth noise characteristics of different types of sensor outputs, a preprocessing step based on bilateral filtering or spatial consistency constraints can be introduced before encoding to further improve the continuity and robustness of the encoded features.
[0059] In different application scenarios, the neighborhood size for local normal vector estimation can be adjusted based on the complexity of the target surface. When the target surface changes gently, the neighborhood size can be appropriately increased to improve smoothness. When the target area boundary is steep, the neighborhood size can be reduced to improve boundary sensitivity. The encoding processing module can be deployed on the edge computing unit for real-time operation or cached after batch processing in the cloud to meet the timeliness and computing resource requirements of different ultrasound inspection workflows.
[0060] Example description: In automatic ultrasound scanning of the thyroid region, the depth map obtained from the RGB-D image of the neck is encoded and processed. First, the horizontal disparity channel is calculated based on the original pixel points of the depth map to capture the subtle displacement changes in the depth direction of different tissue layers of the neck. Then, based on the horizontal reference plane set when the patient is in a supine position, the ground height channel of each point on the neck surface is calculated to extract the height change characteristics of the thyroid region relative to the surrounding tissues. The surface normal vector is further estimated in the local neighborhood and the angle with the direction of gravity is calculated to generate a normal angle channel. The three-channel data is merged and input into the target area recognition module as encoded depth information. In this way, even if the patient's neck fat tissue is thick, the muscle tissue is uneven, or the posture changes slightly, the thyroid region can still be accurately distinguished, ensuring the accuracy of the scan starting point planning and the stability of the subsequent scanning path, thereby improving the consistency of thyroid ultrasound imaging and diagnostic reliability.
[0061] This embodiment encodes depth information into multi-channel data that fuses horizontal parallax, ground height, and normal angle. This not only enriches the spatial geometric representation of the original depth information, but also significantly improves the feature differentiation and positioning accuracy of subsequent target scanning area identification. Horizontal parallax enhances the ability to detect lateral structural changes, ground height provides a stable vertical reference, and the normal angle characterizes subtle differences in surface orientation. The combined encoded depth information of these three effectively compensates for the information loss of single-channel depth maps in complex structure identification. This enables the automatic ultrasonic scanning system to achieve more accurate target positioning and path planning for different individuals, different body positions, and different surface morphologies, improving the overall system's universality and imaging quality.
[0062] S30, identifying a target scanning area within the target area based on the image information and the encoded depth information;
[0063] In this embodiment, in order to accurately determine the starting area of the ultrasonic scan, it is necessary to perform fusion processing based on the image information and the encoded depth information to identify the target scanning area within the target area. As a source of two-dimensional texture and structural features, image information can provide rich surface information, such as contour edges, texture patterns, local brightness changes, etc., which helps to preliminarily locate surface structural features. The encoded depth information supplements the structural properties such as surface undulations, tilt direction and height changes in three-dimensional space by fusing three types of geometric features: horizontal parallax, ground height and normal angle, thus making up for the shortcomings of two-dimensional images in spatial depth perception.
[0064] During the fusion process, features of the image information and the encoded depth information can be extracted separately. The image information can be used for global feature extraction using a convolutional neural network. This neural network uses multiple layers of convolution kernels to capture local texture and edge information, and models large-scale spatial relationships through downsampling and nonlinear transformations. The encoded depth information can be used to extract local geometric features using a spatial transformer network. The spatial transformer network learns affine or nonlinear spatial transformations based on the input data to improve recognition of areas with significant local structural changes.
[0065] The extracted global features are cascaded and fused with local geometric features. The resulting multimodal features simultaneously encode texture, edge, spatial morphology, and orientation information in a unified feature space, resulting in stronger target differentiation and location accuracy. To further enhance recognition contextual understanding, a temporal attention mechanism can be introduced to contextually enhance the multimodal fusion features. The temporal attention mechanism dynamically allocates attention weights within the feature sequence, strengthening the representation of features related to the target area and suppressing noise interference or background information, thereby improving overall recognition accuracy.
[0066] Finally, based on the enhanced fusion features, classification and positioning operations are performed through a fully connected layer, outputting the class prediction probability of the target scan area and the corresponding 3D positioning frame coordinates. The classification probability is used to distinguish the structure of interest from the background tissue, and the positioning frame coordinates are used to determine the scan starting point and planning range, ensuring the accuracy of subsequent scan path planning and probe control actions.
[0067] In practice, a set of pre-trained convolutional neural networks can be used to process image information and extract multi-scale texture features. For example, shallow convolutions extract edge information, while deep convolutions extract semantic generalization features. Simultaneously, the encoded depth information is fed into a spatial transformation network for geometric feature extraction, capturing surface structural variations and pose information in the target area. During the feature fusion stage, the two feature streams can be cascaded using feature concatenation or attention fusion modules to ensure effective alignment of features from different modalities within the same feature space.
[0068] During the context enhancement phase, a bidirectional temporal attention mechanism can be employed, combining forward and backward feature dependencies to strengthen the global relevance of local features. The context-enhanced features are passed through a two-layer fully connected network, which outputs a classification probability vector and 3D spatial localization box parameters. The classification probability vector is normalized using Softmax, and the category with the highest probability is selected as the recognition result. The 3D localization box parameters are then mapped to the real-world coordinate system through a nonlinear transformation to determine the position and size of the target scanning area.
[0069] The receptive field size of the convolutional neural network can be adjusted to account for the tissue characteristics of different target areas. Adding dilated convolution or pyramid pooling modules can expand the feature capture range to accommodate scenes with drastic surface morphology changes or sparse texture distribution. In different application scenarios, the fusion module can choose between weighted fusion or adaptive dynamic fusion strategies, automatically adjusting feature weights based on the characteristics of the input data to further optimize the comprehensive expression of multimodal information.
[0070] Example description: In the application of thyroid ultrasound examination, by collecting image information and encoded depth information of the neck area, a convolutional neural network is first used to extract the neck contour and tissue boundary features from the image information. At the same time, a spatial transformation network is used to extract the geometric undulation features of the thyroid area relative to the surrounding tissues from the encoded depth information. The two features are cascaded and fused and input into the temporal attention module for context enhancement. This can effectively highlight the texture and spatial features of the thyroid area and weaken the influence of interfering tissues such as surrounding muscles and blood vessels. Finally, the classification probability is output through the fully connected layer to accurately identify the thyroid tissue, generate a three-dimensional positioning frame, and determine the initial scanning starting point and range. Even in the case of changes in patient position, reflection or occlusion on the skin surface, it can still maintain high accuracy and high robustness of recognition effects, providing a stable and reliable starting reference for the subsequent standardized automatic scanning process.
[0071] This embodiment extracts and fuses features from image information and encoded depth information, enabling accurate identification of the target scanning area within the target area by leveraging the complementary advantages of two-dimensional texture structure and three-dimensional geometric properties. Image information provides surface details and texture clues, while encoded depth information complements spatial morphology and structural characteristics. The enhanced features resulting from the fusion of the two have stronger target perception capabilities and positioning accuracy. The introduction of the temporal attention mechanism further enhances the correlation and robustness between features, effectively suppressing background noise and artifact interference, enabling the system to stably identify the target scanning area under complex backgrounds and variable posture conditions, providing high-quality initial data assurance for subsequent scanning path planning and posture control, and overall improving the consistency, standardization, and reliability of automatic ultrasonic scanning.
[0072] S40, determining an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0073] In this embodiment, after identifying the target scanning area, the initial scanning point and initial scanning direction are further determined based on the target scanning area's location to provide accurate starting conditions for subsequent ultrasound scanning. The target scanning area's location is typically described by a positioning box in three-dimensional space, including center coordinates, size range, and orientation information. The initial scanning point is determined based on the specific spatial distribution characteristics of the target area. A reasonable starting position is selected to ensure that the scanning starting point covers the center of the area or key structures, while facilitating subsequent path planning.
[0074] The determination of the initial scan point can be derived based on pre-established spatial structural knowledge. Spatial structural knowledge refers to pre-existing cognitive information about the layout patterns of specific tissue structures. In medical scenarios, this can include the common locations, size ratios, and relative orientations of anatomical organs such as the thyroid and carotid arteries. This knowledge can be acquired through clinical statistical data, standard anatomical models, or machine learning methods. In applications, by analyzing the 3D positioning frame data of the target scan area and combining it with spatial structural knowledge, two representative starting points can be selected, located near the specific boundaries or center of the target area, to establish a basic relationship between the scan starting point and the initial scan direction.
[0075] Mapping the 2D image coordinates of the selected initial scan points to 3D coordinates in the 3D point cloud data is a key step in determining the scan starting point. This mapping process combines depth information with camera intrinsics, transforming from 2D to 3D coordinates through backprojection or geometric transformation. This allows the annotated points on the image plane to be accurately projected to their corresponding locations in real space.
[0076] Based on the positional relationship between the two initial scan points in three-dimensional space, the initial scan direction vector can be calculated. This vector defines the initial direction of probe movement. It is typically constructed by pointing the first scan point toward the second scan point and, after normalization, serves as the initial attitude reference for the scan control system. To ensure that the initial orientation aligns with the principal axis of the target region or the direction of structural extension, correction optimization can be performed in conjunction with the long axis of the positioning frame. This ensures that the probe orientation during scanning aligns with the natural extension of the anatomical structure, improving imaging consistency and effectiveness.
[0077] Through the above steps, the initial scanning point and initial scanning direction can be reasonably determined according to the position of the target scanning area, providing accurate starting conditions for subsequent path following, posture adjustment and image acquisition.
[0078] During implementation, the center point coordinates of the target scanning area's three-dimensional positioning frame can be first extracted as the first initial scanning point. To ensure that the scanning range covers the entire area, a second initial scanning point can be selected by offsetting the positioning frame outward by a fixed distance along its long axis. If the target area has an irregular shape, the centroid and the highest point on the boundary can be used as the basis for selecting the two initial scanning points to accommodate the needs of different structural characteristics. The two-dimensional pixel coordinates of the initial scanning point are back-projected into three-dimensional space using the depth map, and the actual three-dimensional space coordinates are calculated using pre-calibrated parameters. During the back-projection process, the depth values need to be filtered to remove isolated points and noise points to ensure the accuracy and stability of the mapping results. The initial scanning direction vector can be obtained by calculating the vector difference between the three-dimensional coordinates of the two initial scanning points and normalizing them to form a standard direction vector. To accommodate differences in patient position and posture, a posture correction mechanism based on the direction of gravity can be introduced to fine-tune the initial direction vector so that it forms a reasonable angle with the direction of gravity to avoid imaging distortion caused by excessive probe tilt. In specific application scenarios, such as when the target area has a distinct directional structure, the optimal initial direction vector can be dynamically derived by analyzing the surface normal distribution of the target area to further optimize the scanning effect. Different implementation methods can adjust the number of initial scanning points based on the complexity of the target structure. For example, when scanning complex organs, multiple starting points can be selected and the scanning path can be planned in stages. In scenarios with a single simple structure, the minimum two-point structure can be retained to improve efficiency.
[0079] Example description: In the task of automatic thyroid ultrasound scanning, the system identifies the three-dimensional positioning frame generated by the neck RGB-D image and depth data, determines the center point of the thyroid area as the first initial scanning point, and further extends a fixed distance in the direction of the long axis of the thyroid to determine the second initial scanning point. The two scanning points are back-projected into the spatial point cloud through the depth map to generate corresponding three-dimensional coordinate points, calculate to form the initial scanning direction vector, and fine-tune it according to the gravity reference direction to ensure that the starting direction of the probe is consistent with the natural extension axis of the thyroid gland. In this way, even under conditions of different patient body shapes, different supine angles or changes in neck curvature, the initial scanning point and direction can still be stably and accurately determined, laying a solid foundation for subsequent high-quality automated ultrasound scanning.
[0080] This embodiment combines the position data of the target scanning area with spatial structure knowledge to determine the initial scanning point and initial scanning direction, thereby establishing accurate and reasonable probe starting conditions before the ultrasound scan begins. The initial scanning point ensures that the scan covers the key target area, and the initial scanning direction conforms to the natural extension trend of the tissue, reducing image distortion or structural loss caused by posture deviation during the scanning process. Using positioning frame derivation and point cloud back-projection technology, stable and consistent starting point and direction planning can be achieved for different individuals, different postures, and different surface morphologies, improving the reliability, consistency, and standardization of automatic ultrasound scanning.
[0081] S50, controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process;
[0082] In this embodiment, after determining the initial scanning point and initial scanning direction, it is necessary to control the scanning probe to start scanning according to these starting conditions and acquire the scanned image in real time during the scanning process to provide input data for subsequent image processing and path adjustment. The primary task of controlling the scanning probe is to accurately move the end effector of the robotic arm to the position of the initial scanning point and adjust the posture of the scanning probe so that it is aligned with the initial scanning direction. The position of the initial scanning point is usually represented by three-dimensional spatial coordinates, and the initial scanning direction is defined in the form of a unit vector. The two together determine the spatial position and orientation of the probe.
[0083] During actual control, the 3D coordinates of the initial scan point must be analyzed as the target position input parameter for the motion control system. Combined with the initial scan direction, the target pose matrix of the robotic arm's end effector can be calculated. This matrix includes a translation component defining the starting position and a rotation component defining the probe's pointing direction. When generating motion control instructions, the joint angle changes must be calculated through an inverse kinematic solution to ensure that the robotic arm can smoothly move to the target position and complete pose adjustments within the physical constraints, avoiding motion failures caused by path singularities or joint limits.
[0084] After the probe is moved to the initial scanning point and adjusted, the ultrasound probe's scanning module is activated, and continuous scanning is performed along the initial scanning direction. Once the scanning module is activated, the probe's spatial position and contact status must be continuously monitored to ensure proper contact between the probe and the target surface, without excessive pressure or separation, to maintain stable image quality.
[0085] During real-time scan image acquisition, the scanning system captures ultrasound image frames at fixed intervals, simultaneously recording the acquisition timestamp and spatial pose data for each frame. Transmitting the real-time scan images to the image processing module allows for cache management and timestamp alignment, providing a consistent data foundation for subsequent image analysis, artifact detection, and path adjustment.
[0086] The entire scanning control link needs to have a certain degree of dynamic adaptability while ensuring scanning stability and image acquisition continuity. For example, when a slight deviation in posture or abnormal contact is detected, the posture can be quickly adjusted or the scanning can be paused to avoid image distortion or data loss.
[0087] In specific implementations, the 3D coordinates of the initial scan point are analyzed through the robotic arm control interface and input as the target position into the motion control module. Combined with the initial scan direction, the target pose is calculated using quaternions or Euler angles, generating a complete end-effector target transformation matrix. Based on the inverse kinematics solution, the robotic arm is controlled to smoothly move along the optimal path to the target position, while dynamically monitoring changes in joint angles to prevent entry into singularity zones or exceeding mechanical limits. As the end-effector approaches the initial scan point, a fine-tuning mechanism based on visual or force feedback can be introduced to further correct the final position and orientation, ensuring good contact between the probe and the target surface and that the pose is consistent with the initial scan direction. Fine-tuning can be achieved through small-scale position shifts, pose rotations, or contact force adjustments. After the probe is positioned, the ultrasound module is activated to acquire real-time images at a set frequency, for example, 20 to 30 ultrasound frames per second. Each image frame is recorded with the corresponding acquisition timestamp and probe pose, and transmitted to the image processing module in real time via wired or wireless communication. To ensure data synchronization, a unified system clock can be used to mark events across modules, or dedicated synchronization signal lines can be introduced for hardware-level timing synchronization. To adapt to the size and morphological changes of different target areas, an adaptive scanning speed adjustment mechanism can be set during the scanning process to dynamically adjust the probe movement speed and image acquisition frequency according to the complexity of the target area to ensure that the image frame spacing covers the area continuously and without omissions.
[0088] Example: During an automated thyroid ultrasound examination, the robotic arm control system moves the ultrasound probe to the initial scanning point identified in the thyroid region, while simultaneously adjusting the probe's posture to align along the long axis of the thyroid gland. A force sensor monitors the contact pressure between the probe and the neck skin in real time to ensure neither compression of the tissue causing deformation nor signal weakening due to insufficient contact. After the ultrasound module is activated, real-time ultrasound images are continuously acquired at a frequency of 25 frames per second, and the probe's spatial posture data is simultaneously recorded. All image frames are timestamp-aligned and cached in the image processing module for subsequent continuous target tracking and artifact region detection. This ensures that even if the patient experiences slight movement during the examination, the system can still obtain high-quality image sequences, ensuring consistency and standardization of thyroid ultrasound imaging.
[0089] This embodiment can significantly improve the standardization of the ultrasound scanning process and the consistency of image acquisition by precisely controlling the scanning probe to start scanning based on the initial scanning point and initial scanning direction, and by acquiring images in real time during the scanning process. The accurate alignment of the probe's starting position and direction ensures the spatial coverage integrity and structural coherence of the initial image data. Real-time image acquisition and synchronous recording of posture provide a complete and reliable data foundation for subsequent target tracking, artifact detection, and path adjustment. Combined with fine-tuning and adaptive mechanisms, it can effectively address image distortion problems caused by posture drift or contact changes during the scanning process, thereby improving the robustness, stability, and imaging quality of the automatic ultrasound scanning system.
[0090] S60, analyzing the real-time scanning image, and identifying a preset target and an artifact area in the real-time scanning image;
[0091] In this embodiment, in order to dynamically evaluate the scanning effect and identify abnormal areas during the ultrasound scanning process, it is necessary to analyze the real-time scanning image and identify the preset targets and artifact areas in the real-time scanning image. The real-time scanning image refers to the ultrasound image frames collected at continuous time intervals during the scanning process, which contains the tissue structure information of the area covered by the current probe. The preset target usually refers to a specific anatomical structure or tissue feature that requires attention, such as thyroid parenchyma, vascular channels or pathological nodules. The artifact area refers to the abnormal image area caused by improper probe posture, abnormal contact pressure, sudden change in tissue acoustic impedance or scanning path deviation, including but not limited to shadow interference, enhanced reflection, lateral artifacts and other phenomena.
[0092] The first step in analyzing real-time scanned images is to preprocess the images, including size standardization, grayscale normalization, and noise suppression, to eliminate the impact of differences in imaging parameters between different frames on the recognition results. Subsequently, image features are extracted for discrimination between preset targets and artifact areas. Preset target recognition can be based on the extraction of local texture features and global spatial structure features by convolutional neural networks, and the output of a target existence probability map through a classifier. Artifact area recognition can be based on brightness thresholds, texture anomaly detection, or deep neural network segmentation modules, by detecting low signal-to-noise ratio areas, areas without structured continuity, or areas with strong reflective anomalies to mark potential artifact areas.
[0093] The identified preset targets and artifact areas need to be spatially localized, typically expressed as bounding boxes, polygonal segmentation masks, or probability heatmaps. To improve recognition stability, a temporal smoothing mechanism can be introduced to filter the recognition results of multiple consecutive frames to suppress jitter or misjudgment caused by single-frame recognition errors.
[0094] Through real-time image analysis and target artifact recognition, the scanning quality can be dynamically evaluated during the scanning process, providing a basis for subsequent posture adjustment, path correction and scanning termination judgment, forming a closed-loop control system.
[0095] In specific implementations, a pre-trained target detection network can be used to infer real-time scanned images and output detection results for pre-set targets. The network can be based on an anchor-free architecture to improve the sensitivity and localization accuracy of small target recognition. To adapt to the texture characteristics of ultrasound images, dilated or directional convolution modules can be added to the feature extraction layer to enhance the perception of fine-grained structural changes. Artifact region detection can be achieved using a multi-scale feature pyramid network, combining local texture consistency and brightness variation characteristics to extract candidate artifact regions. To reduce the false recognition rate, a contextual information modeling module can be introduced into the artifact discrimination stage to analyze the continuity of tissue structures surrounding the candidate regions, thereby determining the boundary clarity and texture coherence between the artifact region and normal tissue. The identified pre-set targets and artifact regions can be mapped to a spatial coordinate system after image plane annotation, and spatial localization can be achieved by combining the real-time probe pose data. The recognition results of consecutive frames can be smoothed in the time domain using a Kalman filter or a weighted sliding average mechanism to improve recognition stability and robustness. In different application scenarios, the operating frequency and recognition accuracy of the image analysis module can be dynamically adjusted according to the complexity of the target area and the artifact distribution characteristics to balance the processing load and recognition accuracy.
[0096] Example description: During the automatic ultrasound scanning of the thyroid gland, the system collects neck ultrasound image frames in real time, detects the thyroid tissue area through a convolutional neural network, and uses the artifact detection module to mark the shadow interference area that appears in the image. When it is detected that the recognition probability of thyroid tissue decreases or the coverage area of the artifact area exceeds the set threshold, the system dynamically adjusts the scanning probe posture or local scanning path to make corrections to ensure that the quality of the subsequently acquired images meets the diagnostic requirements. Even under conditions of patient respiratory movement or slight changes in body position, the system can continue to track the target area and identify abnormal areas, maintain the consistency and reliability of high-quality imaging, and provide strong support for the accurate screening and evaluation of thyroid lesions.
[0097] This embodiment dynamically assesses scanning progress and image quality by analyzing scanned images and identifying preset targets and artifact areas in real time during the scanning process, promptly identifying imaging issues caused by posture deviation, contact anomalies, or path errors. Stable recognition of preset targets ensures the pertinence and consistency of scan content. Real-time detection of artifact areas enables the system to proactively adjust scanning parameters or trajectories when artifact interference is severe, effectively improving the imaging quality and intelligent operation of automated scanning, reducing the risk of invalid images or misdiagnoses due to scanning anomalies, and enhancing the system's adaptive control capabilities and clinical reliability.
[0098] S70, adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0099] In this embodiment, during real-time scanning, after identifying the preset target and artifact areas, the scanning probe's scanning posture and scanning path are dynamically adjusted based on the identification results to optimize imaging and improve scanning quality. The scanning posture refers to the orientation and tilt angle of the scanning probe in three-dimensional space, while the scanning path refers to the path of the scanning probe along the spatial trajectory. The purpose of adjusting the scanning posture and scanning path is to minimize the impact of artifact areas while ensuring that the target area remains within the scanning field of view, thereby avoiding image loss or severe distortion due to deviations in the scanning posture or an unreasonable scanning path.
[0100] Adjusting the scanning posture first requires analyzing the location and area of the artifact area in the image. If the artifact area is concentrated in a specific area of the image, such as above or to one side, it means that the scanning probe posture may have deviated from the optimal angle of incidence, resulting in the sound beam being unable to enter the tissue surface perpendicularly. At this time, it is necessary to calculate the angle changes that need to be fine-tuned in the scanning probe based on the spatial distribution characteristics of the artifact area, including pitch angle adjustment, roll angle adjustment, or yaw angle adjustment, so that the probe is oriented more perpendicular to the surface of the target area, reducing the probability of artifact generation and improving the quality of the echo signal.
[0101] Adjusting the scanning path requires making decisions based on the overall distribution of the artifact area and the continuous tracking status of the preset target. If the artifact area covers the main part of the scanning path, or the preset target is continuously lost, the scanning path needs to be updated to bypass the artifact-intensive area or re-plan a new path covering the target area. The updated scanning path can dynamically generate a new motion trajectory based on the position and shape information of the artifact area in three-dimensional space, combined with the current position of the probe and the distribution of the target area. Continuity must be ensured during the path planning process to avoid sharp turns or excessive acceleration changes to ensure a smooth and stable scanning process and maintain consistency and coherence in image acquisition.
[0102] During the adjustment process, the contact force and spatial posture of the scanning probe need to be monitored in real time to ensure that the adjustment does not cause the probe to separate from the target surface or excessive pressure, and to ensure that the probe and the target surface maintain appropriate contact during the scanning process, laying the foundation for subsequent high-quality image acquisition.
[0103] In practical applications, the real-time artifact distribution analysis module can detect the distribution characteristics of artifact regions in ultrasound images and extract the center position, coverage area, and distribution direction of the artifact regions. Based on the relative offset between the artifact center position and the image center, the required probe posture adjustment is calculated, generating pitch, roll, or yaw angle adjustment commands. The robotic arm control module then smoothly executes small posture adjustments to prevent imaging continuity from being affected by large angle changes. If the artifact area exceeds a set threshold or the preset target is lost in multiple consecutive image frames, the system can trigger a path replanning mechanism. Path replanning can generate a new scanning trajectory using a sampling or search-based local path optimization algorithm based on the current probe spatial position, the target region positioning frame, and the spatial distribution of the artifact region. For example, a heuristic search method can be used to find a motion path with minimal artifact interference and optimal target visibility. During the path adjustment process, path smoothing constraints can be set to limit the curvature change rate and maximum acceleration to avoid image acquisition interruption or distortion caused by sharp turns or sudden deceleration. As the probe moves along the new path, it continues to capture images and identify targets and artifacts in real time, forming a closed-loop adjustment mechanism that ensures dynamic adaptation during the scanning process. Different adjustment strategies can be selected based on the morphological characteristics of the target area in different scenarios. For example, for slender targets, a posture fine-tuning strategy can be prioritized to maintain scanning along the long axis. For blocky targets, a dual strategy of posture adjustment and path optimization can be used to ensure full coverage.
[0104] Example description: In automatic thyroid ultrasound scanning, the collected ultrasound images are analyzed in real time, the thyroid tissue area is identified as the preset target, and the shadow artifact area appearing in the image is detected. When the system detects that the shadow artifact is mainly concentrated in the upper half of the image, it infers that the current scanning probe pitch angle is too large, and immediately generates a fine-tuning instruction to reduce the pitch angle so that the sound beam is more vertically incident on the thyroid surface. If the probability of thyroid tissue recognition decreases and the shadow artifact area continues to expand in several consecutive frames of images, the system triggers the path replanning mechanism, generates a new scanning path that bypasses the artifact-intensive area, and restarts the scan from the new starting point to ensure that the thyroid area is fully covered. Through this real-time recognition and dynamic adjustment mechanism, the high quality of the ultrasound image and the consistency of the scanning process can be maintained even in the event of sudden artifacts caused by changes in the patient's body position or local tissue abnormalities, significantly improving the intelligence level and diagnostic accuracy of automatic thyroid ultrasound examinations.
[0105] This embodiment dynamically adjusts the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and artifact area, and can deal with the problem of image quality degradation caused by tissue morphology changes, probe posture deviation or artifact interference during the scanning process in real time. The posture adjustment ensures the perpendicularity of the probe sound beam direction to the tissue surface, significantly reduces shadow artifacts and reflection distortion, and improves image clarity and detail visibility. Path optimization avoids blind scanning of the probe in artifact-intensive areas, and improves the effective image acquisition rate and target coverage integrity. Overall, the dynamic adjustment of the scanning posture and path enables the automatic ultrasound scanning system to have the ability to adapt to environmental changes, improves the stability, intelligence and imaging quality of the scanning process, and lays a solid foundation for high-precision, standardized automatic ultrasound diagnosis.
[0106] S80, monitoring the continuous existence of a preset target in the real-time scanning image;
[0107] In this embodiment, during ultrasound scanning, to ensure that the scanning probe continuously covers the target area, it is necessary to dynamically monitor the continuous presence of pre-set targets in the real-time scan image. Pre-set targets refer to clinically significant tissue structures that should persist in the image, such as thyroid parenchyma, vascular structures, or pathological nodules. Continuous presence means that the system can stably identify the pre-set targets within a continuous time window, maintaining the spatial continuity and probabilistic stability of the recognition results.
[0108] To monitor the continuous presence of a pre-set target, it is first necessary to perform target recognition on each frame of the real-time scanned image, extracting metrics such as target presence probability, spatial location, and outline morphology. Single-frame recognition results cannot reflect temporal continuity, so the recognition results must be cached in chronological order and analyzed within a sliding time window. The length of the time window can be set based on the scanning speed and image acquisition frequency to ensure coverage of continuous image frames within a given scanning area.
[0109] Within a time window, the stability of the target's presence can be assessed by calculating statistical indicators such as the mean, variance, and minimum value of the target's presence probability across consecutive frames. When the target's presence probability is high and its spatial position changes minimally across several consecutive frames, the target can be considered to be continuously present. Conversely, when the target's presence probability is low, its spatial position changes dramatically, or consecutive frames are lost, the continuous presence is considered interrupted.
[0110] Continuous presence monitoring requires a certain level of fault tolerance to accommodate short-term recognition jitter caused by image quality fluctuations or slight changes in probe posture. For example, a short-term abnormal frame tolerance mechanism can be set. When target recognition fails in a single frame but the surrounding frames remain stable, the continuous state is considered uninterrupted, avoiding misjudgment and termination of the scan due to occasional noise.
[0111] By real-time monitoring of the continuous existence of preset targets, the effectiveness of the scan can be dynamically evaluated, and probe offset, path error or contact anomaly problems during the scan process can be discovered in a timely manner, providing a basis for subsequent adjustment actions or termination of the scan, thus forming a closed-loop quality control during the scanning process.
[0112] During implementation, a cache queue can be configured within the image recognition module to store the target recognition results from the most recent frames of real-time scanned images. Each recognition result includes attributes such as target presence probability, center coordinates, and contour boundaries. After receiving each new frame, the system queues the new recognition result and removes the oldest frame, maintaining a fixed-length sliding window. Within the sliding window, a threshold for target presence probability can be set. For example, if the target presence probability exceeds the threshold in more than 80% of frames and the center coordinates vary within a reasonable range, the target is considered continuously present. If the target presence probability falls below the threshold in several consecutive frames, or if the identified target center position drifts beyond a preset spatial distance, the continuous presence state is considered interrupted. To further enhance the robustness of the determination, a weighting strategy can be introduced to assign different weights to frames at different positions within the time window, such as assigning higher weights to the latest frames, thereby increasing the system's sensitivity to recent trends. In different application scenarios, the sliding window length and the judgment threshold can be adapted based on the target area size, scanning speed, and tissue motion characteristics. For example, a longer window and strict threshold can be used in static organ scanning, while a short window and loose tolerance strategy can be used in dynamic organ scanning to balance recognition stability and real-time requirements.
[0113] For example, the following formula is used to define the continuous detection status indicator:
[0114]
[0115] Each frame of real-time scanning image is judged by the image recognition network whether the thyroid region is detected. If it is detected, k t Set to 1, if not detected, set to 0. In the continuous m frames of image, the system sets the k t The values are accumulated to obtain the sliding window cumulative value C s According to the cumulative value C s Determine the continuous existence of the target. s If the window size exceeds half of the target, the thyroid target is considered stable and present; otherwise, the target is considered missing. The window size m can be dynamically adjusted based on the actual acquisition frame rate, scanning speed, and target feature stability. For example, for high-frame-rate probes, m can be appropriately increased to improve system stability, while for low-frame-rate systems, m can be reduced to improve response speed.
[0116] Scan system status S t Based on the following formula definition:
[0117]
[0118] During the scanning process, the system accumulates the value C according to the sliding window. s The relationship between the time window size m and the current scanning state is dynamically adjusted. If the current state is SB (The target does not appear) and the cumulative detection value is greater than m / 2, the state switches to S C (Target visible); if the current state is S C If the cumulative detection value is less than m / 2, switch to S E (Target disappears). This regular switching of scanning states ensures the system can respond promptly to target appearances and disappearances, avoiding misjudgments caused by transient detection errors. In actual deployments, state transition logic can be implemented in real time through a soft logic control module, supporting dynamic parameter adjustment to adapt to different scenarios.
[0119] According to the current scanning stage state S t , define the target reference position P O as follows:
[0120]
[0121] During the scanning process, according to the current state S t Dynamically select the position reference point. When the target is visible S C When the current detected target center position P t As a position reference, to ensure that the probe follows the target dynamic fine-tuning; when the target does not appear (S B ), it returns to the predefined preset position P a As a reference, it ensures that the probe does not drift or lose control due to recognition failure. a The reference position is typically set to the center of the scan area or a safe starting position manually calibrated by the physician. The reference position switching logic can be dynamically executed by the state machine control module to ensure the continuity and stability of the probe movement.
[0122] Example: In the thyroid automatic ultrasound scanning task, the system analyzes the recognition results of thyroid tissue in each frame of the ultrasound image in real time, recording the probability of thyroid existence and the center coordinate position. By maintaining a sliding window covering the most recent thirty frames of images, the average value of the thyroid existence probability and the change in the center position are calculated. When the probability of thyroid existence is detected to be continuously higher than the set threshold and the center position is stable, the system determines that the scan continues to cover the target area and continues to advance along the current path. When the probability of thyroid existence decreases or the center position drifts abnormally in five consecutive frames of images, the system determines that the target continuous existence is interrupted and triggers path adjustment or scan termination. Through this real-time monitoring of the continuous existence state, the system can detect abnormalities and respond in time even in the event of slight patient movement, breathing changes, or deviation of the scanning probe posture, ensuring continuous imaging and image quality consistency of the thyroid area, thereby improving the stability and diagnostic reliability of automatic ultrasound scanning.
[0123] By monitoring the continuous presence of a preset target in real time during the scanning process, this embodiment dynamically assesses whether the scanning probe maintains continuous coverage of the target area, effectively preventing imaging loss due to path deviation, contact failure, or posture changes. This continuous presence monitoring mechanism provides real-time quality control during the scanning process, enabling the system to make timely adjustments or terminate the operation when scanning anomalies are identified. This ensures the consistency and diagnostic effectiveness of ultrasound image acquisition, and enhances the intelligence and clinical applicability of automated ultrasound scanning.
[0124] S90: If the preset target is not recognized in the continuous multiple frames of images, terminate the scanning.
[0125] In this embodiment, during ultrasound scanning, if the preset target is not recognized in multiple consecutive frames, it indicates that the scanning probe has deviated from the target area, or imaging failure has occurred due to posture misalignment, contact abnormality, tissue obstruction, etc. To avoid wasting time, accumulating invalid images, and occupying device resources due to continued invalid scanning, it is necessary to terminate the scanning operation promptly after confirming that the target has been lost continuously.
[0126] The determination of missing a target across multiple consecutive frames is based on the sequence of recognition results from the real-time scanned images. During the real-time image recognition process, each frame generates data such as the target's presence probability, spatial position, and outline. If the target's presence probability in several consecutive frames falls below a predetermined threshold within a set sliding time window, and spatial position tracking fails, the target is considered lost.
[0127] To improve judgment accuracy, a redundancy tolerance mechanism can be configured. For example, individual abnormal frames can be tolerated, but termination can be triggered as soon as the overall trend indicates continuous target loss. Parameters such as the continuous loss judgment window size, probability threshold, and maximum permissible drift range can be configured based on the accuracy requirements of the scanning task, tissue type, and imaging characteristics to adapt to different application scenarios.
[0128] When the conditions are met, the system issues a scan termination command, immediately halting the robotic arm's movement and ultrasound image acquisition. It also records the current scan status and the cause of the abnormality, providing a basis for subsequent system self-tests, logging, or manual review. Terminating the scan requires ensuring that the probe enters a safe state after termination, such as slowly withdrawing the probe or keeping it stationary, to avoid causing additional interference or damage to the scanned object.
[0129] By terminating invalid scans in a timely manner, the efficiency of system resource utilization can be significantly improved, the amount of invalid data can be reduced, the risk of misdiagnosis can be lowered, and the intelligence, standardization and efficiency of the scanning process can be ensured.
[0130] In a specific implementation, the real-time image recognition module continuously outputs data on the presence probability of a preset target in each image frame, maintaining a fixed-length sliding window to record the recognition results of the most recent frames. A continuous loss determination is triggered when the target presence probability falls below a set threshold for more than a set number of consecutive frames within the window, and the spatial tracking module is unable to associate a valid target pose trajectory. Once the continuous loss determination is established, the system immediately sends a termination command to the robotic arm motion control module, halting the current scanning task. Termination commands can include stopping motion, shutting down the ultrasound acquisition module, locking the current probe pose, or slowly returning the probe to its initial position, ensuring a smooth and safe termination process. The parameters of the termination condition can be adjusted to suit different applications. For example, in thyroid scans, termination can be determined after twenty consecutive frames have been lost. In vascular scans, where targets are small and prone to drift, the tolerance can be increased, for example, allowing for a short period of ten consecutive frames to be lost before determining the target. For scanning tasks in high-risk areas, redundant verification of the termination determination, such as multiple rounds of detection and confirmation, can be added to further improve reliability. In the system log module, the trigger time, termination reason, loss detection parameters, and the last valid identification location and image of the termination event can be synchronously recorded for subsequent auditing and optimization.
[0131] For example, the logic for automatically terminating the scan is based on the cumulative detection value C s :
[0132]
[0133] Continuously monitor the target recognition in each frame of the image, and accumulate the target presence mark k in m frames t , get the cumulative existence value C s When C s When the m value is less than a preset threshold (e.g., less than m / 2), it is considered a continuous loss, and the system automatically triggers a stop-scan command, halting the robotic arm's movement and ultrasound data stream acquisition. Different m values and termination thresholds can be set based on target stability characteristics in different application scenarios. For example, m can be shortened to speed up the response when scanning superficial structures, while m can be increased to avoid misjudgments when scanning deep tissue.
[0134] Example: During a thyroid ultrasound scan, ultrasound imaging equipment, an RGB-D camera, and a robotic arm with seven degrees of freedom are deployed. The multimodal image processing and analysis unit processes the captured RGB and depth images of the patient's neck in real time. The system uses the RGB-D camera to synchronously acquire color image information and depth information of the patient's neck. It then uses a pre-calibrated internal parameter matrix to align the pixel coordinate systems of the two types of information. It then uses neighbor interpolation to fill in invalid pixels in the depth information to generate a complete depth map. This is then mapped into 3D point cloud data using perspective projection, resulting in a set of 3D coordinate points on the surface of the target area.
[0135] When processing the collected depth information, the system calculates the horizontal disparity channel based on the depth sensor's disparity model, the ground height channel based on the gravity reference axis, and the normal angle channel based on the surface normal estimation algorithm. These three channels are then fused into multi-channel encoded depth information. When identifying the target area, the system extracts the global features of the image information through a convolutional neural network, while simultaneously extracting the local geometric features of the encoded depth information through a spatial transformation network. After cascading these two, contextual enhancement processing is performed using a temporal attention mechanism. Finally, the fully connected layer outputs the classification probability and 3D positioning frame coordinates of the target scan area, accurately locating the thyroid gland in the patient's neck.
[0136] After acquiring the position of the thyroid gland, the system defines the first and second initial scanning points of the target area based on pre-set spatial structure knowledge, which includes prior knowledge of medical anatomy. The system then converts the two-dimensional coordinates of the initial scanning point into first and second three-dimensional coordinate points in a three-dimensional point cloud using a mapping function, and calculates the initial scanning direction vector from the first point to the second point. The three-dimensional coordinates of the initial scanning point are then parsed, and motion control instructions for the robotic arm end effector are generated based on the three-dimensional coordinates and the initial scanning direction. The robotic arm end effector is driven to move to the initial scanning point and the scanning posture of the ultrasound probe is adjusted. The scanning module of the ultrasound probe is activated in this adjusted posture, continuously scanning along the initial scanning direction and acquiring ultrasound images in real time. The scanned image data is then transmitted to the image processing module for caching and timestamp attachment for subsequent synchronous processing.
[0137] During real-time scanning, the system continuously analyzes the recognition results of preset targets and artifact areas in the ultrasound image. If an artifact area is detected in the real-time scanning image, the artifact area distribution parameters are generated by detecting the artifact position and coverage area. Based on this parameter, the pixel offset of the scanning probe in the image coordinate system is calculated and mapped into three-dimensional offset coordinates. The system calculates the in-plane adjustment distance based on the current probe position and offset. When the adjustment distance is less than the set threshold, the posture of the scanning probe is slightly corrected to optimize the contact angle. When the adjustment distance is greater than or equal to the set threshold, the scanning path is updated to avoid the artifact area, and the contact pressure between the probe and the target area is dynamically adjusted based on the hybrid position and force control strategy to ensure that the probe always fits the patient's skin evenly with a reasonable torque, thereby improving image quality.
[0138] The system also monitors the continuous presence of pre-set thyroid targets in ultrasound images in real time. If the thyroid target cannot be identified in multiple consecutive image frames, the system determines that the thyroid region scan is complete and triggers an automatic termination mechanism, controlling the robotic arm to safely return the ultrasound probe to its initial position, completing the ultrasound scan. Through the control and adjustment of this entire process, this automated thyroid ultrasound scanning process effectively improves the accuracy of thyroid positioning, ensures high-standardized ultrasound image quality, reduces the uncertainty and occupational injury risks associated with manual manipulation, and significantly enhances the stability and intelligence of the scan.
[0139] This embodiment effectively avoids wasted time and the accumulation of invalid images caused by blindly scanning the probe outside the target area by promptly terminating the scan when the preset target is continuously lost. This termination mechanism can identify scan failures in real time and take action, significantly improving the intelligent management of the scanning process, increasing system resource utilization and the validity of diagnostic data, while reducing equipment operating load and operational risks. The timely and smooth termination of the action further ensures patient safety and experience, meeting the technical goals of an efficient, intelligent, and reliable automated ultrasound scanning system.
[0140] The present invention relates to the field of computer vision technology, and discloses a thyroid ultrasound robot automatic scanning method, device, equipment and medium based on RGB images and depth information, including: obtaining image information and depth information of a target area, encoding the depth information to generate encoded depth information, identifying a target scanning area within the target area based on the image information and the encoded depth information, determining an initial scanning point and an initial scanning direction according to the position of the target scanning area, controlling a scanning probe to scan according to the initial scanning point and the initial scanning direction and collect a real-time scanning image, analyzing the real-time scanning image to identify a preset target and an artifact area, adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area, monitoring the continuous existence state of the preset target in the real-time scanning image, and terminating the scanning when the preset target is not recognized in multiple consecutive frames of images. The present invention accurately identifies the target scanning area by fusing image information and coded depth information, controls the scanning probe to carry out standardized scanning in combination with the initial scanning point and scanning direction, and dynamically adjusts the scanning posture and scanning path based on the real-time image recognition results during the scanning process. At the same time, the scanning termination timing is determined according to the continuous detection status of the preset target, thereby achieving accurate positioning, stable image quality, and intelligent optimization of the scanning path during the ultrasonic scanning process, thereby improving the automatic ultrasonic imaging effect and standardization of the thyroid gland and similar superficial organs.
[0141] In one embodiment, the above step S10 includes:
[0142] S101, capturing image information and depth information of a target area through a color depth camera;
[0143] S102, aligning the depth information with the pixel coordinate system of the image information based on a pre-calibrated parameter matrix;
[0144] S103, performing neighboring interpolation filling on invalid pixels in the depth information to generate complete depth information;
[0145] S104 , mapping the complete depth information to generate three-dimensional space point cloud data through perspective projection transformation, where the three-dimensional space point cloud data is a set of three-dimensional coordinate points on the surface of the target area.
[0146] In this embodiment, to accurately obtain the spatial location information and surface structural features of the target area, it is necessary to simultaneously obtain image information and depth information of the target area. The color depth camera has the ability to simultaneously capture visible light images and corresponding pixel depth values, and can simultaneously generate color image data and depth map data during a single capture process. The image information provides texture, brightness, and structural information of the target area on a two-dimensional plane, while the depth information describes the relative spatial distance corresponding to each pixel.
[0147] Because image and depth information come from two different photosensitive modules, there is often parallax, so they need to be aligned in pixel coordinate systems based on a pre-calibrated parameter matrix. This pre-calibrated parameter matrix includes an intrinsic parameter matrix and an extrinsic parameter matrix. The intrinsic parameter matrix describes the imaging characteristics of the camera itself, such as focal length and principal point position, while the extrinsic parameter matrix describes the spatial relationship between the color image sensor and the depth sensor. Using these matrices, the pixels in the depth image can be mapped to the coordinate system of the color image, completing the spatial alignment between the two.
[0148] During the acquisition process, depth information is easily affected by lighting conditions, surface materials, and obstructions, resulting in missing or abnormal pixel values. Therefore, invalid pixels in the depth information need to be repaired. Neighboring interpolation is an efficient repair method. By analyzing the valid pixel values in the neighborhood of an invalid pixel, it uses weighted averaging, nearest neighbor values, or local surface fitting to supplement them, generating coherent and complete depth information and eliminating spatial structural breaks caused by missing pixels.
[0149] After completing depth information restoration, the combination of each pixel's 2D image coordinates and depth value is converted into a 3D coordinate point through perspective projection, forming 3D point cloud data. Based on the camera's imaging model, perspective projection back-projects each pixel into 3D space according to the known focal length, principal point position, and depth value. 3D point cloud data is a collection of spatially discrete points on the surface of the target area, accurately reflecting the 3D morphology and spatial structural characteristics of the target area's surface, providing a spatial foundation for subsequent target positioning, path planning, and probe control.
[0150] In the specific implementation, a color depth camera with integrated RGB and infrared depth sensors can be used to simultaneously capture image and depth information of the target area. After acquisition, the depth image is first remapped using calibrated internal and external parameters to align each pixel position with the color image. Bilinear interpolation can be used during the alignment process to ensure pixel density continuity during the mapping process. After coordinate alignment, invalid pixels in the depth image are filled using neighboring interpolation. Weighted nearest neighbor interpolation can be used as the interpolation method, searching for the nearest valid pixel set around the invalid pixel and generating a fill value based on a distance-weighted average. To improve interpolation quality, a maximum interpolation radius and a minimum valid pixel count threshold can be set to ensure the coherence and spatial rationality of the inpainted area. After filling, each pixel is back-projected using a perspective projection formula. Combining the depth value and the camera's internal parameters, the pixel position in the image coordinate system is transformed into a real-world coordinate in three-dimensional space. The resulting three-dimensional point cloud data can be stored in a sparse matrix or point cloud format, containing the spatial location attributes of each 3D coordinate point, which can be used for subsequent target scanning area identification, initial scanning point setting, and path planning. In different application scenarios, depth cameras with different resolutions can be selected according to the complexity and size of the target area, and the depth map resolution and projection accuracy parameters can be adjusted to balance the amount of collected data and spatial accuracy requirements.
[0151] This embodiment achieves accurate restoration of the three-dimensional surface structure of the target area by synchronously capturing and aligning image and depth information using a color depth camera. Interpolation and restoration of depth information ensures data continuity and integrity, effectively eliminating structural fractures caused by acquisition defects. The three-dimensional point cloud data generated through perspective projection transformation provides a reliable spatial basis for subsequent target recognition, path planning, and scanning probe control, significantly improving the automatic scanning system's spatial perception capabilities and target positioning accuracy, laying the foundation for high-quality autonomous scanning.
[0152] In one embodiment, the above step S20 includes:
[0153] S201, calculating a horizontal disparity channel of the depth information based on a depth sensor disparity model;
[0154] S202, calculating a ground height channel of the depth information according to a gravity direction reference axis;
[0155] S203, generating a normal angle channel of the depth information based on a surface normal vector estimation algorithm;
[0156] S204: Merge the horizontal disparity channel, the ground height channel, and the normal angle channel into multi-channel coded depth information.
[0157] In this embodiment, to enrich the expressiveness of depth information, the original depth map is subjected to feature encoding processing to extract the spatial structure information implicit in the depth map from different physical and geometric perspectives. The horizontal disparity channel is a feature calculated based on the depth sensor's disparity model. The disparity value is inversely proportional to depth and can reflect the relative distance changes between different locations on the object's surface and the camera. By analyzing the disparity corresponding to each pixel in the depth map, the hierarchical structure and local relief characteristics of the object's surface can be revealed.
[0158] The ground height channel is a feature calculated based on a gravity-referenced axis, describing the vertical height of each pixel relative to the ground. By establishing a spatial coordinate system aligned with gravity, depth information can be normalized to a unified vertical reference, improving comparability across different scan subjects and body positions. Ground height not only provides absolute height information but also aids in determining the spatial hierarchical relationships between different anatomical structures.
[0159] The normal angle channel is a feature generated by a surface normal estimation algorithm. It reflects the angle between the surface tangent plane normal and the direction of gravity or other reference directions. By calculating the surface gradient within each pixel neighborhood, the local surface orientation can be estimated, and the normal angle can be further derived. This channel effectively captures details such as changes in surface tilt and curvature, playing an important role in distinguishing structural boundaries and detecting unusual protrusions or depressions.
[0160] After extracting the features from the three channels, the horizontal disparity channel, ground height channel, and normal angle channel are fused in a fixed order to form multi-channel encoded depth information. This multi-channel encoded depth information not only contains the original depth distance information but also integrates the structural hierarchy, height distribution, and surface orientation characteristics, greatly enriching the spatial feature expression capability and providing more comprehensive feature support for subsequent target recognition, pose estimation, and path planning.
[0161] In the specific implementation, the original depth map can first be back-projected. The depth values are then converted into horizontal disparity values using a known disparity formula to form a horizontal disparity channel. The conversion process must take into account the sensor's built-in disparity reference distance and focal length parameters to ensure the accuracy of the disparity calculation. To calculate the ground height channel, each 3D point cloud coordinate can be projected onto the gravity axis based on the camera's gravity reference direction obtained during initial calibration, and the vertical height from the point to the reference plane is calculated. To accommodate different patient positions, the position and orientation of the reference plane can be dynamically adjusted to ensure the consistency and robustness of the ground height feature. To generate the normal angle channel, a local plane fitting method or a gradient operator method can be used. For each valid depth pixel, several neighboring pixels are selected within its neighborhood. A local tangent plane is fitted using least squares. The normal vector direction of this plane is calculated, and the angle between the plane and the preset reference direction is then calculated. To improve the stability of the normal vector estimation, a minimum neighborhood radius and a maximum tolerance error range can be set to filter out noise and outliers. After the three feature channels are calculated, they are normalized to the same numerical range to ensure that the feature scales of different channels are consistent, facilitating subsequent fusion processing. During the fusion process, the three channels can be directly stacked according to the channel dimension to form a three-channel encoded depth map, or appropriate dimensionality increase or decrease can be performed based on the input requirements of the specific recognition network.
[0162] In different application scenarios, different normal vector estimation methods and disparity calculation accuracy can be selected according to the surface complexity of the target area. For example, high-precision small neighborhood fitting can be used when detecting small vascular structures, and low-resolution fast estimation can be used when imaging large-scale organs to balance feature accuracy and computational overhead.
[0163] This embodiment significantly enhances the original depth map's ability to express surface layers, spatial positions, and local geometry by performing multi-channel encoding on depth information and extracting spatial features such as horizontal parallax, ground height, and normal angle. Multi-channel encoded depth information provides richer, more stable, and more discriminative feature input in the subsequent target recognition process, improving the accuracy and robustness of target area positioning and recognition, reducing the risk of recognition bias due to insufficient single depth information, and providing a more reliable spatial basis for ultrasound scanning path planning and probe control.
[0164] In one embodiment, the above step S30 includes:
[0165] S301, extracting global features of the image information through a convolutional neural network;
[0166] S302, extracting local geometric features of the encoded depth information through a spatial transformation network;
[0167] S303, cascadingly fusing the global features with the local geometric features to generate a multimodal fusion feature;
[0168] S304, performing context enhancement on the multimodal fusion feature based on a temporal attention mechanism to generate an enhanced fusion feature;
[0169] S305 , based on the enhanced fusion features, generating the classification probability and three-dimensional positioning frame coordinates of the target scanning area through a fully connected layer.
[0170] In this embodiment, in order to accurately identify the target scanning area within the target area, it is necessary to extract multimodal features based on the image information and the encoded depth information and perform fusion processing. First, the image information is extracted using a convolutional neural network. The convolutional neural network has a strong local receptive field and multi-level feature expression capabilities. It can extract global features containing texture, edge, shape and other information from the original image. The global features can comprehensively reflect the overall spatial layout and structural pattern of the target area.
[0171] To fully exploit the local geometric structure of the encoded depth information, a spatial transformer network can be used for feature extraction. The spatial transformer network can adaptively learn the geometric changes of local regions. Through the spatial transformer module, the depth encoding features are locally aligned and normalized, so that the extracted local geometric features can accurately reflect the detailed structure and local deformation of the target area surface.
[0172] After extracting the global features of the image and the local geometric features of the depth information, they need to be cascaded and fused to generate multimodal fusion features. This cascade fusion operation concatenates feature vectors from different sources along the feature dimension, preserving their respective information content while establishing a cross-modal joint representation that takes into account both image texture and spatial structure.
[0173] To further enhance the temporal consistency and contextual awareness of fused features, we can perform contextual enhancement on multimodal fused features based on the temporal attention mechanism. This mechanism dynamically assigns importance weights to different features based on their temporal correlations, strengthening key features and suppressing noisy features, thereby generating enhanced fused features that contain rich contextual information.
[0174] After obtaining the enhanced fusion features, the target scanning area is identified through a fully connected layer. This involves generating a classification probability for the target scanning area and the corresponding 3D positioning frame coordinates. The classification probability is used to determine whether the current feature corresponds to the target area, and the 3D positioning frame coordinates are used to determine the precise location and range of the target area in 3D space, providing an accurate basis for subsequent initial scanning point determination and path planning.
[0175] In a specific implementation, a pre-trained convolutional neural network can be used to extract features from color image information. For example, a convolutional module with a deep receptive field can be used to extract multi-scale texture and structural features. For the encoded depth information, a spatial transformer network can be introduced. A small-scale convolutional transformation is applied to the input, combined with an affine parameter prediction module, to automatically spatially normalize and align local depth features, extracting local geometric features that are rotation- and scale-invariant. During feature fusion, global image features and local depth features can be concatenated into a unified multimodal fused feature representation through a channel-wise concatenation operation. To enhance the contextual relevance of the fused features, a temporal attention mechanism based on recurrent neural network units can be introduced to establish temporal dependencies across feature sequences and highlight key feature responses through a weighted summation of attention. Finally, the enhanced fused features are input to a two-layer fully connected network. The first layer performs dimensionality reduction and feature compression, while the second layer outputs classification probabilities and 3D localization box parameters. The classification probabilities are normalized using a sigmoid or softmax function. The localization box parameters include center point coordinates and size parameters, all defined in the 3D point cloud space.
[0176] This embodiment utilizes multimodal fusion feature extraction and context enhancement based on image and encoded depth information, effectively combining two-dimensional texture information with three-dimensional structural information to improve the accuracy and robustness of target scanning area recognition. The synergistic effect of convolutional neural networks, spatial transformer networks, and temporal attention mechanisms enhances the stability and discriminative power of feature representation, ensuring the system can accurately locate target areas despite various interference conditions, such as complex backgrounds, changing lighting, and posture shifts. This provides a solid spatial and categorical foundation for subsequent automatic ultrasound scanning path planning and execution.
[0177] In one embodiment, the above step S40 includes:
[0178] S401, defining a first initial scanning point and a second initial scanning point of a target scanning area based on preset spatial structure knowledge;
[0179] S402, converting the two-dimensional coordinates of the first initial scanning point and the second initial scanning point into a first three-dimensional space coordinate point and a second three-dimensional space coordinate point in three-dimensional space point cloud data through a mapping function;
[0180] S403 : Calculate an initial scanning direction vector from the first three-dimensional space coordinate point to the second three-dimensional space coordinate point based on the first three-dimensional space coordinate point and the second three-dimensional space coordinate point.
[0181] In this embodiment, in order to determine the starting position and movement direction of the scanning task, it is necessary to further deduce the starting scanning plan in space based on the positioning results of the target scanning area in the image and depth data. First, two key initial scanning points in the target scanning area are determined based on the preset spatial structure knowledge. The preset spatial structure knowledge refers to the structural spatial cognitive information for the scanning object area pre-defined in the system modeling stage, which especially includes medical anatomical prior knowledge in medical scenarios. Medical anatomical prior knowledge can provide spatial relationships, morphological characteristics and typical scanning start and end point selection criteria between tissues and organs, so that the definition of the initial scanning point is more in line with actual anatomical logic and ultrasound imaging requirements.
[0182] After initially determining the 2D image coordinates of the first and second initial scan points, a mapping function is needed to convert these two points from the 2D image plane to 3D space. Based on a perspective projection model, this mapping function uses the known camera intrinsic parameter matrix, depth value, and pixel coordinates to infer the corresponding 3D coordinates. This accurately maps the initial points on the image plane to their actual locations in the 3D point cloud data. This mapping process ensures the accuracy of the physical position of the starting point in space, avoiding positioning errors caused by parallax, distortion, and other factors.
[0183] After completing the three-dimensional coordinate transformation, the initial scanning direction vector is calculated by calculating the vector between the first three-dimensional space coordinate point and the second three-dimensional space coordinate point. This initial scanning direction vector defines the starting trajectory of the scanning probe movement and the reference for attitude adjustment. The vector starts at the first three-dimensional space coordinate point and ends at the second three-dimensional space coordinate point. Its clear direction makes it easier for the robotic arm control system to formulate motion instructions and adjust the probe attitude based on this vector, ensuring the continuity of the scanning motion and the consistency of the imaging.
[0184] Through operation, the initial scanning point and scanning direction suitable for the automatic scanning task can be accurately derived according to the position of the target scanning area, which not only conforms to the characteristics of the target tissue structure, but also meets the technical requirements of path planning and ultrasound imaging.
[0185] In specific applications, a first initial scan point can be selected near the geometric center of the target scanning area, and a second initial scan point can be selected at the edge of the intended scanning trajectory. To improve the accuracy of the initial point, the initial scan point position can be adaptively adjusted based on the patient's anatomical features, incorporating pre-annotated data templates from a medical imaging database, to ensure coverage of the primary imaging requirements of the target area. The mapping of 2D image coordinates to 3D spatial points can be achieved using an internal parameter matrix back-projection method. Specifically, a derivation method based on homogeneous coordinate transformation can be used to inversely transform each 2D point based on its corresponding depth value to obtain the complete 3D position coordinates. If the acquisition device has external parameter matrix correction information, mapping errors can be further corrected to improve 3D spatial positioning accuracy. To calculate the initial scanning direction vector, a vector subtraction operation can be performed directly: the first 3D spatial coordinate point is subtracted from the second 3D spatial coordinate point to obtain the direction vector, which is then normalized to unit vector. The normalized direction vector is not only used for starting pose setting but also serves as the initial direction input for subsequent scanning path planning, improving motion path continuity and planning efficiency. In different application scenarios, the distance between the initial scanning points can be adjusted according to the spatial scale of the scanned object to ensure that the scanning starting point and direction meet the image coverage requirements while taking into account the continuity of the robotic arm movement and the standardized acquisition requirements of ultrasound images.
[0186] For example, the RGB image is extracted by the VGG16 network, and the depth image is extracted by AlexNet through convolution, and then the global features are extracted separately. and local features Next, the fusion module completes the integration of modal features through cascade operations:
[0187]
[0188] This formula defines how to extract features from RGB images and the encoded depth map features Splice and obtain the fused global features With local features Used for subsequent classification and positioning tasks.
[0189] in, Represents the global features extracted from RGB images; Represents the local features extracted from RGB images; Represents the global features for deep information extraction; Represents the local features extracted from depth information; concat represents the feature concatenation operation; Represents the fused multimodal features.
[0190] In order to enhance the contextual robustness of recognition, the LSTM-based global attention mechanism (GlobalContext-based Attention Model) is introduced to perform temporal perception processing on the fusion features to generate the final global feature F G In order to improve the accuracy of local target frame recognition, multiple randomly initialized spatial transformer networks (STNs) are introduced to perform geometric transformation enhancement on the local image, and then the two fully connected layers output F L The classification probability prediction formula is as follows:
[0191] p=softmax(f cls (concat(F L ,F G )))
[0192] The positioning frame offset is calculated from the local features:
[0193] t * =f bbox (F L )
[0194] The loss function adopts a typical multi-task structure:
[0195] L(p,c,t c ,v)=L cls (p,c)+λ·[c≥1]·L bbox (t c ,v)
[0196] Among them, L cls is the cross entropy loss, L bbox is the Smooth L1 loss function, and λ is the balance factor.
[0197] Based on the coordinates [x1, y1], [x2, y2] returned by the recognition frame, two initial scanning points are defined according to medical knowledge. The three-dimensional starting position P in the point cloud is obtained through the mapping function P = M(C). s1 , and its corresponding scanning direction is:
[0198]
[0199] Finally define the starting transformation matrix T s1 :
[0200]
[0201] This embodiment effectively eliminates the problems of scan path redundancy, imaging area offset, and insufficient scanning accuracy caused by improper selection of the starting position and direction in traditional methods by determining the initial scan point and direction based on the location of the target scan area. By combining pre-defined spatial structure knowledge to define the starting point, the initial scan point is ensured to conform to the anatomical characteristics of the target tissue. A mapping function accurately restores the three-dimensional spatial position, and the initial scan direction vector is calculated to determine the starting motion trend of the scan. This provides an accurate and coherent spatial reference for subsequent automatic scan path generation and dynamic adjustment, significantly improving the standardization of the scanning process and the consistency of imaging quality.
[0202] In one embodiment, the above step S50 includes:
[0203] S501, analyzing the three-dimensional coordinates of the initial scanning point;
[0204] S502, generating a motion control instruction for the end effector of the robot arm based on the three-dimensional coordinates and the initial scanning direction;
[0205] S503, controlling the end effector of the robotic arm to move to the initial scanning point according to the motion control instruction, and adjusting the scanning posture of the scanning probe according to the initial scanning direction;
[0206] S504, based on the adjusted scanning posture, starting a scanning module of the ultrasound probe, scanning along the initial scanning direction, and acquiring a real-time scanning image;
[0207] S505: Transmit the real-time scanned image to an image processing module for buffering and time stamp alignment.
[0208] In this embodiment, before the automatically controlled scanning probe begins scanning and acquiring real-time images, the 3D coordinates of the initial scanning point must first be resolved. The initial scanning point is obtained through preliminary identification and derivation. Its 3D coordinates contain a specific description of its spatial position and serve as fundamental data for robotic arm motion control. Resolving the 3D coordinates of the initial scanning point involves extracting the spatial position parameters of that point from established 3D point cloud data or spatial mapping results. This includes the representation of the position vector in a global coordinate system, which typically consists of three components corresponding to horizontal, vertical, and depth distances.
[0209] Based on the resolved 3D coordinates and initial scanning direction, motion control instructions for the robotic arm's end effector need to be generated. These instructions include target pose information, namely the target position and target attitude. The position is derived from the 3D coordinates of the initial scanning point, and the attitude is derived from the directional requirements of the initial scanning direction vector. By combining the 3D coordinates and the translation direction, a target pose description can be constructed. This is then converted into an instruction format recognizable by the robotic arm's motion controller, which is used to drive the end effector to perform precise movement and attitude adjustments.
[0210] After receiving the motion control command, the robotic arm's end effector moves to the initial scanning point along the planned path, while adjusting the scanning probe's scanning posture according to the initial scanning direction. Scanning posture adjustment refers to adjusting the relative orientation of the probe and the target surface during movement or after reaching the target position, so that the probe's ultrasound beam is incident on the target area surface vertically or at a specified angle to obtain high-quality imaging data. The scanning posture directly affects the clarity of the ultrasound image, the visibility of tissue stratification, and the effectiveness of artifact suppression. Therefore, it is necessary to precisely control the probe's orientation and pitch angle based on the initial scanning direction vector.
[0211] Based on the adjusted scanning posture, the ultrasound probe's scanning module is activated, beginning continuous scanning along the initial scanning direction, capturing ultrasound images of the target area in real time. The scanning module controls the probe's transmission and reception of ultrasound signals, while simultaneously capturing echo signals to form continuous B-mode images or other types of imaging data. During scanning, the probe maintains continuous motion along the initial direction to ensure spatial continuity and anatomical consistency between image frames.
[0212] During the acquisition of real-time scanned images, image data must be promptly transmitted to the image processing module for caching and time stamping, achieving timestamp alignment. Timestamp alignment helps subsequent data processing modules accurately match image frames with probe motion during image sequence reconstruction, posture correction, or dynamic analysis, ensuring data consistency and timing synchronization, and improving the overall system's automated scanning accuracy and imaging continuity.
[0213] In practical applications, the 3D coordinates of the initial scan point can be directly extracted by searching the previously generated 3D point cloud index to directly retrieve the spatial position data of the corresponding point. To improve parsing speed, a spatial hash table or KD tree structure can be constructed during the point cloud data preprocessing phase to quickly index the 3D coordinate information corresponding to the initial scan point. When generating motion control instructions for the robotic arm end effector, an inverse kinematics solution can be used. Combining the target pose and the current state of the robotic arm, the rotation angle change of each joint is calculated and specific joint trajectory planning instructions are generated. During this calculation process, a redundancy optimization strategy can be introduced to optimize the robotic arm's motion path while meeting pose requirements, avoiding singularities or excessive joint deflection. When controlling the robotic arm end effector to move to the initial scan point, linear interpolation or optimal trajectory tracking algorithms can be used to ensure smooth transitions in velocity and acceleration, reduce motion jitter, and maintain stable probe contact with the target surface. After reaching the initial scan point, the end effector's posture is adjusted based on the initial scan direction vector. This can be achieved through quaternion rotation or Euler angle posture adjustment to ensure that the ultrasonic probe is pointing in the direction of the target surface normal or in accordance with the preset scan angle. After starting the ultrasound probe's scanning module, the ultrasound system's continuous operating mode can be controlled, and appropriate transmission frequency, frame rate, and gain parameters can be set to ensure continuous scanning along the initial direction and real-time acquisition of high-quality images. Real-time scanned images are immediately transmitted to the image processing module after acquisition. DMA transmission acceleration technology can be used to reduce transmission delays and improve system response speed. During the timestamp alignment process, a high-precision system clock can be used to mark the acquisition time of each frame of the image, and the timestamp information is synchronously embedded in the image data header for subsequent processing modules to parse and synchronize, ensuring precise alignment of the ultrasound image sequence and the robotic arm motion data on the timeline.
[0214] Example: During automated thyroid ultrasound scanning, the system first analyzes the three-dimensional coordinates of the initial scan point. Using a mapping function, it quickly retrieves the spatial position from the point cloud data and calculates the target pose based on the identified initial scan direction vector. The system then invokes the inverse kinematics module to generate motion control commands for the robotic arm's end effector, controlling the probe's smooth movement along a predetermined path to the initial scan point. During this movement, the probe's posture is simultaneously adjusted to maintain an appropriate angle between the scanning plane and the neck surface, ensuring that the imaging direction is orthogonal to the tissue plane or meets the optimal angle of incidence. After reaching the initial point and completing posture adjustments, the system initiates the ultrasound probe's continuous scanning mode, performing a stable scan along the initial scan direction. High-resolution ultrasound images of the thyroid region are acquired in real time, and each frame is timestamped and cached synchronously in the image processing module. This precise control and timing management enables the system to continuously and standardizedly perform automated ultrasound scans of the thyroid region, effectively improving image quality and scanning efficiency while significantly reducing the risk of imaging inconsistencies due to operator experience.
[0215] This embodiment can achieve high-precision spatial positioning and dynamic alignment between the scanning probe and the target area by analyzing the three-dimensional coordinates of the initial scanning point, generating motion control instructions, and controlling the precise movement and posture adjustment of the robotic arm end effector. Through continuous scanning control and real-time image acquisition based on the initial scanning direction, the standardization of the ultrasound imaging process and the stability of image quality are effectively improved. Through real-time image transmission and timestamp alignment, the consistency and timing synchronization of image data and motion trajectory are ensured, providing a reliable foundation for subsequent automatic path adjustment, image reconstruction and diagnostic analysis, and overall improving the intelligence level and application reliability of the ultrasound automatic scanning system.
[0216] In one embodiment, the above step S70 includes:
[0217] S701, detecting the position and coverage area of the artifact region in the real-time scan image, and generating artifact region distribution parameters;
[0218] S702, calculating a pixel offset of the scanning probe based on the artifact area distribution parameter;
[0219] S703, mapping the pixel offset to three-dimensional offset coordinates in the three-dimensional space point cloud data;
[0220] S704, calculating the in-plane adjustment distance between the three-dimensional offset coordinates and the current scanning probe position;
[0221] S705, if the in-plane adjustment distance is less than a preset distance threshold, correcting the scanning posture of the scanning probe based on the three-dimensional offset coordinates;
[0222] S706 , if the in-plane adjustment distance is not less than a preset distance threshold, updating the scanning path of the scanning probe based on the three-dimensional offset coordinates;
[0223] S707 , adjusting the contact pressure between the scanning probe and the target area to a preset range based on a hybrid position and force control strategy.
[0224] In this embodiment, in order to dynamically adjust the position and posture of the scanning probe during the scanning process, it is necessary to perform real-time adjustment and control based on the identified preset target and artifact area information. First, it is necessary to detect the position of the artifact area in the real-time scanning image and its coverage area. The artifact area usually refers to the image low signal or high artifact area caused by acoustic wave obstruction, reflection abnormality or tissue interface abnormality. The position and area of the artifact area can be detected by image segmentation algorithm or brightness and texture analysis method, thereby generating artifact area distribution parameters that describe the spatial distribution of the artifact. The distribution parameters usually include information such as the center position, area size, and contour shape of the artifact area.
[0225] Based on the artifact area distribution parameters, the scanning probe's pixel offset must be calculated. This represents the two-dimensional offset difference between the current scanning probe imaging area and the preset target optimal imaging area. This offset can be calculated based on the pixel coordinate difference between the artifact center and the image center or the preset target center, further reflecting the actual required probe displacement direction and amplitude.
[0226] The calculated pixel offset is mapped to the three-dimensional point cloud data to generate three-dimensional offset coordinates. The mapping process needs to combine depth information or point cloud spatial distribution characteristics to convert the two-dimensional pixel displacement into a displacement vector in the actual space. The three-dimensional offset coordinates provide a quantitative basis for spatial position correction.
[0227] After obtaining the 3D offset coordinates, the in-plane adjustment distance between the offset and the actual position of the current scanning probe needs to be calculated. The in-plane adjustment distance is the straight-line distance between the current probe position and the desired correction position on the reference plane of the target surface. Calculating the in-plane adjustment distance helps select the appropriate adjustment strategy based on the offset, ensuring efficient and accurate adjustment operations.
[0228] If the in-plane adjustment distance is less than the preset distance threshold, it means that the deviation between the probe and the ideal scanning trajectory is small. At this time, the scanning posture of the scanning probe can be directly corrected based on the three-dimensional offset coordinates. The correction of the scanning posture includes fine-tuning the pitch angle, roll angle or yaw angle to optimize the incident direction of the ultrasound beam and improve image quality and artifact suppression effect.
[0229] If the in-plane adjustment distance is greater than or equal to the preset distance threshold, the scanning probe's scanning path needs to be updated, replanning the motion trajectory from the current position to the target position to avoid image loss or image area jumps caused by excessive offset. Updating the scanning path generates a new trajectory curve or straight line segment using the path planning algorithm to ensure a smooth transition of the probe to the target scanning area.
[0230] After completing posture correction or path update, the contact pressure between the scanning probe and the target area surface must be dynamically adjusted based on a hybrid position and force control strategy to maintain the contact pressure within a preset range. This hybrid position and force control strategy comprehensively considers position errors and contact force variations, adjusting probe motion in real time through a dual-channel control loop. This avoids the risk of image quality degradation or probe separation from the target surface due to abnormal contact pressure, further improving scanning stability and imaging continuity.
[0231] The hybrid position and force control strategy is a composite control method used in robotic control. It is used to simultaneously balance the spatial position accuracy of the robotic arm's end-effector and the regulation of the force applied during contact with the environment during operation. The basic idea of this strategy is to divide the control task into two subspaces within the task space: the position control subspace and the force control subspace. Within the position control subspace, the control system primarily ensures that the end-effector accurately moves along a preset trajectory and precisely follows the target path. In the force control subspace, the control system dynamically adjusts the magnitude and direction of the output force based on the interactive force feedback generated by contact between the end-effector and the environment, keeping the contact force within a preset range. This prevents probe damage to soft tissue or degradation of imaging quality due to poor contact. In practice, the hybrid position and force control strategy is typically based on a task framework, projecting motion commands and force feedback into two orthogonal subspaces, processing them separately through different control laws, and ultimately superimposing them to generate the overall control command. For the automatic scanning system of ultrasound probes, hybrid position and force control can adjust the contact pressure of the probe on the human body surface in real time while maintaining precise tracking of the scanning path, achieving dual optimization of imaging stability and tissue safety. It is particularly suitable for application scenarios that require stable adhesion to the surface of biological tissue for continuous imaging.
[0232] For example, define the image pixel offset e t :
[0233]
[0234] According to the current reference position P O The horizontal position in the image is determined to determine whether it is in the center area of the image (W / 3 to 2W / 3). If it is in the center area, no adjustment is required, and the pixel offset e t Take 0; if it deviates from the center, calculate P O The difference between the pixel space and the image center W / 2 is multiplied by the mapping coefficient ρ to convert it into the actual number of pixels that need to be adjusted. The mapping coefficient ρ is dynamically set based on the ultrasound image resolution, the depth sensor internal parameters, and the probe size to achieve accurate mapping from pixel space to world space, thereby guiding the probe's small-scale posture correction.
[0235] The hybrid control torque τ is defined as:
[0236]
[0237] Among them, the position error and force error are calculated as follows:
[0238] e f =(IS p )(F r -F c )
[0239] During the probe scanning process, τ represents the comprehensive control torque signal that is ultimately used to drive the end effector of the robotic arm (ultrasound probe), J represents the Jacobian matrix of the robotic arm at the current position, which is used to map the motion and force between the joint space and the Cartesian space, and J -1 Represents the inverse matrix of the Jacobian matrix, which is used to map the desired end velocity to the joint velocity command, J T Represents the transposed matrix of the Jacobian matrix, which is used to map the desired end force to joint torque.
[0240] Represents the position control proportional gain matrix, which is used to adjust the strength of the control response according to the size of the position error. Represents the position control differential gain matrix, which is used to suppress system overshoot or oscillation according to the rate of change of position error. represents the force control proportional gain matrix, which is used to adjust the probe force according to the contact force error. represents the force control differential gain matrix, which is used to improve the dynamic response of the force control according to the change rate of the contact force error.
[0241] e p Represents the position error after processing by the selection matrix, P r represents the target expected position vector, P c Represents the current actual position vector, S p represents the degree of freedom selection matrix, which is used to indicate which directions perform position control and which directions perform force control; e f Represents the force error after processing by the selected matrix, F r Denotes the target desired contact force vector, F c Represents the current actual contact force vector, and I represents the identity matrix, which is used to ensure the consistency of the degree of freedom dimensions.
[0242] Represents the derivative of the position error with respect to time, which is used to capture the dynamic trend of the actual position change. It represents the derivative of the force error with respect to time and is used to capture the dynamic trend of contact force changes.
[0243] In practical applications, the location and coverage of artifact regions can be detected using image segmentation using a convolutional neural network model, or using threshold segmentation combined with grayscale histogram analysis and local texture feature extraction. Artifact centerpoints can be extracted through centroid calculation or maximum connected component detection. Pixel offset calculations use the image processing module's built-in coordinate system to calculate the horizontal and vertical offsets between the artifact region center and the current scanned image center in real time, generating a two-dimensional displacement vector. Mapping pixel offsets to three-dimensional space can be done based on the established image-to-3D point cloud correspondence, combined with depth values or point cloud indices, to reflect the pixel offsets in the 3D point cloud space and generate the 3D offset coordinates required for practical operations. The in-plane adjustment distance can be calculated by calculating the Euclidean distance between the probe's current position and the 3D offset target position on the tangent plane. When the distance is less than a preset threshold, the probe's attitude parameters are fine-tuned by small angles, adjusting pitch, roll, or yaw to optimize the imaging direction. When the distance is greater than or equal to the threshold, the probe's motion is redirected to the desired area through local trajectory planning. Path planning can employ shortest path search, smooth trajectory interpolation, or a hybrid trajectory optimization algorithm. The hybrid position and force control strategy can adopt a force-position hybrid control framework, in which the position control channel is used to ensure the scanning trajectory tracking accuracy, and the force control channel is used to dynamically adjust the contact force between the probe and the tissue surface to prevent the probe from lifting off or excessively compressing the target area, thereby improving imaging quality and safety.
[0244] This embodiment dynamically adjusts the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and artifact area, which can effectively deal with image artifact problems caused by factors such as tissue deformation, movement, and sound wave obstruction during the scanning process, and optimize the quality of ultrasound images in real time. Through the precise mapping of pixel offset to three-dimensional space and the threshold judgment of the adjustment distance in the plane, the two types of adjustment strategies, posture fine-tuning and path replanning, are reasonably distinguished, thereby improving the accuracy and flexibility of the dynamic control of the probe. By continuously adjusting the probe contact pressure through a hybrid position and force control strategy, the dual protection of imaging stability and tissue safety is achieved, thereby significantly improving the imaging consistency, repeatability and overall diagnostic reliability during the automatic scanning process.
[0245] In one embodiment, a scanning control device based on image analysis is provided, which corresponds one-to-one to the thyroid ultrasound robot automatic scanning method based on RGB image and depth information in the above embodiment. Figure 3 , Figure 3This is a functional module diagram of a preferred embodiment of a scanning control device based on image analysis according to the present invention. It includes an image and depth information acquisition module 10, a depth information encoding module 20, a target scanning area identification module 30, an initial scan planning module 40, a scan execution and image acquisition module 50, a real-time image analysis module 60, a scanning posture and path adjustment module 70, a preset target monitoring module 80, and a scan termination control module 90. Each functional module is described in detail below:
[0246] Image and depth information acquisition module 10, used to obtain image information and depth information of the target area;
[0247] a depth information encoding module 20, configured to encode the depth information to obtain encoded depth information;
[0248] A target scanning area identification module 30 is configured to identify a target scanning area within the target area based on the image information and the encoded depth information;
[0249] An initial scanning planning module 40 is used to determine an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0250] The scanning execution and image acquisition module 50 is used to control the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and to obtain a real-time scanning image during the scanning process;
[0251] A real-time image analysis module 60 is configured to analyze the real-time scan image and identify preset targets and artifact areas in the real-time scan image;
[0252] A scanning posture and path adjustment module 70 is configured to adjust the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0253] A preset target monitoring module 80 is used to monitor the continuous existence of a preset target in the real-time scanning image;
[0254] The scanning termination control module 90 is configured to terminate the scanning if the preset target is not recognized in the continuous multiple frames of images.
[0255] In one embodiment, the image and depth information acquisition module 10 is specifically configured to:
[0256] Capture image information and depth information of the target area through a color depth camera;
[0257] Aligning the depth information with the pixel coordinate system of the image information based on a pre-calibrated parameter matrix;
[0258] Performing neighboring interpolation filling on invalid pixels in the depth information to generate complete depth information;
[0259] The complete depth information is mapped to generate three-dimensional space point cloud data through perspective projection transformation. The three-dimensional space point cloud data is a set of three-dimensional coordinate points on the surface of the target area.
[0260] In one embodiment, the depth information encoding module 20 is specifically configured to:
[0261] Calculating a horizontal disparity channel of the depth information based on a depth sensor disparity model;
[0262] Calculating a ground height channel of the depth information according to a gravity direction reference axis;
[0263] Generate a normal angle channel of the depth information based on a surface normal vector estimation algorithm;
[0264] The horizontal disparity channel, ground height channel and normal angle channel are combined into multi-channel encoded depth information.
[0265] In one embodiment, the target scanning area identification module 30 is specifically configured to:
[0266] Extracting global features of the image information through a convolutional neural network;
[0267] Extracting local geometric features of the encoded depth information through a spatial transformation network;
[0268] Cascadingly fusing the global features with the local geometric features to generate multimodal fusion features;
[0269] Performing context enhancement on the multimodal fusion features based on a temporal attention mechanism to generate enhanced fusion features;
[0270] Based on the enhanced fusion features, the classification probability and three-dimensional positioning frame coordinates of the target scanning area are generated through a fully connected layer.
[0271] In one embodiment, the initial scan planning module 40 is specifically configured to:
[0272] Based on the preset spatial structure knowledge, defining a first initial scanning point and a second initial scanning point of the target scanning area;
[0273] Converting the two-dimensional coordinates of the first initial scanning point and the second initial scanning point into a first three-dimensional space coordinate point and a second three-dimensional space coordinate point in the three-dimensional space point cloud data through a mapping function;
[0274] An initial scanning direction vector pointing from the first three-dimensional space coordinate point to the second three-dimensional space coordinate point is calculated based on the first three-dimensional space coordinate point and the second three-dimensional space coordinate point.
[0275] In one embodiment, the scanning execution and image acquisition module 50 is specifically configured to:
[0276] Analyzing the three-dimensional coordinates of the initial scanning point;
[0277] generating a motion control instruction for the end effector of the robotic arm based on the three-dimensional coordinates and the initial scanning direction;
[0278] Controlling the end effector of the robotic arm to move to the initial scanning point according to the motion control instruction, and adjusting the scanning posture of the scanning probe according to the initial scanning direction;
[0279] Based on the adjusted scanning posture, starting the scanning module of the ultrasound probe, scanning along the initial scanning direction, and acquiring a real-time scanning image;
[0280] The real-time scanned image is transmitted to the image processing module for buffering and time stamp alignment.
[0281] In one embodiment, the scanning posture and path adjustment module 70 is specifically configured to:
[0282] Detecting the position and coverage area of the artifact area in the real-time scanning image, and generating artifact area distribution parameters;
[0283] Calculating a pixel offset of the scanning probe based on the artifact area distribution parameter;
[0284] Mapping the pixel offset to three-dimensional offset coordinates in three-dimensional space point cloud data;
[0285] Calculating an in-plane adjustment distance between the three-dimensional offset coordinates and the current scanning probe position;
[0286] If the in-plane adjustment distance is less than a preset distance threshold, correcting the scanning posture of the scanning probe based on the three-dimensional offset coordinates;
[0287] If the in-plane adjustment distance is not less than a preset distance threshold, updating the scanning path of the scanning probe based on the three-dimensional offset coordinates;
[0288] Based on a hybrid position and force control strategy, the contact pressure between the scanning probe and the target area is adjusted to a preset range.
[0289] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of a thyroid ultrasound robot automatic scanning method based on RGB images and depth information.
[0290] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a method for automatic scanning of a thyroid ultrasound robot based on RGB images and depth information.
[0291] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0292] Obtain image information and depth information of the target area;
[0293] Encoding the depth information to obtain encoded depth information;
[0294] identifying a target scanning area within the target area based on the image information and the encoded depth information;
[0295] Determining an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0296] Controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process;
[0297] Analyzing the real-time scanning image to identify preset targets and artifact areas in the real-time scanning image;
[0298] Adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0299] Monitoring the continuous presence of a preset target in the real-time scanning image;
[0300] If the preset target is not recognized in the continuous multiple frames of images, the scanning is terminated.
[0301] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0302] Obtain image information and depth information of the target area;
[0303] Encoding the depth information to obtain encoded depth information;
[0304] identifying a target scanning area within the target area based on the image information and the encoded depth information;
[0305] Determining an initial scanning point and an initial scanning direction according to the position of the target scanning area;
[0306] Controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process;
[0307] Analyzing the real-time scanning image to identify preset targets and artifact areas in the real-time scanning image;
[0308] Adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area;
[0309] Monitoring the continuous presence of a preset target in the real-time scanning image;
[0310] If the preset target is not recognized in the continuous multiple frames of images, the scanning is terminated.
[0311] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0312] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0313] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0314] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A thyroid ultrasound robot automatic scanning method based on RGB image and depth information, characterized in that: The following steps are involved: Obtain image information and depth information of the target area; Encoding the depth information to obtain encoded depth information; identifying a target scanning area within the target area based on the image information and the encoded depth information; Determining an initial scanning point and an initial scanning direction according to the position of the target scanning area; Controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process; Analyzing the real-time scanning image to identify preset targets and artifact areas in the real-time scanning image; Adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area; Monitoring the continuous presence of a preset target in the real-time scanning image; If the preset target is not recognized in the continuous multiple frames of images, the scanning is terminated.
2. The thyroid ultrasound robot automatic scanning method based on RGB image and depth information according to claim 1, characterized in that: Obtain image information and depth information of the target area, including: Capture image information and depth information of the target area through a color depth camera; Aligning the depth information with the pixel coordinate system of the image information based on a pre-calibrated parameter matrix; Performing neighboring interpolation filling on invalid pixels in the depth information to generate complete depth information; The complete depth information is mapped to generate three-dimensional space point cloud data through perspective projection transformation. The three-dimensional space point cloud data is a set of three-dimensional coordinate points on the surface of the target area.
3. The thyroid ultrasound robot automatic scanning method based on RGB image and depth information according to claim 1, characterized in that: Encoding the depth information to obtain encoded depth information includes: Calculating a horizontal disparity channel of the depth information based on a depth sensor disparity model; Calculating a ground height channel of the depth information according to a gravity direction reference axis; Generate a normal angle channel of the depth information based on a surface normal vector estimation algorithm; The horizontal disparity channel, ground height channel and normal angle channel are combined into multi-channel encoded depth information.
4. The thyroid ultrasound robot automatic scanning method based on RGB image and depth information according to claim 1, characterized in that: Identifying a target scanning area within the target area based on the image information and the encoded depth information includes: Extracting global features of the image information through a convolutional neural network; Extracting local geometric features of the encoded depth information through a spatial transformation network; Cascadingly fusing the global features with the local geometric features to generate multimodal fusion features; Performing context enhancement on the multimodal fusion features based on a temporal attention mechanism to generate enhanced fusion features; Based on the enhanced fusion features, the classification probability and three-dimensional positioning frame coordinates of the target scanning area are generated through a fully connected layer.
5. The thyroid ultrasound robot automatic scanning method based on RGB image and depth information according to claim 1, characterized in that: Determining an initial scanning point and an initial scanning direction according to the position of the target scanning area includes: Based on the preset spatial structure knowledge, defining a first initial scanning point and a second initial scanning point of the target scanning area; Converting the two-dimensional coordinates of the first initial scanning point and the second initial scanning point into a first three-dimensional space coordinate point and a second three-dimensional space coordinate point in the three-dimensional space point cloud data through a mapping function; An initial scanning direction vector pointing from the first three-dimensional space coordinate point to the second three-dimensional space coordinate point is calculated based on the first three-dimensional space coordinate point and the second three-dimensional space coordinate point.
6. The thyroid ultrasound robot automatic scanning method based on RGB image and depth information according to claim 1, characterized in that: Controlling the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and acquiring a real-time scanning image during the scanning process, including: Analyzing the three-dimensional coordinates of the initial scanning point; generating a motion control instruction for the end effector of the robotic arm based on the three-dimensional coordinates and the initial scanning direction; Controlling the end effector of the robotic arm to move to the initial scanning point according to the motion control instruction, and adjusting the scanning posture of the scanning probe according to the initial scanning direction; Based on the adjusted scanning posture, starting the scanning module of the ultrasound probe, scanning along the initial scanning direction, and acquiring a real-time scanning image; The real-time scanned image is transmitted to the image processing module for buffering and time stamp alignment.
7. The method for automatic thyroid ultrasound robot scanning based on RGB images and depth information according to claim 1, characterized in that: Adjusting the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area includes: Detecting the position and coverage area of the artifact area in the real-time scanning image, and generating artifact area distribution parameters; Calculating a pixel offset of the scanning probe based on the artifact area distribution parameter; Mapping the pixel offset to three-dimensional offset coordinates in three-dimensional space point cloud data; Calculating an in-plane adjustment distance between the three-dimensional offset coordinates and the current scanning probe position; If the in-plane adjustment distance is less than a preset distance threshold, correcting the scanning posture of the scanning probe based on the three-dimensional offset coordinates; If the in-plane adjustment distance is not less than a preset distance threshold, updating the scanning path of the scanning probe based on the three-dimensional offset coordinates; Based on a hybrid position and force control strategy, the contact pressure between the scanning probe and the target area is adjusted to a preset range.
8. A thyroid ultrasound robot automatic scanning device based on RGB image and depth information, characterized in that: The scanning control device based on image analysis includes: Image and depth information acquisition module, used to obtain image information and depth information of the target area; A depth information encoding module, configured to encode the depth information to obtain encoded depth information; A target scanning area identification module, configured to identify a target scanning area within the target area based on the image information and the encoded depth information; An initial scanning planning module, configured to determine an initial scanning point and an initial scanning direction according to the position of the target scanning area; A scanning execution and image acquisition module, configured to control the scanning probe to start scanning according to the initial scanning point and the initial scanning direction, and to acquire a real-time scanning image during the scanning process; A real-time image analysis module, configured to analyze the real-time scanned image and identify preset targets and artifact areas in the real-time scanned image; A scanning posture and path adjustment module, configured to adjust the scanning posture and scanning path of the scanning probe according to the recognition results of the preset target and the artifact area; A preset target monitoring module, used to monitor the continuous existence state of the preset target in the real-time scanning image; The scanning termination control module is used to terminate the scanning if the preset target is not recognized in the continuous multiple frames of images.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image analysis-based scanning control program stored in the memory and capable of running on the processor. When the image analysis-based scanning control program is executed by the processor, the steps of the thyroid ultrasound robot automatic scanning method based on RGB images and depth information are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a scanning control program based on image analysis, and when the scanning control program based on image analysis is executed by the processor, the steps of the thyroid ultrasound robot automatic scanning method based on RGB images and depth information as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Intelligent collaborative robot system for radioexamination and control method thereof
CN120884307A
Intelligent planning method for personalized scanning path of ultrasonic robot
CN121196605A
Crane complex scene obstacle recognition method and system based on artificial intelligence
CN121305529A
Ultrasonic scanning system and method for thyroid gland, computer equipment and storage medium
CN121337400A
A thyroid ultrasound scanning system, method, computer equipment, and storage medium
CN121337400B