An ultrasonic scanning-oriented multi-modal semantic-driven robot arm autonomous learning method
By identifying and comparing anatomical differences in ultrasound image data, the robotic arm is driven to perform precise or expanded scans, solving the problems of low scanning efficiency and poor image quality in existing robotic arm systems when facing non-standard anatomical structures. This enables autonomous learning and adaptive scanning, improving diagnostic accuracy and efficiency.
Patent Information
- Application Number
- CN202511240874.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing ultrasound scanning robotic arm systems are unable to effectively identify anatomical differences when faced with patients with non-standard anatomical structures or abnormalities, resulting in low scanning efficiency, poor image quality, and even the potential to learn maladaptive scanning strategies.
By acquiring ultrasound image data, the boundaries of the target area are identified, the current structural contour information is compared with the standard structural contour information, the difference metric is calculated, and the robotic arm is driven to perform precise scanning based on the difference metric. This includes fine scanning under standard conditions and expanding the scanning range under abnormal conditions, thereby reducing the sensitivity of image quality assessment.
It significantly improves scanning efficiency and diagnostic accuracy when dealing with patients with complex or abnormal anatomical structures, avoids problems such as repeated scanning and poor image quality, and enables the robotic arm to learn autonomously and perform adaptive scanning.
Smart Images

Figure CN120791788B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mechanical arm control, and particularly relates to a multi-modal semantic-driven mechanical arm autonomous learning method for ultrasonic scanning. BACKGROUND
[0002] In the field of medical health, in order to realize the automation and standardization of ultrasonic examination, the prior art has attempted to use a mechanical arm system to assist in completing ultrasonic scanning. Such a system usually includes a high-precision mechanical arm, an end ultrasonic probe, a high-resolution visual sensor, and an internal control unit. The control unit includes a semantic processing unit, an image analysis unit, and an autonomous learning unit. In clinical operation, medical professionals will issue instructions to the system, such as "please perform a standard cross-sectional scan on the target area".
[0003] After the system receives the instruction, the semantic processing unit uses its knowledge of human anatomy and scanning experience to preliminarily infer the approximate location of the target area on the surface of the patient's body. At the same time, the high-resolution visual sensor captures the fine shape of the patient's body surface, combined with the preliminary inference of the semantic processing unit, to more accurately locate the target scanning area and guide the mechanical arm to move the ultrasonic probe to the approximate starting position. When the ultrasonic probe contacts the patient's surface, the system enters the formal scanning stage. The probe continuously acquires real-time ultrasonic images and transmits them to the image analysis unit. According to the real-time image feedback, the mechanical arm finely adjusts the probe posture, tilt angle, and applied pressure to obtain high-quality, diagnostic-value standard images. The autonomous learning unit continuously observes the relationship between the mechanical arm motion trajectory, the body surface features captured by the visual sensor, and the ultrasonic image quality, constantly optimizes the judgment logic, improves the scanning path and parameters, and improves the correspondence accuracy between external body features and internal organ actual positions, aiming to improve the efficiency and diagnostic accuracy of subsequent examinations.
[0004] However, when the system encounters a patient with significant anatomical differences or abnormalities, its limitations will be revealed. In this case, when the operator issues a standard instruction, the semantic processing process proceeds normally, and subtle differences that do not match expectations are detected. Due to this semantic and visual mismatch, the mechanical arm encounters difficulties when trying to position the probe to find the "best" scanning starting point. At this time, the autonomous learning component receives contradictory information, and it will try to adjust the probe placement position, but due to the fundamental anatomical mismatch, it is actually trying to force a standard reference onto a non-standard reality, resulting in continuous poor real-time ultrasonic image quality.
[0005] This persistent inability to obtain clear images can trigger a problem feedback loop within the ultrasound image analysis unit, further causing the autonomous learning unit to fail to recognize this as a novel or abnormal anatomical configuration and instead attempt to "learn" from these repeated "failures" to obtain a "standard" image. If this abnormal data is incorporated into the autonomous learning unit's general reference information without proper classification or weighting, it can potentially contaminate the system's overall understanding of anatomy, degrading the autonomous learning unit's performance when subsequently processing "standard" patients. Ultimately, the system becomes trapped in a cycle of attempting to impose standard anatomical references onto non-standard realities, resulting in significantly prolonged scan times, patient discomfort from repeated and often unnecessary probe adjustments, and ultimately, the failure to obtain diagnostic-quality images of the target organ. In this specific and abnormal clinical context, the "autonomous learning" mechanism, which was intended to continuously improve, instead becomes the source of errors and inefficiencies.
[0006] Therefore, in view of the above problems, the prior art needs to be improved. SUMMARY
[0007] In view of the above deficiencies of the prior art, the present application provides a multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning, aiming to solve the problem that the existing ultrasound scanning robotic arm system cannot effectively identify anatomical differences when facing patients with non-standard anatomical structures or abnormal situations, resulting in low scanning efficiency, poor image quality, and even potentially learning inappropriate scanning strategies.
[0008] In a first aspect, a multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning is provided, the method comprising the steps of:
[0009] S1: obtaining a user's scanning instruction and performing exploratory collection on a target region of a patient according to the scanning instruction to obtain ultrasound image data;
[0010] S2: identifying the boundary of the target region according to the ultrasound image data to obtain current structure contour information;
[0011] S3: retrieving pre-stored standard structure contour information according to the type of the target region;
[0012] S4: comparing the current structure contour information with the standard structure contour information to obtain a difference measure value;
[0013] S5: driving the robotic arm to perform precise scanning on the target region according to the difference measure value.
[0014] The application provides an ultrasonic scanning-oriented multimodal semantic driving mechanical arm autonomous learning method, which can effectively identify the anatomical differences of a patient and adjust a scanning strategy according to a difference measurement value, thereby overcoming the low scanning efficiency and poor image quality of a mechanical arm system in the prior art when facing non-standard anatomical structures, realizing autonomous learning and adaptive scanning of the mechanical arm, and improving the accuracy and efficiency of diagnosis.
[0015] Further, the step S5 comprises:
[0016] S51: if the difference measurement value is less than or equal to a preset allowable threshold, it is judged that the exploration type scanning result is a standard case, the mechanical arm is driven to perform fine scanning on the target region, and the probe posture and pressure are adjusted according to real-time image quality feedback;
[0017] S52: if the difference measurement value is greater than the preset allowable threshold, it is judged that the exploration type scanning result is an abnormal case, the mechanical arm is driven to expand the scanning range, and the sensitivity of image quality evaluation is reduced.
[0018] The application provides an ultrasonic scanning-oriented multimodal semantic driving mechanical arm autonomous learning method, which can intelligently judge whether the current scanning condition is standard or abnormal according to the difference measurement value of the exploration type scanning result, and adopt different scanning strategies, thereby avoiding blind fine scanning in an abnormal condition and improving the adaptability and efficiency of scanning.
[0019] Further, in the step S51, the driving of the mechanical arm to perform fine scanning on the target region and the adjustment of the probe posture and pressure according to real-time image quality feedback comprise the steps of:
[0020] S511: when the mechanical arm is driven to perform fine scanning on the target region, real-time ultrasonic image streams are acquired, and the edge sharpness, contrast and signal-to-noise ratio of each frame of image in the ultrasonic image streams are calculated, so as to obtain real-time image quality feedback;
[0021] S512: if the edge sharpness is less than a preset sharpness, the tilt angle of the probe is adjusted to scan the next frame of image;
[0022] S513: if the contrast is less than a preset contrast, the applied pressure of the probe is increased to scan the next frame of image;
[0023] S514: if the signal-to-noise ratio is less than a preset signal-to-noise ratio, the position of the probe is adjusted to scan the next frame of image;
[0024] S515: until the image quality feedback reaches a preset image quality standard.
[0025] The application provides an ultrasound scanning-oriented multi-modal semantic driving robotic arm autonomous learning method, which can finely adjust the posture, pressure and position of a probe according to real-time image quality feedback, so as to ensure that high-quality ultrasound images are obtained under standard conditions and image quality is avoided due to operation factors.
[0026] Further, in step S52, the driving the robotic arm to expand the scanning range and reduce the sensitivity of image quality evaluation includes the steps of:
[0027] S521: acquiring a difference direction and a difference value of the difference measure value relative to the preset allowable threshold;
[0028] S522: determining an expansion direction according to the difference direction and determining an expansion ratio according to the difference value;
[0029] S523: driving the robotic arm to expand the scanning range and reduce the sensitivity of image quality evaluation according to the expansion direction and the expansion ratio.
[0030] The application provides an ultrasound scanning-oriented multi-modal semantic driving robotic arm autonomous learning method, which can intelligently expand the scanning range and reduce the sensitivity of image quality evaluation according to the specific direction and value of the difference when an abnormal condition is detected, so as to avoid scanning failure due to excessive pursuit of image quality under abnormal conditions, and improve the adaptability and robustness to abnormal conditions.
[0031] Further, step S1 includes:
[0032] S11: acquiring a scanning instruction of a user and performing semantic analysis on the instruction to obtain preliminary information of a target region;
[0033] S12: acquiring three-dimensional point cloud data of a patient's torso surface and extracting macro topographic features of the torso surface based on the three-dimensional point cloud data;
[0034] S13: matching the preliminary information of the target region with the macro topographic features to estimate a projection region of the target region on the patient's body surface;
[0035] S14: determining a target region range and an initial probe position of the exploratory acquisition according to the projection region;
[0036] S15: driving a robotic arm to perform exploratory acquisition according to the target region range and the initial probe position to obtain ultrasound image data.
[0037] Further, step S13 includes:
[0038] S131: obtaining a deformable model of the target region according to the preliminary information;
[0039] S132: adjusting the deformable model according to the macroscopic topographic features, so as to minimize the difference between the deformable model and the macroscopic topographic features;
[0040] S133: estimating the projection region of the target region on the patient's body surface according to the adjusted deformable model.
[0041] Further, step S14 comprises:
[0042] S141: obtaining the shape and size of the projection region;
[0043] S142: generating an exploratory scanning path covering the projection region according to the shape and size of the projection region, so as to determine the target region range of the exploratory acquisition;
[0044] S143: taking the starting point of the scanning path as the initial probe position.
[0045] Further, step S2 comprises:
[0046] S21: regionally segmenting the ultrasound image data to divide the ultrasound image data into a plurality of regions with different acoustic characteristics;
[0047] S22: identifying a segmented region corresponding to the target region from the plurality of regions with different acoustic characteristics;
[0048] S23: extracting the external contour of the segmented region corresponding to the target region as the boundary of the target region, to obtain the current structural contour information.
[0049] Further, step S4 comprises:
[0050] S41: aligning the current structural contour information and the standard structural contour information;
[0051] S42: calculating the difference measure value according to the non-overlapping region between the aligned current structural contour information and the standard structural contour information.
[0052] Further, step S42 comprises:
[0053] S421: calculating the area of the non-overlapping region;
[0054] S422: normalizing the area of the non-overlapping region with respect to the area constituted by the standard structural contour information, to obtain the difference measure value.
[0055] Beneficial effects: The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning proposed in the present application effectively solves the problem that the robotic arm system in the prior art cannot effectively identify anatomical differences when facing patients with non-standard anatomical structures or abnormal conditions, resulting in low scanning efficiency, poor image quality, and even the learning of inappropriate scanning strategies. The scanning efficiency and diagnostic accuracy when facing patients with complex or abnormal anatomical structures are significantly improved. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 A flowchart of the multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning proposed in the present application.
[0057] Figure 2 A simple schematic diagram of the multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning proposed in the present application. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and indicated in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0059] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0060] Please refer to FIG. 1, Figure 2 A multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scanning, the method comprising the steps of:
[0061] S1: obtaining a user's scanning instruction, and performing exploratory collection on a target region of a patient according to the scanning instruction to obtain ultrasound image data;
[0062] S2: identifying the boundary of the target region according to the ultrasound image data to obtain current structure contour information;
[0063] S3: retrieve pre-stored standard structure contour information according to the type of the target region;
[0064] S4: compare the current structure contour information with the standard structure contour information to obtain a difference measure value;
[0065] S5: drive the mechanical arm to perform precise scanning on the target region according to the difference measure value.
[0066] The present application introduces a multi-modal semantic driving and autonomous learning mechanism, enabling the mechanical arm to adaptively adjust according to the actual anatomical structure differences of the patient, thereby effectively solving the problems of low scanning efficiency, poor image quality, and ineffective autonomous learning mechanism of the mechanical arm system in the prior art when facing non-standard anatomical structure patients. By comparing the current structure contour information with the standard structure contour information and driving the mechanical arm to perform precise scanning according to the difference measure value, the present application can significantly improve the accuracy and efficiency of ultrasound scanning, while avoiding unnecessary repeated scanning and discomfort to the patient.
[0067] Specifically, the method proposed in the present application aims to realize the automation and intelligentization of ultrasound scanning.
[0068] Ultrasound scanning refers to the process of using ultrasound waves to image and diagnose internal organs or tissues in the human body.
[0069] Multi-modal semantic driving refers to the system's ability to integrate information from different modalities (e.g., user's semantic instructions, three-dimensional point cloud data of the patient's body surface, real-time ultrasound image data, etc.) and guide the actions of the mechanical arm through semantic understanding.
[0070] The mechanical arm is a programmable, multi-functional mechanical device that can accurately perform pre-set or feedback-adjusted actions, and is mainly used in the present application to control the ultrasound probe for scanning.
[0071] Autonomous learning refers to the system's ability to optimize its behavior strategies and parameters through continuous interaction with the environment and receiving feedback, to improve the accuracy and efficiency of scanning.
[0072] In step S1, the user's scanning instructions can be provided to the system through voice input, text input, or graphical interface selection, etc. For example, the user can dictate "please scan the target region", or select the corresponding target region scanning option on the touch screen. After receiving the instruction, the system will analyze it to determine the target scanning region.
[0073] Exploratory acquisition refers to the preliminary and covering scanning of the target region by the robotic arm to obtain the overall ultrasound image data of the region. For example, the robotic arm can move the probe on the surface of the target region according to a preset grid path while continuously acquiring ultrasound images. The ultrasound image data can be a sequence of two-dimensional images or three-dimensional volume data, containing acoustic information of the internal structure of the target region.
[0074] In step S2, after obtaining the ultrasound image data, the data needs to be processed to identify the accurate boundary of the target region. For example, an image segmentation algorithm such as a threshold-based segmentation technique can be used to analyze the ultrasound images. The image segmentation algorithm is a mature technology in the prior art. In this application, the image segmentation algorithm is used to divide the regions with different acoustic characteristics in the image, thereby distinguishing the boundary of the target region. The identified boundary information constitutes the current structure contour information, which reflects the actual anatomical structure of the patient.
[0075] In step S3, after identifying the current structure contour information, the system needs a reference standard for comparison. This standard structure contour information is pre-stored in the system database and represents the anatomical morphology of the target region under normal or typical conditions. For example, if the target region is the liver, the system will retrieve the standard liver contour model from the database according to the "liver" type. These standard models can be obtained by statistical analysis and modeling of a large number of ultrasound image data of normal people, or can be constructed based on anatomical atlas.
[0076] In step S4, the current structure contour information is compared with the standard structure contour information, and the difference between the two is quantified. The larger the difference measure value, the greater the deviation between the current structure and the standard structure, which may indicate anatomical variation or abnormality. The specific comparison method can overlap the two, calculate the area of the non-overlapping region as a proportion of the total area of the standard structure contour information, and then quantify the difference measure value.
[0077] In step S5, according to the difference measure value obtained in step S4, the system intelligently adjusts the scanning strategy of the robotic arm. If the difference measure value is small, it indicates that the current structure is basically consistent with the standard structure, and the robotic arm can perform fine scanning to obtain high-quality diagnostic images. For example, the robotic arm can move according to a preset fine scanning path and fine-tune the probe attitude and pressure according to real-time image quality feedback (such as image clarity, contrast, signal-to-noise ratio, etc.) to ensure that the image quality is optimal.
[0078] If the difference measure value is large, indicating that the current structure and the standard structure have significant differences, the system will judge it as an abnormal situation. At this time, the mechanical arm will expand the scanning range to cover the target area that may deviate, and reduce the sensitivity of image quality evaluation. For example, the mechanical arm can expand the scanning area in the deviated direction according to the direction and value of the difference, and relax the real-time requirements for image quality, giving priority to ensuring that complete information of the target area can be captured, rather than immediately pursuing the ultimate image quality. This strategy can effectively deal with anatomical variations or abnormal situations, and avoid the system from falling into ineffective repeated scanning and adjustment.
[0079] Compared with the prior art, the advantage of the present application is that by acquiring the current structure contour information in real time and comparing it with the standard information, the scanning strategy can be adaptively adjusted according to the actual difference, which significantly improves the adaptability of the system to individual differences of patients. The autonomous learning mechanism of the present application no longer blindly learns from "failure", but adjusts based on the understanding of the actual anatomical structure difference. When a significant difference is detected, the system will adopt strategies such as expanding the scanning range, rather than just making ineffective compensatory movements, which makes the learning process more effective and targeted, and avoids the formation of "maladaptive" strategies.
[0080] Because the traditional ultrasonic scanning method may not fully consider the actual situation represented by different difference measure values when driving the mechanical arm to perform precise scanning according to the difference measure value, for example, when there is a large deviation between the current structure contour of the target area and the standard structure contour, if a single precise scanning strategy is still used, it may lead to low scanning efficiency or ineffective coverage of abnormal areas. If the above problems are not solved, the autonomous learning and adaptability of the mechanical arm will be limited when facing complex or abnormal scanning scenarios, thereby affecting the quality of ultrasonic image acquisition and diagnostic efficiency. To this end, the present application further proposes a multi-modal semantic driven mechanical arm autonomous learning method for ultrasonic scanning, which intelligently judges the exploration type scanning result as a standard situation or an abnormal situation according to the size of the difference measure value, and adopts different precise scanning strategies respectively to improve the adaptability and robustness of scanning. Specifically, step S5 comprises:
[0081] S51: If the difference measure value is less than or equal to the preset allowable threshold, the exploration type scanning result is judged as a standard situation, the mechanical arm is driven to perform fine scanning on the target area, and the mechanical arm end probe posture and pressure are adjusted according to real-time image quality feedback;
[0082] S52: If the difference measure value is greater than the preset allowable threshold, the exploration type scanning result is judged as an abnormal situation, the mechanical arm is driven to expand the scanning range, and the sensitivity of image quality evaluation is reduced.
[0083] wherein the preset allowable threshold is a pre-set numerical value used to define the range within which the difference measure value is considered to be an acceptable "standard case", and beyond which is considered to be an "abnormal case". The threshold value can be set according to the specific application scenario, the characteristics of the scanning site, and the required image accuracy.
[0084] When the difference measure value is less than or equal to the preset allowable threshold, it indicates that the exploratory scan has obtained a preliminary result that is highly consistent with the standard case, at which point the robot arm is driven to perform a refined scan. The refined scan aims to further improve the details and clarity of the image, by real-time acquisition of image quality feedback such as edge sharpness, contrast, signal-to-noise ratio, etc., to fine-tune the probe's attitude and applied pressure at the end of the robot arm, to ensure the best ultrasound image quality in the target area.
[0085] Conversely, when the difference measure value is greater than the preset allowable threshold, it indicates that the exploratory scan result deviates significantly from the standard case, which may mean that the actual morphology or location of the target area deviates greatly from the expected, or there is an abnormal structure. In this case, the robot arm is driven to expand the scanning range to ensure complete coverage of the actual target area. At the same time, in order to avoid frequent adjustments due to poor image quality in abnormal cases, the sensitivity of image quality assessment is reduced, allowing the robot arm to focus more on expanding the coverage range rather than excessively pursuing perfect quality of local images, thereby improving adaptability to abnormal cases.
[0086] In the above manner, the robot arm is driven to perform a refined scan on the target area, and the probe's attitude and pressure at the end of the robot arm are adjusted according to real-time image quality feedback. However, in actual implementation, merely mentioning "adjustment according to real-time image quality feedback" may not ensure that the ultrasound image quality meets the best diagnostic standards, and the specific image quality assessment indicators and corresponding adjustment strategies are not explicitly stated, which may result in low efficiency of the adjustment process or ineffective solution to image quality problems.
[0087] Therefore, further, in step S51, driving the robot arm to perform a refined scan on the target area, and adjusting the probe's attitude and pressure according to real-time image quality feedback includes the steps of:
[0088] S511: while driving the robot arm to perform a refined scan on the target area, real-time acquisition of an ultrasound image stream, and calculation of the edge sharpness, contrast, and signal-to-noise ratio of each frame of image in the ultrasound image stream, thereby obtaining real-time image quality feedback;
[0089] S512: if the edge sharpness is less than the preset sharpness, adjust the tilt angle of the probe to scan the next frame of image;
[0090] S513: If the contrast is less than the preset contrast, increase the applied pressure of the probe to scan the next frame of image;
[0091] S514: If the signal-to-noise ratio is less than the preset signal-to-noise ratio, adjust the position of the probe to scan the next frame of image;
[0092] S515: Until the image quality feedback reaches the preset image quality standard.
[0093] Specifically, in step S511, real-time acquisition of the ultrasound image stream means that image data is continuously received from the ultrasound device to form a continuous image sequence during the fine scanning process of the mechanical arm. The edge sharpness, contrast, and signal-to-noise ratio of each frame of image in the ultrasound image stream are calculated to quantify the quality of the current image.
[0094] Among them, the edge sharpness reflects the clarity of the object edge in the image, the contrast represents the difference degree of brightness or color in different regions of the image, and the signal-to-noise ratio measures the proportion of effective signal and noise in the image. The calculation of these indicators can be realized by image processing algorithm,
[0095] In the processing of ultrasound images, the image processing unit receives the real-time ultrasound image stream. In order to calculate the edge sharpness, the unit can be configured with a digital signal processor (DSP) which runs the Laplacian operator algorithm internally. When a frame of image data is input, the DSP performs convolution operation on the image, and the response value output reflects the intensity and clarity of the image edge, and the average value of the response value can be used as the quantitative value of the edge sharpness.
[0096] For contrast calculation, the image processing unit can use a microcontroller which executes a gray level histogram analysis algorithm. The microcontroller first generates a gray level histogram of the image, and then calculates the standard deviation of the gray level values in the histogram, which can be used as the quantitative value of the contrast.
[0097] For the estimation of signal-to-noise ratio, the image processing unit can include a special noise analysis module which estimates the noise variance by analyzing the pixel value fluctuation of the background region of the image, and at the same time calculates the signal variance of the entire image, and then performs ratio operation on the two to obtain the signal-to-noise ratio.
[0098] These calculation results are transmitted to the mechanical arm control unit in real time as the basis for adjusting the probe attitude and pressure. In this way, a comprehensive real-time image quality feedback can be obtained.
[0099] The specific adjustment strategy performed according to the real-time image quality feedback is: when the edge sharpness is less than the preset sharpness, it indicates that the image edge is blurred, which may be caused by poor contact or improper angle of the probe with the body surface. At this time, the sound beam incidence direction can be optimized by adjusting the tilt angle of the probe, thereby improving the edge definition. The adjustment amount of the tilt angle is set in advance by the person skilled in the art, for example, the tilt angle of the probe is driven by a stepping motor to 0.5 degrees. After the mechanical arm performs this tilt action, the system continues to acquire the next frame of ultrasound image, and re-evaluates the edge sharpness of the image to verify the adjustment effect.
[0100] When the contrast is less than the preset contrast, it may mean that the sound wave penetration is insufficient or the echo difference of the tissue is not obvious. By increasing the applied pressure of the probe, the contact between the probe and the body surface can be improved, the interference of air between the probe and the body surface is reduced, and the sound wave penetration and echo signal are enhanced, thereby improving the contrast. The adjustment amount of the applied pressure is set in advance by the person skilled in the art, for example, from 10 newtons to 15 newtons. The increase of the pressure makes the probe and the skin surface fit more closely, and squeezes out the small bubbles or air layer that may exist between the probe and the skin.
[0101] When the signal-to-noise ratio is less than the preset signal-to-noise ratio, the noise is large, which may be related to improper position of the probe or environmental interference. At this time, the position of the probe can be adjusted to find the best sound window, reduce noise interference, and improve signal quality. For example, the mechanical arm can move the probe several millimeters to the head side or foot side along the patient's body surface, or several millimeters to the medial side or lateral side, to try to avoid the rib bone shadow, intestinal gas and other areas that may cause signal attenuation or noise increase.
[0102] These adjustments are made for the next frame of image, forming a closed-loop feedback control. And the adjustment process continues until the real-time image quality feedback (i.e. the comprehensive indicators of edge sharpness, contrast, signal-to-noise ratio, etc.) reaches the preset image quality standard. The preset image quality standard can be set according to the specific diagnostic needs.
[0103] Further, in step S52, driving the mechanical arm to expand the scanning range and reduce the sensitivity of image quality evaluation includes steps:
[0104] S521: obtaining a difference direction and a difference value of the difference measure value relative to the preset allowable threshold;
[0105] S522: determining an expansion direction according to the difference direction, and determining an expansion ratio according to the difference value;
[0106] S523: driving the mechanical arm to expand the scanning range and reduce the sensitivity of image quality evaluation according to the expansion direction and the expansion ratio.
[0107] Specifically, in step S521, after comparing the current structure profile information with the standard structure profile information, not only the difference degree (i.e. the difference measure value) between the two is obtained, but also the specific manifestation of the difference needs to be further analyzed. For example, by analyzing the geometric center, the principal axis direction or the maximum deviation point of the non-overlapping area, the offset direction (difference direction) of the current structure profile relative to the standard structure profile and the offset distance or area size (difference value) can be determined. The purpose is to provide accurate guidance information for subsequent scanning range adjustment.
[0108] In step S522, the difference information obtained in the previous step is converted into a scanning strategy of the mechanical arm. For example, if the difference direction indicates that the target area has deviated in a certain specific direction (such as upward, downward, left or right), the expansion direction will be set to the offset direction. At the same time, the size of the difference value will directly affect the expansion ratio, the larger the difference value, the greater the degree of deviation, and the expansion ratio will also increase accordingly to ensure that the actual target area can be covered. For example: the expansion direction of the mechanical arm is determined to be the lower right; according to the difference value, the expansion ratio is determined, for example, set to increase the scanning range by 20% in the lower right area.
[0109] Thus, in step S523, according to the determined expansion direction and expansion ratio, the scanning path and range are adjusted accordingly. At the same time, the sensitivity of the image quality evaluation is reduced, specifically, the threshold of edge sharpness, contrast or signal-to-noise ratio is appropriately relaxed according to the actual specific situation to allow the mechanical arm to continue searching in the expanded area, even if the initial image quality is poor. It can be accepted until the complete target area is found and covered. The degree of relaxation is gradually adjusted by the technician during use. In this way, even if the target organ position deviates or the morphology is abnormal, the mechanical arm can efficiently and accurately complete the scanning task.
[0110] Further, step S1 includes:
[0111] S11: Obtain the scanning instruction of the user, and perform semantic analysis on the instruction to obtain preliminary information of the target area;
[0112] S12: Obtain three-dimensional point cloud data of the surface of the patient's torso, and extract macro-topographic features of the surface of the torso based on the three-dimensional point cloud data;
[0113] S13: Match the preliminary information of the target area with the macro-topographic features to estimate the projection area of the target area on the patient's body surface;
[0114] S14: Determine the target area range and the initial probe position of the exploration type collection according to the projection area;
[0115] S15: According to the target region range and the initial probe position, drive the mechanical arm to perform exploratory acquisition to obtain ultrasound image data.
[0116] The preliminary information can include typical anatomical position, size, and shape of the target region.
[0117] The depth camera scans the patient's body surface, and the three-dimensional point cloud data contains the precise geometric information of the patient's torso surface. Based on the three-dimensional point cloud data, the macro topographic features of the torso surface are extracted, which refers to identifying and quantifying feature points or regions with anatomical significance on the body surface, such as rib edges, xiphoid, navel, iliac bone, etc. These feature points can be used as reference benchmarks for subsequent matching.
[0118] The projection region is a two-dimensional or three-dimensional geometric region indicating the range of the body surface that the ultrasound probe should cover.
[0119] Specifically, taking the liver as the target region, the user's scan instruction can be obtained through various ways such as voice input, text input, or graphical interface selection. When the system receives the advanced scan instruction issued by the medical professional (for example, "standard scan of the liver"), its control core can be a high-performance embedded processor (for example, a multi-core processor based on ARM Cortex-A series, such as NVIDIA Jetson Xavier NX), which is responsible for receiving and analyzing the instruction. The semantic understanding unit can be a pre-configured natural language processing module that converts the instruction into machine-executable task parameters. This module queries a pre-stored anatomical structure information library that contains the average size, shape of common organs (such as the liver), and the approximate position range relative to the body surface landmarks (such as the costal arch and xiphoid) under standard body type. For example, for "liver scan", the system will obtain the projection region of the liver on the right upper abdomen.
[0120] At the same time, the visual sensor array can use multiple high-resolution stereo vision cameras (for example, Intel RealSense D435i) to obtain three-dimensional point cloud data of the patient's body surface based on the triangulation principle. These point cloud data can be processed to reconstruct a detailed geometric model of the patient's body surface.
[0121] The preliminary judgment of the semantic unit (for example, the approximate rectangular area of the liver on the body surface) is matched with the body surface model reconstructed by the visual system. For example, the system identifies landmark points such as the lower edge of the costal arch, the xiphoid process, etc. on the body surface model, and calculates the initial contact point (and the initial probe position) of the ultrasound probe on the patient's body surface and the approximate scanning area (i.e. the target area range) according to the relative position relationship between these landmark points and the preset liver. The motion controller of the mechanical arm (for example, a six-axis mechanical arm controlled by an EtherCAT bus-based servo driver) will accurately move the ultrasound probe (for example, a linear array or convex array probe) above this preliminary positioning area according to these calculation results, and prepare to make contact. This step is the starting point of the system's real contact with the individual patient, and is to deal with the potential differences between "standard anatomical expectations" and "individual anatomical realities", and to avoid subsequent "false attribution".
[0122] Further, step S13 comprises:
[0123] S131: Obtain a deformable model of the target area according to the preliminary information;
[0124] S132: Adjust the deformable model according to the macro-topographic features, so that the difference between the deformable model and the macro-topographic features is minimized;
[0125] S133: Estimate the projection area of the target area on the patient's body surface according to the adjusted deformable model.
[0126] The deformable model can be a pre-defined three-dimensional model that can be deformed according to external data, such as a Free-Form Deformation (FFD) model. Free-Form Deformation (FFD) is a computer graphics and image processing technique that deforms an object by defining a control point grid in three-dimensional space. Moving these control points can non-rigidly change the shape of the object inside the grid without directly manipulating the geometric vertices of the object. In this application, the model aims to capture the typical morphology of the target area and allow it to deform within a certain range to adapt to individual differences in patients.
[0127] Specifically, the macro-topographical features are extracted from the three-dimensional point cloud data of the patient's torso surface, reflecting the geometric shape and curvature information of the body surface. The adjustment process can be achieved through an optimization algorithm, such as the Iterative Closest Point (ICP) algorithm. The Iterative Closest Point (ICP) algorithm is a widely used algorithm for three-dimensional point cloud registration. The basic idea is to find the best transformation (including rotation and translation) between two point sets through iteration, so that the distance metric between the points on one point set and the nearest points on the other point set is minimized. In each iteration, the algorithm first finds the nearest point in the target point set for each point in the source point set as the corresponding point, then calculates the best transformation according to these corresponding points, and applies the transformation to update the source point set, repeating this process until convergence. In this scheme, the best geometric transformation between the current structure contour point set and the standard structure contour point set is found through iteration, so that the two are accurately aligned in space. This ensures the accuracy of the subsequent difference metric calculation, so that the difference value can truly reflect the essential deviation of the anatomical morphology, rather than the position or attitude deviation caused by non-essential factors. The purpose of this adjustment is to make the deformable model as consistent as possible with the actual macro-topographical features in space, so as to accurately locate the target region on the patient's body surface.
[0128] Once the deformable model is successfully adjusted and aligned with the macro-topographical features, the adjusted model can accurately represent the relative position and morphology of the target region in the patient's body. Thus, the projection area of the target region on the patient's body surface can be obtained by calculating the two-dimensional projection of the adjusted model on the patient's body surface.
[0129] Further, step S14 includes:
[0130] S141: Obtain the shape and size of the projection area;
[0131] S142: Generate an exploration scanning path covering the projection area according to the shape and size of the projection area, to determine the target region range of the exploration acquisition;
[0132] S143: Take the starting point of the scanning path as the initial probe position.
[0133] Wherein, the projection area is estimated after matching the preliminary information of the target region with the macro-topographical features of the patient's torso surface. The shape can refer to the geometric contour of the projection area, which can be a rectangle, an ellipse, or an irregular polygon, etc.; the size can refer to the length, width, area, or perimeter of the projection area, etc. These information provides basic data for subsequent scanning path generation.
[0134] The scanning path is designed to ensure that the probe at the end of the robotic arm can systematically and comprehensively cover the surface projection range of the target region during the exploration acquisition. For example, the scanning path can be designed as a spiral, zigzag, or grid shape to adapt to different shapes and sizes of the projection area. For example, if the projection area is approximately rectangular, the system can choose to generate a grid-shaped scanning path, with the probe moving back and forth at a fixed row and column distance, ensuring comprehensive coverage within the rectangular area. For circular or elliptical projection areas, a spiral scanning path can be generated, with the probe moving in a spiral from the center outward or from the outside inward. These path types are generated by the control unit and converted into motion instructions for the robotic arm to drive the probe at the end of the robotic arm to move along the preset trajectory. By generating this scanning path, the target region range of the exploration acquisition is clearly defined, guiding the preliminary scanning action of the robotic arm.
[0135] The scheme of the present application provides a quantitative basis for subsequent scan planning by first accurately obtaining the shape and size of the projection area of the target region on the surface of the patient. On this basis, by generating an exploration scanning path covering the projection area, it is ensured that the robotic arm can systematically and comprehensively cover the entire target region during the preliminary acquisition, avoiding missed or repeated scanning. At the same time, the starting point of the scanning path is used as the initial probe position, so that the robotic arm can start the exploration acquisition task from a clear and pre-set starting point, thereby optimizing the motion planning of the robotic arm and laying a foundation for subsequent acquisition of ultrasound image data.
[0136] Further, step S2 comprises:
[0137] S21: region segmentation of the ultrasound image data, dividing the ultrasound image data into multiple regions with different acoustic characteristics;
[0138] S22: identifying the segmentation region corresponding to the target region in the multiple regions with different acoustic characteristics;
[0139] S23: extracting the external contour of the segmentation region corresponding to the target region as the boundary of the target region, obtaining the current structure contour information.
[0140] The ultrasound image data can be input to an image processing module, which is configured to execute a region segmentation algorithm. Through the image segmentation algorithm, the ultrasound image data is divided into multiple discrete regions, and the pixel points in each region share similar acoustic characteristics, such as the same pixel gray value. Different acoustic characteristics reflect the performance of different tissue types or structures in the ultrasound image.
[0141] In a specific application, a classifier can be pre-trained to determine whether a segmented region represents a target region according to its shape, size or location. For example, a round cyst can be distinguished from an irregular tumor by shape features; a too small or too large noise region can be excluded by size features; a location feature can be used to locate in combination with anatomical prior knowledge. The comprehensive use of these features improves the accuracy of target region recognition and ensures the effectiveness of subsequent processing, thereby solving the technical problem of accurately identifying a target region from a complex ultrasound image.
[0142] Once the segmented region corresponding to the target region is determined, its external contour will be extracted as the boundary of the target region. The extraction of the external contour can utilize standard image processing techniques, such as the commonly used Sobel operator edge detection algorithm. The extracted external contour can be represented as a series of coordinate points. The current structure contour information is the basis for subsequent comparison with standard structure contour information.
[0143] By the above technical solution, the accuracy and robustness of target region boundary recognition in an ultrasound image can be significantly improved.
[0144] Further, step S4 includes:
[0145] S41: align the current structure contour information and the standard structure contour information;
[0146] S42: calculate a difference measure value according to the non-overlapping region between the aligned current structure contour information and the standard structure contour information.
[0147] Specifically, aligning the current structure contour information and the standard structure contour information means that the two are as much as possible to coincide in spatial position through geometric transformation (such as translation, rotation, scaling, etc.). This alignment process aims to eliminate non-essential differences caused by factors such as scan starting position, angle or patient position, to ensure that the subsequent comparison is based on the morphological differences of the structure itself rather than the positional differences. The alignment method can use the Iterative Closest Point (ICP) algorithm. After alignment, a difference measure value is calculated according to the non-overlapping region between the aligned current structure contour information and the standard structure contour information. The non-overlapping region refers to the part that does not coincide between the two contours after alignment. The purpose of calculating this non-overlapping region is to quantify the degree of deviation between the current scan result and the standard morphology. For example, the non-overlapping region can be calculated by geometric methods or distance measure-based methods, thereby obtaining a numerical value representing the degree of difference.
[0148] The scheme of the present application effectively eliminates the deviation caused by non-structural factors such as the initial position, angle or patient position of the mechanical arm during the scanning process by aligning the current structure contour information with the standard structure contour information, thereby ensuring the accuracy of the subsequent difference calculation. It is precisely due to this preprocessing that the subsequent difference measurement value based on the non-overlapping area can truly reflect the morphological changes of the target area, rather than just the spatial position shift. By quantifying the non-overlapping area, the deviation between the current scanning result and the preset standard can be intuitively and quantitatively evaluated, thereby providing a reliable basis for subsequent precise scanning or scanning range adjustment of the mechanical arm.
[0149] Further, step S42 comprises:
[0150] S421: calculating the area of the non-overlapping area;
[0151] S422: normalizing the area of the non-overlapping area with the area constituted by the standard structure contour information to obtain the difference measurement value.
[0152] In step S421, after aligning the current structure contour information and the standard structure contour information, the non-overlapping part between them is determined, and the two-dimensional space size of this part is quantified. For example, the area of the non-overlapping area can be accurately measured by pixel counting, geometric calculation or image processing algorithm.
[0153] Further, in step S422, the purpose of normalization processing is to eliminate the influence of different sizes of target areas, so that the difference measurement value has comparability. For example, the area of the non-overlapping area can be divided by the total area constituted by the standard structure contour information, thereby obtaining a ratio value between 0 and 1, which is the difference measurement value. This normalization processing ensures that regardless of the absolute size of the target area, the difference measurement value can accurately reflect the relative deviation between the current structure and the standard structure.
[0154] The scheme of the present application can provide an effective method for quantifying the difference between the current structure contour information and the standard structure contour information by calculating the area of the non-overlapping area and performing normalization processing. The area of the non-overlapping area intuitively reflects the degree of inconsistency between the two, and the normalization processing makes the difference measurement value not affected by the absolute size of the target area, thereby ensuring that the difference measurement value can accurately and objectively reflect the precision of the mechanical arm scanning and the actual situation of the target area in different scanning scenarios.
[0155] To sum up, the application can realize autonomous learning and adaptive scanning of the mechanical arm, significantly improving the scanning efficiency and diagnostic accuracy when facing complex or abnormal anatomical structure patients. It avoids the repeated adjustment and scanning failure caused by the mismatch between semantics and vision in the prior art, reduces the discomfort of the patient, and ensures that the final ultrasound image is obtained. In addition, through the identification and adaptive adjustment of the difference, the autonomous learning mechanism of the application can avoid learning "maladaptive" strategies, thereby maintaining the accuracy of the system's understanding of anatomy and improving its overall performance in subsequent processing of various patients.
[0156] The above merely describes the embodiments of the application and is not intended to limit the protection scope of the application. For those skilled in the art, the application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. An ultrasound scan-oriented multi-modal semantic-driven robotic arm autonomous learning method, characterized in that, The method comprises steps of: S1: obtaining a user's scanning instruction, and performing an exploratory scan on a target region of a patient according to the scanning instruction to obtain ultrasound image data; S2: identifying a boundary of the target region according to the ultrasound image data to obtain current structure contour information; S3: calling pre-stored standard structure contour information according to a type of the target region; S4: comparing the current structure contour information with the standard structure contour information to obtain a difference measurement value; S5: driving a mechanical arm to perform a precise scan on the target region according to the difference measurement value; Step S5 comprises: S51: if the difference measurement value is less than or equal to a preset allowable threshold, judging that the exploratory scan result is a standard case, driving the mechanical arm to perform a refined scan on the target region, and adjusting a mechanical arm end probe posture and pressure according to real-time image quality feedback; Step S51 comprises: S511: while driving the mechanical arm to perform the refined scan on the target region, obtaining an ultrasound image stream in real time, and calculating an edge sharpness, a contrast, and a signal-to-noise ratio of each frame of image in the ultrasound image stream to obtain the real-time image quality feedback; S512: if the edge sharpness is less than a preset sharpness, adjusting an inclination angle of the probe to scan a next frame of image; S513: if the contrast is less than a preset contrast, increasing an applied pressure of the probe to scan the next frame of image; S514: if the signal-to-noise ratio is less than a preset signal-to-noise ratio, adjusting a position of the probe to scan the next frame of image; S515: until the image quality feedback reaches a preset image quality standard.
2. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 1, wherein, Step S5 further comprises: S52: if the difference measurement value is greater than the preset allowable threshold, judging that the exploratory scan result is an abnormal case, driving the mechanical arm to expand a scan range and reduce a sensitivity of image quality evaluation.
3. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 2, wherein, In step S52, the driving of the mechanical arm to expand the scan range and reduce the sensitivity of the image quality evaluation comprises steps of: S521: obtaining a difference direction and a difference value of the difference measurement value relative to the preset allowable threshold; S522: determining an expansion direction according to the difference direction and determining an expansion ratio according to the difference value; S523: driving the mechanical arm to expand the scan range and reduce the sensitivity of the image quality evaluation according to the expansion direction and the expansion ratio.
4. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 1, wherein, Step S1 comprises: S11: obtaining a user's scanning instruction, and performing semantic analysis on the instruction to obtain preliminary information of a target region; S12: obtaining three-dimensional point cloud data of a patient torso surface, and extracting macroscopic topographic features of the patient torso surface based on the three-dimensional point cloud data; S13: matching the preliminary information of the target region with the macroscopic topographic features to estimate a projection region of the target region on the patient torso surface; S14: determining a target region range and an initial probe position of the exploratory scan according to the projection region; S15: driving the mechanical arm to perform the exploratory scan according to the target region range and the initial probe position, to obtain the ultrasound image data.
5. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 4, wherein, Step S13 comprises: S131: obtaining a deformable model of the target region according to the preliminary information; S132: adjusting the deformable model according to the macroscopic topographic features, so as to minimize the difference between the deformable model and the macroscopic topographic features; S133: estimating the projection region of the target region on the surface of the patient's torso according to the adjusted deformable model.
6. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 4, wherein, Step S14 comprises: S141: obtaining the shape and size of the projection region; S142: generating an exploratory scan path covering the projection region according to the shape and size of the projection region, to determine the target region range of the exploratory scan; S143: taking the starting point of the scan path as the initial probe position.
7. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 1, wherein, Step S2 comprises: S21: regionally segmenting the ultrasound image data, to divide the ultrasound image data into multiple regions with different acoustic characteristics; S22: identifying a segmented region corresponding to the target region from the multiple regions with different acoustic characteristics; S23: extracting the external contour of the segmented region corresponding to the target region as the boundary of the target region, to obtain the current structure contour information.
8. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 1, wherein, Step S4 comprises: S41: aligning the current structure contour information and the standard structure contour information; S42: calculating the difference measure value according to the non-overlapping region between the aligned current structure contour information and the standard structure contour information.
9. The multi-modal semantic-driven robotic arm autonomous learning method for ultrasound scan according to claim 8, wherein, Step S42 comprises: S421: calculating the area of the non-overlapping region; S422: normalizing the area of the non-overlapping region with respect to the area constituted by the standard structure contour information, to obtain the difference measure value.
Citation Information
Patent Citations
Ultrasound diagnostic apparatus and control method of ultrasound diagnostic apparatus
US20250107772A1