Automatic positioning method of ultrasonic probe based on visual recognition

By combining macroscopic semantic coarse localization with contact visual servo fine localization, the problem of adaptability of automatic ultrasound probe localization methods to individual differences and soft tissue deformation is solved, and rapid and robust probe initial localization and high-quality image acquisition are achieved.

CN122163254APending Publication Date: 2026-06-09JIANGSHAN PEOPLES HOSPITAL (JIANGSHAN PEOPLES HOSPITAL MEDICAL COMMUNITY JIANGSHAN OCCUPATIONAL DISEASE PREVENTION & CONTROL HOSPITAL)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSHAN PEOPLES HOSPITAL (JIANGSHAN PEOPLES HOSPITAL MEDICAL COMMUNITY JIANGSHAN OCCUPATIONAL DISEASE PREVENTION & CONTROL HOSPITAL)
Filing Date
2026-04-02
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing automatic positioning methods for ultrasound probes rely on high-precision pre-calibration and rigid registration, which cannot adapt to individual differences and soft tissue deformation. They also lack real-time closed-loop feedback under contact conditions, resulting in decreased positioning accuracy and poor image quality.

Method used

It adopts a two-layer architecture that combines macroscopic semantic coarse localization with contact visual servo fine localization. It identifies anatomical semantic regions through macroscopic vision and uses contact visual servo for closed-loop control to adjust the probe pose in real time to adapt to individual differences and soft tissue deformation.

Benefits of technology

It achieves rapid and robust initial probe positioning, improves positioning accuracy and image quality, simplifies operation steps, reduces system costs, and ensures good contact between the probe and the skin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122163254A_ABST
    Figure CN122163254A_ABST
Patent Text Reader

Abstract

This invention relates to the field of medical device automation technology, and in particular to an automatic ultrasound probe positioning method based on visual recognition. The method includes: acquiring macroscopic visual information containing a target body surface region; identifying at least one anatomical semantic region corresponding to the target body surface region based on the macroscopic visual information; controlling a robotic arm to move the ultrasound probe to an initial positioning position determined based on the anatomical semantic region; acquiring microscopic visual information of the contact area between the ultrasound probe and the body surface; extracting visual features reflecting the probe-body surface contact state based on the microscopic visual information; and generating control commands based on the difference between the visual features and the preset target state, and driving the robotic arm to adjust the ultrasound probe's pose. This invention employs a two-layer architecture combining macroscopic semantic coarse positioning and contact visual servo fine positioning. It does not rely on complex pre-calibration, can adapt to individual differences and soft tissue deformation, and can achieve adaptive adjustment after contact, thus achieving rapid, robust, and adaptive initial probe positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical equipment automation technology, and in particular to an automatic positioning method for ultrasound probes based on visual recognition. Background Technology

[0002] Ultrasound imaging is widely used in clinical diagnosis and screening due to its advantages such as no radiation, real-time operation, and low cost. With the development of automation technology, automatic ultrasound probe scanning systems (such as fully automated breast ultrasound and thyroid scanning robots) have gradually become a research and application hotspot. The primary key technology of such systems lies in how to automatically and accurately position the ultrasound probe to the initial position of the target scanning area.

[0003] Existing visual recognition-based automatic ultrasound probe positioning methods mainly suffer from the following problems: Relying on high-precision pre-calibration and rigid registration: Mainstream methods usually require obtaining a high-precision 3D model of the patient's body surface in advance through a 3D scanner (such as structured light or laser scanning), and then using complex iterative nearest point (ICP) and other registration algorithms with the real-time point cloud data collected by the vision system and the pre-stored model during each scan to calculate the position where the probe should be placed. This process is cumbersome and time-consuming, has extremely high requirements for the consistency of the patient's position, and the system hardware is expensive. Poor adaptability to individual differences and soft tissue deformation: There are significant individual differences in human body surface morphology, subcutaneous fat thickness, and target organ location; soft tissue will deform when the probe touches the skin; existing positioning methods based on rigid body surface feature points (such as the tip of the Adam's apple and the midpoint of the clavicle) will have a serious decrease in positioning accuracy when the above differences and deformations occur, or even lead to positioning failure. Lack of real-time closed-loop feedback under contact conditions: Most methods end after the vision system guides the probe to the predetermined spatial coordinates, which is an "open-loop" control; however, the actual contact quality between the probe and the skin (such as the uniformity of the coupling agent and the fit between the probe surface and the skin) directly affects the quality of the ultrasound image; existing technologies cannot make fine adjustments based on the real-time contact state after the probe makes contact to optimize the acoustic coupling effect.

[0004] Therefore, to address the above problems, we propose an automatic ultrasound probe positioning method based on visual recognition. Through a two-layer architecture that combines macroscopic semantic coarse positioning with contact visual servo fine positioning, it does not rely on complex pre-calibration, can adapt to individual differences and soft tissue deformation, and can achieve adaptive adjustment after contact, thus realizing fast, robust and adaptive probe initial positioning. Summary of the Invention

[0005] In order to overcome the problems of existing ultrasonic probe positioning methods, such as reliance on high-precision pre-calibration and rigid registration, poor adaptability to individual differences and soft tissue deformation, and lack of real-time closed-loop feedback under contact conditions.

[0006] The technical solution of this invention is: an automatic positioning method for an ultrasound probe based on visual recognition, comprising the following steps: S1: Macroscopic visual positioning step: acquire macroscopic visual information containing the target body surface area, identify at least one anatomical semantic region corresponding to the target body surface area based on the macroscopic visual information; control the robotic arm to drive the ultrasound probe to move to the initial positioning position determined based on the anatomical semantic region; S2: Contact Visual Servo Positioning Step: After the ultrasound probe contacts the body surface, acquire microscopic visual information of the contact area between the ultrasound probe and the body surface; based on the microscopic visual information, extract visual features reflecting the contact state between the probe and the body surface; based on the difference between the visual features and the preset target state, generate control commands and drive the robotic arm to adjust the pose of the ultrasound probe so that the visual features approach the preset target state.

[0007] Preferably, in step S1, the anatomical semantic region is a two-dimensional or three-dimensional region defined by anatomical landmarks on the surface of the larynx, used to cover or be adjacent to the target anatomical structure.

[0008] Preferably, in step S2, the contact visual servo positioning step is a closed-loop control process, wherein the visual features serve as feedback signals, and the control commands are used to adjust the position of the ultrasonic probe in a plane parallel to the contact surface and / or the tilt angle of the ultrasonic probe relative to the contact surface.

[0009] Preferably, in step S2, the visual features reflecting the contact state between the probe and the body surface include at least one of the following: motion features calculated based on the skin texture of the contact area, morphological features calculated based on the morphology of the coupling agent in the contact area, and symmetry features calculated based on the geometric relationship between the probe edge and the skin interface.

[0010] Preferably, the motion features calculated based on the skin texture of the contact area are used to generate position adjustment commands parallel to the plane of the contact surface; the symmetry features calculated based on the geometric relationship between the probe edge and the skin interface are used to generate commands to adjust the tilt angle of the ultrasound probe.

[0011] Preferably, in step S2, the contact force information between the ultrasound probe and the body surface is also acquired simultaneously; the generation of the control command is also based on the comparison result of the contact force information and the preset contact force range.

[0012] Preferably, in step S2, the process of generating control commands and driving adjustments based on the difference between visual features and preset target states continues until a preset convergence condition is met. The convergence condition includes the visual features remaining stable over multiple consecutive control cycles.

[0013] Preferably, in step S1, after controlling the robotic arm to move the ultrasound probe to the initial positioning position, the ultrasound probe is controlled to slowly contact the body surface in a preset manner.

[0014] Preferably, in step S1, the macroscopic visual information includes depth information; the initial positioning position determined based on the anatomical semantic region is calculated based on the plane and center position fitted by the three-dimensional spatial point cloud corresponding to the anatomical semantic region.

[0015] Preferably, in step S1, the identification of anatomical semantic regions based on macroscopic visual information is achieved by inputting the macroscopic visual information into a pre-trained semantic segmentation neural network model, and the output of the model is a map representing the probability of different anatomical semantic regions.

[0016] The beneficial effects of this invention are: 1. The method of the present invention eliminates the cumbersome individualized 3D modeling and registration process, and can be started with only a monocular or RGB-D camera for rapid semantic recognition, which greatly simplifies the operation steps and shortens the preparation time; 2. In terms of robustness and adaptability, the method of the present invention uses semantic "regions" that are insensitive to deformation and individual differences instead of precise "points" on a macroscopic level, and utilizes visual servoing to adapt to soft tissue deformation and skin sliding after contact in real time on a microscopic level, so that the system can cope with patients of various body types and different body positions. 3. In terms of positioning quality and adaptability, the contact vision servo closed loop of the present invention can actively optimize the fit between the probe and the skin and the distribution of coupling agent, directly ensuring the image quality of the initial scanning point from the physical interaction level, which is something that traditional open-loop positioning methods cannot achieve. 4. In terms of safety, the method of the present invention avoids excessive mechanical pressure by integrating compliant contact and force sensation control. Attached Figure Description

[0017] Figure 1 The diagram shows the flowchart of the automatic positioning method for ultrasonic probes based on visual recognition according to the present invention. Detailed Implementation

[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0019] Example 1 Please see Figure 1This invention provides an embodiment of an automatic ultrasound probe positioning method based on visual recognition, comprising the following steps: S1: Macroscopic visual positioning steps: acquire macroscopic visual information containing the target body surface area, identify at least one anatomical semantic region corresponding to the target body surface area based on the macroscopic visual information; control the robotic arm to drive the ultrasound probe to move to the initial positioning position determined based on the anatomical semantic region; S2: Contact Visual Servo Positioning Steps: After the ultrasound probe contacts the body surface, acquire the microscopic visual information of the contact area between the ultrasound probe and the body surface; based on the microscopic visual information, extract the visual features that reflect the contact state between the probe and the body surface; based on the difference between the visual features and the preset target state, generate control commands and drive the robotic arm to adjust the pose of the ultrasound probe so that the visual features approach the preset target state.

[0020] Furthermore, in step S1, the anatomical semantic region is a two-dimensional or three-dimensional region defined by anatomical landmarks on the surface of the larynx, used to cover or be adjacent to the target anatomical structure.

[0021] Specifically, the anatomical semantic region in step S1 refers to a two-dimensional image region or three-dimensional spatial region defined by visible or inferable anatomical landmarks (such as muscle outlines and bone protrusions) on the body surface of the larynx, neck, or chest, used to cover or be adjacent to target anatomical structures (such as the thyroid gland or breast quadrant). Unlike the localization of a single precise feature point (such as a coordinate point) in the prior art, this invention localizes a region with tolerance space. Among them, a lightweight neural network or traditional image processing algorithm is used to segment stable semantic regions such as "sternocleidomastoid muscle-anterior cervical triangle" and "cricoid cartilage horizontal band" from the macroscopic RGB or RGB-D image. Compared with a single feature point, these regions have better stability and recognizability when there are individual morphological changes and slight posture deviations. This reduces the requirements for patient positioning standardization and feature point clarity, and improves the robustness and success rate of the coarse localization stage.

[0022] Furthermore, in step S2, the contact visual servo positioning step is a closed-loop control process, wherein visual features serve as feedback signals, and control commands are used to adjust the position of the ultrasonic probe in a plane parallel to the contact surface and / or the tilt angle of the ultrasonic probe relative to the contact surface.

[0023] Specifically, the contact vision servo positioning step in step S2 is a dynamic closed-loop control process. In this process, the features extracted in real time from the microscopic visual information are used as the system's feedback quantity and compared with the internally set ideal contact state (target value). The deviation is calculated by the control algorithm (such as proportional-integral-derivative PID algorithm or model predictive control MPC) to generate the speed or position increment command of the robotic arm end in Cartesian space. This control loop is driven directly based on the intuitive manifestation of physical interaction of contact vision features, rather than pursuing a preset, static spatial coordinate. This achieves online dynamic optimization of the probe pose and can automatically compensate for the effects of soft tissue deformation, skin slippage, and initial coarse positioning errors, ensuring that the probe always maintains a good contact state.

[0024] Furthermore, in step S2, the visual features reflecting the probe-skin surface contact state include at least one of the following: motion features calculated based on the skin texture of the contact area, morphological features calculated based on the coupling agent morphology of the contact area, and symmetry features calculated based on the geometric relationship between the probe edge and the skin interface; wherein: Skin texture motion characteristics: By performing optical flow calculations on microscopic contact area images of consecutive frames, the motion vector field of skin texture is obtained; its mean and variance reflect whether there is relative slippage between the probe and the skin; this feature is mainly used to drive the movement of the probe in the contact surface plane (XY direction) to eliminate slippage; Coupling agent morphology characteristics: By segmenting the image, the uniformity and continuity of the area and width of the crescent-shaped region formed by the ultrasonic coupling agent at the probe edge are extracted; the ideal coupling agent morphology should be a continuous and uniform bright band; this feature can be used to determine the tightness of the contact and whether the coupling agent distribution is uniform. Interface symmetry feature: The boundary lines formed by the left and right sides of the probe and the skin are extracted by edge detection, and the curvature, angle or distance from the center line of the two boundary lines is calculated to determine the symmetry. This feature directly reflects whether the probe axis is aligned with the normal direction of the skin surface, and is used to drive the rotation adjustment (pitch and roll) of the probe around the X and Y axes.

[0025] Furthermore, the motion characteristics calculated based on the skin texture of the contact area are related to the relative sliding speed between the probe and the skin. When the motion characteristic value is greater than zero, the control algorithm will generate a translation command opposite to the direction of the average motion vector, driving the probe to track the skin, thereby eliminating slippage, until the motion characteristic approaches zero. The symmetry characteristics calculated based on the geometric relationship between the probe edge and the skin interface are used to generate a command to adjust the tilt angle of the ultrasound probe. For example, when the gap on the left side of the interface is greater than that on the right side, it indicates that the probe is tilted to the right. The control algorithm will generate a small rotation command around the Y-axis (or the corresponding axis) to straighten the probe to the left, until the symmetry characteristic values ​​on both sides are equal or close.

[0026] Furthermore, in step S2, the contact force information between the ultrasound probe and the body surface is also acquired simultaneously (through a six-dimensional force sensor); the generation of control commands is also based on the comparison results of the contact force information and the preset contact force range; for example, when the visual characteristics have stabilized but the contact force is lower than the optimal range, the probe can be controlled to press down slightly (moving in the negative Z direction); otherwise, it can be raised slightly; thus realizing multimodal fusion control of force and vision, ensuring the safety of operation while ensuring contact quality, and avoiding excessive pressure that may cause patient discomfort.

[0027] Furthermore, in step S2, the process of generating control commands based on the difference between visual features and preset target states and driving the adjustment continues until the preset convergence conditions are met. The convergence conditions include, but are not limited to, all monitored visual feature values ​​being stable within their respective target threshold ranges for N consecutive control cycles. This provides a clear and reliable basis for terminating the servo process and prevents the system from entering an oscillation or infinite adjustment state.

[0028] Furthermore, in step S1, after the robotic arm drives the ultrasonic probe to the initial positioning position, the ultrasonic probe is controlled to contact the body surface perpendicularly or along the surface normal direction at a preset slow speed and constant small contact force; thus achieving a compliant and safe initial contact, avoiding the risk of the probe hitting the body surface at high speed, and providing a stable starting point for subsequent visual servo adjustments.

[0029] Furthermore, in step S1, the macroscopic visual information includes depth information; the initial positioning position determined based on the anatomical semantic region is obtained by fitting a plane using the Random Sample Consensus (RANSAC) algorithm based on the three-dimensional spatial point cloud corresponding to the anatomical semantic region, and calculating the center or centroid of the point cloud; the initial positioning position is usually set at a certain distance above the plane, and the initial direction of the probe is roughly aligned with the normal direction of the plane; the depth information provides a rough but reasonable spatial orientation and height for coarse positioning, making subsequent approach and contact movements more efficient and safe.

[0030] Furthermore, in step S1, the identification of anatomical semantic regions based on macroscopic visual information is achieved by inputting the macroscopic visual information into a pre-trained semantic segmentation neural network model (such as U-Net, DeepLab series). This model is trained using a large number of medical surface images labeled with different anatomical semantic regions, and its output is a probability map of the same size as the input image. The value of each pixel represents the confidence level of its belonging to a specific anatomical semantic region. By utilizing the powerful feature extraction and generalization capabilities of deep learning, the target semantic region can be robustly and accurately segmented from complex and varied surface images.

[0031] When the system is in operation, after it is started, it acquires macroscopic images of the patient's neck through a top-down camera, and the semantic segmentation network identifies semantic regions such as the "anterior cervical triangle" and calculates the initial approach pose. After the robotic arm controls the probe to move safely to this position, it gently contacts the skin along the normal direction; at the moment of contact, the side-viewing miniature camera is activated and begins to analyze the skin texture optical flow, coupling agent morphology and edge symmetry at the contact edge; The servo controller calculates translation and rotation compensation commands in real time based on the deviations between these characteristic values ​​and the ideal values, and sends them to the robotic arm. The probe is adjusted at the micrometer level under servo drive until the texture is still, the coupling agent is uniform, and the edges are symmetrical. At this point, the system determines that the optimal contact state has been reached, and the positioning is completed. The entire process, from identification to precise positioning, can be completed automatically within tens of seconds without human intervention.

[0032] Through the above steps, the method of this invention eliminates the cumbersome process of individualized 3D modeling and registration, requiring only a monocular or RGB-D camera for rapid semantic recognition to start, greatly simplifying the operation and shortening the preparation time. In terms of robustness and adaptability, macroscopically, semantic "regions" that are insensitive to deformation and individual differences are used instead of precise "points," and microscopically, visual servoing is used to adapt to soft tissue deformation and skin sliding after contact in real time, enabling the system to cope with patients of various body types and different body positions. In terms of positioning quality and adaptability, the contact visual servoing closed loop can actively optimize the fit between the probe and the skin and the distribution of coupling agent, directly ensuring the image quality of the initial scanning point from the physical interaction level, which is impossible to achieve with traditional open-loop positioning methods. In terms of safety, compliant contact and force fusion control avoid excessive mechanical pressure.

[0033] Example 2 Optionally, this embodiment provides a basic method for automatic positioning of an ultrasound probe based on visual recognition, specifically including the following steps: Step S101: System Initialization and Preparation Place the patient on the scanning bed and adjust their position (e.g., supine with neck slightly extended) so that the target scanning area (e.g., the anterior cervical region) is exposed to the field of view of the macroscopic vision device; the system starts up and performs a self-check, including the robotic arm returning to zero and the calibration and verification of the macroscopic camera (e.g., an RGB-D camera) and the microscopic camera integrated on the probe side; Step S102: Macroscopic visual localization – information acquisition and semantic region recognition A macroscopic camera is controlled to acquire macroscopic RGB-D images containing the target body surface region of the patient; the acquired RGB images are input into a pre-trained semantic segmentation neural network model; the model outputs a probability heatmap of at least one target anatomical semantic region; for example, for a thyroid scan, the target semantic region can be the "anterior cervical triangle"; this region is not a point, but a two-dimensional image region that is approximately triangular in shape and bounded by the anterior borders of the two sternocleidomastoid muscles and the midline of the neck; the recognition process is to find the set of pixels in the image that belong to this region; Step S103: Macroscopic visual localization – Calculate initial pose and move Based on the semantic region pixel set obtained in step S102, and combined with the synchronously acquired depth image, the corresponding 3D spatial point cloud is extracted. Use the Random Sample Consensus (RANSAC) algorithm from the point cloud Fit an optimal plane And calculate point cloud center point Initial positioning position Set in plane normal vector In terms of direction, distance from the center point Above A pose at a location, and the direction of the probe axis is... Roughly aligned; then, the robotic arm is controlled to plan a collision-free path to move the ultrasonic probe to the desired position. ; Step S104: Contact Visual Servo Localization – Initial Contact and Feature Acquisition Control the robotic arm along The probe is driven to contact the skin at a low speed and with a constant contact force threshold. Once the probe contacts the skin (triggered by a six-dimensional force sensor at the end of the robotic arm), the microscopic camera on the probe side (such as a 2-megapixel, 60fps miniature CMOS camera) and its ring LED illumination are activated, and continuous acquisition of microscopic images of the probe-skin contact edge area begins. ,in The frame number; Step S105: Contact visual servo positioning – visual feature extraction For each frame of microscopic image Visual features reflecting the contact state are extracted; this embodiment extracts two types of features: Skin texture motion features For two consecutive frames of images and The grayscale image was used to calculate the dense optical flow field using the Farneback optical flow method; motion characteristics. Defined as the average value of the optical flow vector magnitude, the calculation formula is:

[0034] in, To calculate the total number of pixels in the region, For the first Optical flow vector of 1 pixel; It characterizes the relative sliding speed between the probe and the skin; Probe edge symmetry characteristics Image extraction using the Canny edge detection algorithm The boundary lines formed by the left and right sides of the probe and the skin are calculated; the average distances from the left and right boundary lines to the theoretical center line of the probe are calculated respectively. and Symmetry characteristics Defined as:

[0035] It characterizes the alignment deviation between the probe axis and the local normal direction of the skin; Step S106: Contact vision servo positioning – servo control and pose adjustment The preset target state is: motion characteristics. (No slippage), symmetry characteristics (Completely symmetric); based on current features With target features To address the deviation, a proportional-integral (PI) control algorithm is used to generate adjustment speed commands for the robotic arm's end effector in the task space. ; Define deviation , ; The control commands are calculated as follows:

[0036]

[0037]

[0038] in, and This refers to the translational speed command of the probe in the X and Y directions within the contact plane; This is the angular velocity command for the probe around the Y-axis (pitch axis); The direction angle of the average optical flow vector in the current frame; and For the proportional and integral gain coefficients of the corresponding control loop; control commands It is sent to the robotic arm controller for execution, driving the probe to perform real-time pose fine-tuning; Step S107: Iteration and Convergence Judgment Repeat steps S105 and S106 to form a closed-loop control of "image acquisition - feature extraction - deviation calculation - instruction generation - drive adjustment - new image acquisition"; simultaneously, at each step, determine the convergence condition: if continuous Frames All less than the threshold (like (pixels / frame) and All less than the threshold (like If the number of pixels is less than or equal to the number of pixels, the system is considered to have reached the optimal contact state, and servo positioning is complete; otherwise, iterative adjustments continue.

[0039] Example 3 Optionally, this embodiment further optimizes the robustness of macroscopic positioning and the feature dimensions of microscopic servoing based on embodiment 2.

[0040] In step S102, the semantic segmentation network model outputs a probability heatmap of multiple anatomical semantic regions. Taking thyroid localization as an example, it simultaneously identifies three semantic regions: "anterior border of the left sternocleidomastoid muscle", "anterior border of the right sternocleidomastoid muscle", and "horizontal reference zone of the cricoid cartilage". The final "target body surface region" is the logical intersection or weighted combination of these three regions. This method can more accurately define the thyroid body surface projection area, and is especially suitable for patients with short or thick necks. In step S103, the plane is fitted and the center is calculated based on this more accurate combined region point cloud; In step S105, in addition to extraction and New features: Coupling agent morphology characteristics Microscopic images Threshold segmentation and morphological operations were performed to extract the "crescent-shaped" bright spot region formed by the ultrasound coupling agent; the area of ​​this region was then calculated. and the area of ​​its left half and the area of ​​the right half Define morphological features Indicators of area uniformity:

[0041] Simultaneously, the continuity of the coupling agent region is calculated (e.g., whether there is a break); the target state is... And continuous; In step S106, the controller's deviation vector is expanded to... ,in Correspondingly, the generation of control commands also takes into account The impact; for example, when At that time, it may generate tiny rotation commands around the Z-axis (roll axis). To compensate for uneven distribution of coupling agent; this embodiment introduces more semantic regions and richer contact visual features, making the positioning process more adaptable and more accurate to complex anatomical structures and contact conditions.

[0042] Example 4 Optionally, based on Embodiment 2 or 3, this embodiment adds force information fusion and a more intelligent convergence judgment mechanism to further improve the security and automation of positioning.

[0043] During the contact process in step S104 and subsequent servo adjustments, the contact force between the probe and the skin, measured by the six-dimensional force sensor, is continuously and in real time read. and ; In step S106, the generated control command is based not only on visual feature deviation but also incorporates force feedback. A force control sub-objective is defined: maintaining the component of the contact force in the normal direction. Stable within the expected range Internal (e.g., [4N, 8N]); the control law is modified as follows:

[0044] in, This is the translational velocity command of the probe along the normal direction (Z direction) of the contact surface. The desired contact force (e.g., 6N); when the vision servo tends to stabilize but When it is low, If positive, the probe is slightly depressed; otherwise, it is slightly raised. At this point, the complete control command is: ; In step S107, the convergence condition is upgraded to a multimodal condition: simultaneously satisfying (a) visual feature stability (same as in Example 2), and (b) contact force. continuous Frame stabilized Within the range, and (c) contact torque The pressure is less than the set threshold; the positioning is considered complete only when all conditions are met; this embodiment achieves precise control of contact pressure while optimizing the contact state through force-vision fusion control, avoiding patient discomfort caused by poor coupling due to insufficient pressure or excessive pressure, and ensuring the comprehensive optimization and high reliability of the positioning results through more stringent convergence judgment.

[0045] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. An automatic positioning method for ultrasound probes based on visual recognition, characterized in that: Includes the following steps: S1: Macroscopic visual positioning step: acquire macroscopic visual information containing the target body surface area, identify at least one anatomical semantic region corresponding to the target body surface area based on the macroscopic visual information; control the robotic arm to drive the ultrasound probe to move to the initial positioning position determined based on the anatomical semantic region; S2: Contact visual servo positioning step: After the ultrasound probe contacts the body surface, acquire the microscopic visual information of the contact area between the ultrasound probe and the body surface; based on the microscopic visual information, extract visual features that reflect the contact state between the probe and the body surface; Based on the difference between the visual features and the preset target state, control commands are generated and the robotic arm is driven to adjust the position of the ultrasound probe so that the visual features approach the preset target state.

2. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S1, the anatomical semantic region is a two-dimensional or three-dimensional region defined by anatomical landmarks on the surface of the larynx, used to cover or be adjacent to the target anatomical structure.

3. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S2, the contact visual servo positioning step is a closed-loop control process, wherein the visual features serve as feedback signals, and the control commands are used to adjust the position of the ultrasonic probe in a plane parallel to the contact surface and / or the tilt angle of the ultrasonic probe relative to the contact surface.

4. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 3, characterized in that: In step S2, the visual features reflecting the contact state between the probe and the body surface include at least one of the following: motion features calculated based on the skin texture of the contact area, morphological features calculated based on the morphology of the coupling agent in the contact area, and symmetry features calculated based on the geometric relationship between the probe edge and the skin interface.

5. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 4, characterized in that: The motion features calculated based on the skin texture of the contact area are used to generate position adjustment commands parallel to the plane of the contact surface; the symmetry features calculated based on the geometric relationship between the probe edge and the skin interface are used to generate commands to adjust the tilt angle of the ultrasound probe.

6. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S2, the contact force information between the ultrasound probe and the body surface is also acquired simultaneously; the generation of the control command is also based on the comparison result between the contact force information and the preset contact force range.

7. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S2, the process of generating control commands and driving adjustments based on the difference between visual features and preset target states continues until a preset convergence condition is met. The convergence condition includes the visual features remaining stable over multiple consecutive control cycles.

8. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S1, after controlling the robotic arm to move the ultrasound probe to the initial positioning position, the ultrasound probe is controlled to slowly contact the body surface in a preset manner.

9. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S1, the macroscopic visual information includes depth information; the initial positioning position determined based on the anatomical semantic region is calculated based on the plane and center position fitted by the three-dimensional spatial point cloud corresponding to the anatomical semantic region.

10. The automatic positioning method for an ultrasonic probe based on visual recognition according to claim 1, characterized in that: In step S1, the identification of anatomical semantic regions based on macroscopic visual information is achieved by inputting the macroscopic visual information into a pre-trained semantic segmentation neural network model, and the output of the model is a map representing the probability of different anatomical semantic regions.