A multimodal perception and ai-driven intelligent ultrasound detection system and method
By using a multimodal perception and AI-driven intelligent ultrasound detection system, combined with multiple sensors and algorithms, the problems of low organ positioning accuracy and imprecise pressure control in traditional ultrasound examinations have been solved, achieving high-precision, safe, and reproducible standardized examinations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-09
AI Technical Summary
Traditional ultrasound examinations rely on the operator's experience, resulting in low accuracy in organ localization, imprecise control of pressure, and a lack of standardization in the examination process, making it difficult to achieve high-precision, high-safety, and reproducible standardized examinations.
The intelligent ultrasound detection system, which employs multimodal perception and AI-driven technology, combines a depth camera, an infrared scanner, an inertial measurement unit, and a thin-film pressure sensor. Through a Kalman filter fusion processor and a PID control algorithm, it achieves real-time organ localization, path planning, and force adjustment, ensuring image clarity.
It has achieved full automation and standardization of the ultrasound examination process, improved the accuracy of organ positioning and the precision of pressure control, and ensured the safety of the examination and the consistency of image quality.
Smart Images

Figure CN122163252A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent medical testing technology, and in particular relates to a multimodal perception and AI-driven intelligent ultrasound testing system and method. Background Technology
[0002] The quality of ultrasound examinations is highly dependent on the experience level of the operator. Traditional ultrasound examinations involve medical staff moving a handheld probe across the patient's body surface, judging the accuracy of the examination site and the appropriateness of the pressure applied by observing real-time images. This operating mode has the following shortcomings: First, the accuracy of organ localization is limited by the operator's anatomical knowledge and ability to interpret images in real time, making it difficult to accurately track deep organs or minute lesions; second, the pressure applied by the probe depends entirely on the operator's feel, with too little pressure resulting in blurry images and too much pressure potentially causing patient discomfort or even injury; third, the examination procedure lacks standardization, making it difficult to reproduce the examination path and techniques used by different operators, or even the same operator at different times, affecting the consistency of diagnosis.
[0003] While some existing assistive ultrasound devices have incorporated robotic arms or image processing technology, they mostly focus on optimizing a single function, such as only achieving automatic probe clamping or only providing image enhancement processing. They fail to systematically solve the synergistic problems of multiple aspects, such as dynamic organ positioning, adaptive pressure adjustment, and standardized examination pathways, making it difficult to meet the clinical demand for high-precision, high-safety, and reproducible standardized ultrasound examinations. Summary of the Invention
[0004] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides a multimodal perception and AI-driven intelligent ultrasound detection system and method. Through the synergistic integration of multimodal perception fusion, AI image diagnosis, path planning and force feedback control, the entire process of ultrasound examination is automated and standardized, solving the problems of traditional ultrasound examination, such as strong dependence on operator skills, low organ positioning accuracy and imprecise pressure control.
[0005] Technical Solution: To achieve the above objectives, the present invention provides a multimodal perception and AI-driven intelligent ultrasonic detection system, comprising:
[0006] The base platform serves as the foundation for the system, on which a movable slide bed is installed;
[0007] A multi-degree-of-freedom robotic arm system is mounted on the base platform, with an ultrasound probe detachably connected to its end, for driving the ultrasound probe to move on the human body surface;
[0008] The multimodal sensing unit includes a depth camera and an infrared scanner fixed on a base platform, and an inertial measurement unit (IMU) that can be fixed to the surface of the subject's body; the depth camera is used to acquire 3D point clouds of the human torso, the infrared scanner is used to acquire the coordinates of the human body surface, and the inertial measurement unit (IMU) is used to capture the subject's motion posture.
[0009] The force sensing unit includes a thin-film pressure sensor disposed between the end of the multi-degree-of-freedom robotic arm system and the ultrasonic probe, for real-time detection of the contact pressure between the ultrasonic probe and the human body;
[0010] The main control console is electrically connected to the ultrasonic probe, the multi-degree-of-freedom robotic arm system, the multimodal sensing unit, and the force sensing unit; the main control console integrates:
[0011] An ultrasound image acquisition module, used to receive the raw ultrasound signals acquired by the ultrasound probe and generate digitized ultrasound image data; and
[0012] The central control module is configured as follows:
[0013] a) Fusion localization: Receives 3D point cloud data from the depth camera and motion data from the inertial measurement unit (IMU), fuses them through Kalman filtering, estimates the subject's skeletal posture in real time, and calculates the real-time position of internal organs based on pre-registered anatomical atlas.
[0014] b) Path planning: Receive the human body surface coordinates from the infrared scanner, combine them with the real-time position of the organ, and plan a collision-free examination path for the ultrasound probe;
[0015] c) Force-position coordinated control: The multi-degree-of-freedom robotic arm system is controlled to move along the inspection path with the ultrasonic probe, while receiving pressure data from the force sensing unit. According to the preset target contact force range, the pressure of the ultrasonic probe is adjusted through a PID control algorithm.
[0016] d) Intelligent diagnostic feedback: Receives digitized ultrasound image data from the ultrasound image acquisition module, identifies lesion features and judges image clarity in real time through the EfficientNet model, and dynamically adjusts the target contact force based on clarity feedback until the image is clear.
[0017] Furthermore, it includes a slide drive mechanism for driving the movable slide to move longitudinally on the base platform.
[0018] Furthermore, the multi-degree-of-freedom robotic arm system includes: a longitudinal moving guide rail disposed on the base platform, a gantry mounted on the longitudinal moving guide rail, a transverse moving guide rail disposed below the gantry, a telescopic rod disposed on the transverse moving guide rail, and a main robotic arm mounted at the end of the telescopic rod; the telescopic rod can extend and retract in the vertical direction.
[0019] Furthermore, it also includes a quick-change disc, which includes a quick-change disc mother disc and a quick-change disc plate that can be locked and released by a locking mechanism. The quick-change disc mother disc is fixedly connected to the end of the main robotic arm, the ultrasonic probe is fixedly connected to the quick-change disc plate, and the thin-film pressure sensor is disposed on the assembly extrusion contact surface of the quick-change disc mother disc and the quick-change disc plate.
[0020] The number of quick-change trays is multiple, and the lower part of each quick-change tray is connected to an ultrasonic probe of a different frequency; the lower part of each quick-change tray is also provided with a hanging block with a shape that is larger at the top and smaller at the bottom, and the inner side wall of the base platform is provided with a side wall support frame, and the quick-change tray can be suspended on the side wall support frame by the hanging block.
[0021] Furthermore, the assembly extrusion contact surface is an annular end face where the quick-change disc mother plate and the quick-change disc plate are fitted together in the locked state. The thin-film pressure sensor is embedded in this annular end face and is used to directly sense the pressure transmitted from the ultrasonic probe to the quick-change disc plate and then to the quick-change disc mother plate.
[0022] Furthermore, it also includes an automatic glue application unit, which comprises:
[0023] A side-mounted robotic arm fixed to the inner wall of the base platform;
[0024] A peristaltic pump is fixed on the base platform, with one end of the pump's feed tube connected to a coupling agent tank and the other end connected to the side-mounted robotic arm.
[0025] The side-mounted robotic arm is used to clamp the material tube and automatically applies coupling agent to the surface of the ultrasonic probe after the ultrasonic probe is switched.
[0026] A method for intelligent ultrasonic detection control in a multimodal sensing and AI-driven intelligent ultrasonic detection system includes the following steps:
[0027] Step S1: Acquire 3D point cloud of the subject's torso using a depth camera, acquire motion data of the subject using an inertial measurement unit (IMU), fuse point cloud and IMU data using a Kalman filter, estimate the subject's skeletal posture in real time, and infer the real-time position of internal organs based on pre-registered anatomical atlas.
[0028] Step S2: Obtain the coordinates of the human body surface using an infrared scanner, and combine them with the real-time position of the organs to plan a collision-free examination path for the ultrasound probe;
[0029] Step S3: During the movement of the ultrasound probe along the examination path, the ultrasound image acquisition module acquires digital ultrasound images in real time, and the trained EfficientNet model is used to judge the image clarity and identify lesion features.
[0030] Step S4: Based on the preset target contact force range of the organ being examined and the image clarity judgment result in step S3, adjust the pressing force applied by the main robotic arm through the PID control algorithm so that the image is clear and the contact force does not exceed the safety threshold.
[0031] Furthermore, the specific process of Kalman filter fusion in step S1 includes:
[0032] IMU data is used as input to the prediction model to obtain the predicted attitude;
[0033] The observation pose is obtained by iteratively registering the point cloud acquired by the depth camera with the point cloud of the previous moment using the nearest point ICP.
[0034] The predicted attitude and the observed attitude are fused using Kalman gain to obtain the corrected attitude estimate. The truncated symbolic range field (TSDF) method is then used to fuse the current point cloud into the global 3D model to generate a surface mesh.
[0035] Furthermore, the PID control algorithm described in step S4 incorporates image sharpness feedback, specifically as follows:
[0036] Preset target contact force F target ;
[0037] Real-time acquisition of the filtered force value F from the thin-film pressure sensor filt (k), calculation error e(k) = F target -F filt (k);
[0038] If the current ultrasound image clarity is below the preset threshold, then gradually increase F within the allowable force range. target Continue until the image is clear;
[0039] The control quantity u(k) is calculated using the PID formula and mapped to the joint motor of the main robotic arm to adjust the pressing force.
[0040] Simultaneously, the pressure depth of the ultrasonic probe is monitored by an infrared scanner. If the pressure exceeds the set upper limit, the driving force is immediately withdrawn.
[0041] A storage medium storing an executable program, which, when executed by a processor, enables an intelligent ultrasonic detection control method for a multimodal sensing and AI-driven intelligent ultrasonic detection system.
[0042] Beneficial effects: This invention uses a depth camera and IMU to correct organ misalignment in real time, ensuring the probe is always aligned with the target area and improving tracking accuracy. Combined with the human body's curved surface coordinates planned by an infrared scanner, a standardized examination path is planned, ensuring process reproducibility. Simultaneously, a thin-film pressure sensor directly senses contact pressure, which is precisely adjusted via a PID algorithm. Using ultrasound image clarity as feedback, the pressure is automatically fine-tuned until clear when the image is blurry, while the infrared scanner simultaneously monitors and stops if the pressure depth exceeds the limit. This dual closed-loop control of image quality and pressure threshold ensures both safety and optimal image quality, overcoming the limitations of traditional examinations that rely on operator feel. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the overall structure of the intelligent ultrasonic testing system;
[0044] Figure 2 This is a schematic diagram of the main structure of an intelligent ultrasonic testing system.
[0045] Figure 3 This is a schematic diagram of the quick-change master disk structure;
[0046] Figure 4 A schematic diagram of the quick-change tray structure;
[0047] Figure 5 This is a schematic diagram of the automatic glue application unit. Detailed Implementation
[0048] The invention will now be further described with reference to the accompanying drawings.
[0049] like Figure 1 and Figure 2 As shown, a multimodal sensing and AI-driven intelligent ultrasound inspection system includes: a base platform 1, which serves as the foundation for the system, on which a movable slide bed 10 is mounted. The base platform 1 provides a stable mounting base for the entire system, and the movable slide bed 10 is used to carry the examinee and can move longitudinally on the base platform 1 to move the examinee into or out of the inspection area.
[0050] A multi-degree-of-freedom robotic arm system 2 is mounted on the base platform 1, with an ultrasound probe 3 detachably connected to its end, used to drive the ultrasound probe 3 to move across the human body surface. This robotic arm system possesses multi-degree-of-freedom motion capabilities, enabling precise positioning in three-dimensional space and movement along a planned path, achieving comprehensive scanning of various parts of the human body.
[0051] The multimodal sensing unit includes a depth camera 4 and an infrared scanner 5 fixed on a base platform 1, and an inertial measurement unit (IMU) that can be fixed to the subject's body surface. The depth camera 4 is used to acquire 3D point cloud data of the human torso, the infrared scanner 5 is used to acquire the human body surface coordinates, and the IMU is used to capture the subject's motion posture. Specifically, the depth camera 4 uses structured light to project a coded pattern onto the human body surface and receives reflected light. It calculates the depth value of each pixel based on the degree of pattern deformation, thereby generating 3D point cloud data of the human torso. This point cloud data accurately reflects the subject's external contours and surface features. The infrared scanner 5 emits infrared light onto the human body surface and measures the reflection time, calculating the depth information of each pixel. Combined with the camera's internal parameters, it generates a point cloud set of human body surface coordinates for subsequent path planning. The IMU, through its built-in accelerometer and gyroscope, captures motion posture data at fixed points in real time using a high-frequency sampling method, including three-axis acceleration and three-axis angular velocity, to sense the subject's micro-movements or changes in body position during the detection process.
[0052] The force sensing unit includes a thin-film pressure sensor 11 disposed between the end of the multi-degree-of-freedom robotic arm system 2 and the ultrasonic probe 3, for real-time detection of the contact pressure between the ultrasonic probe 3 and the human body. The thin-film pressure sensor 11 has a very small thickness and can be directly embedded in the assembly contact surface, realizing direct measurement of contact pressure and avoiding the influence of interference factors such as joint friction and inertial force of the robotic arm on the pressure measurement.
[0053] The main control console 6, electrically connected to the ultrasonic probe 3, the multi-degree-of-freedom robotic arm system 2, the multimodal sensing unit, and the force sensing unit, serves as the system's control center. The main control console 6 integrates: an ultrasonic image acquisition module, used to receive the raw ultrasonic signals acquired by the ultrasonic probe 3 and generate digitized ultrasonic image data; and a central control module, configured as follows:
[0054] a) Fusion Localization: The system receives 3D point cloud data from the depth camera 4 and motion data from the inertial measurement unit (IMU). Through Kalman filtering, it fuses the data to estimate the subject's skeletal posture in real time and infers the real-time positions of internal organs based on a pre-registered anatomical atlas. Specifically, the depth camera 4 provides high-precision static point cloud data at a lower frame rate (e.g., 30Hz), while the IMU provides continuous posture change data at a higher frame rate (e.g., 200Hz). The Kalman filter uses the IMU data as a state prediction model and the point cloud registration results from the depth camera 4 as observation updates. Through predictive-corrective recursive calculations, it outputs a smooth and accurate posture estimate. This posture transformation matrix is then applied to the pre-registered anatomical atlas model with the subject's point cloud, allowing the organ mesh in the atlas to update its position in real time as the subject moves, thus solving the problem of organ positioning drift caused by patient breathing or slight movements.
[0055] b) Path planning: The system receives the human body surface coordinates from the infrared scanner 5 and, combined with the real-time position of the organ, plans a collision-free examination path for the ultrasound probe 3. Specifically, the human body surface coordinates generated by the infrared scanner 5 form a model of the subject's body surface. The central control module maps the organ region to be examined onto this surface model, searches for a continuous path from the probe's current position to the target area under surface constraints, and detects in real time whether the path interferes with other parts of the body, ensuring that the ultrasound probe 3 maintains a safe distance from the body surface during the robotic arm's movement, only pressing down to make contact when reaching the examination position.
[0056] c) Force-position coordinated control: The multi-degree-of-freedom robotic arm system 2 moves the ultrasonic probe 3 along the inspection path while receiving pressure data from the force sensing unit. Based on the preset target contact force range, the pressure of the ultrasonic probe 3 is adjusted through a PID control algorithm to achieve closed-loop tracking of force, ensuring that the ultrasonic probe 3 adheres to the body surface with stable pressure and avoiding image quality degradation caused by pressure fluctuations.
[0057] d) Intelligent Diagnostic Feedback: The system receives digitized ultrasound image data from the ultrasound image acquisition module, uses the EfficientNet model to identify lesion features and assess image clarity in real time, and dynamically adjusts the target contact force based on clarity feedback until the image is clear. The EfficientNet model adaptively enhances diagnostically valuable texture features in the image. This model is pre-trained using a large number of labeled images and can simultaneously output the lesion category (e.g., cyst, tumor, normal) and image clarity score. When the clarity score is below a preset threshold, the central control module gradually increases the target contact force within the allowable force range for the current organ until the image clarity meets the standard. If the image is still unclear even after reaching the upper force limit, a repositioning or alarm process is triggered, enabling the system to adapt to individual differences in body fat percentage, muscle tension, etc., among different patients, ensuring that images of optimal diagnostic quality are always acquired.
[0058] Through the coordinated operation of the above four functions, the system achieves full automation of the entire process from patient positioning, path planning, pressure control to image diagnosis, effectively solving the technical problems of traditional ultrasound examinations that rely on manual operation, have low positioning accuracy, and lack precise pressure control.
[0059] This invention includes a slide bed drive mechanism 7 for driving the movable slide bed 10 to move longitudinally on the base platform 1. The multi-degree-of-freedom robotic arm system 2 includes: a longitudinal moving guide rail 21 disposed on the base platform 1; a gantry frame 22 disposed on the longitudinal moving guide rail 21; a transverse moving guide rail 23 disposed below the gantry frame 22; a telescopic rod 24 disposed on the transverse moving guide rail 23; and a main robotic arm 25 installed at the end of the telescopic rod 24; the telescopic rod 24 is extendable and retractable in the vertical direction. In this invention, the slide bed drive mechanism 7, the longitudinal moving guide rail 21, and the transverse moving guide rail 23 are all servo linear drive mechanisms, and the telescopic rod 24 is a two-stage or multi-stage nested servo electric telescopic rod. Wherein:
[0060] The slide bed drive mechanism 7 automatically sends the patient into or out of the main robotic arm 25 workspace while the patient is lying down, avoiding the inconvenience of manually pushing the patient bed. At the same time, the precise position feedback of the moving slide bed 10 can serve as the reference for the system coordinate system, providing a consistent starting reference point for subsequent path planning.
[0061] The longitudinal moving guide rail 21 is arranged along the length of the base platform 1, and adopts a combination structure of ball screw and linear guide rail. The servo motor drives the screw to rotate through a harmonic reducer, thereby driving the gantry 22 to be precisely positioned longitudinally. The gantry 22 has an inverted U-shaped structure, spanning above the moving slide bed 10, providing a stable mounting base for the transverse moving guide rail 23. The transverse moving guide rail 23 is fixed below the crossbeam of the gantry 22, and adopts the same screw drive structure as the longitudinal moving guide rail 21, allowing the telescopic rod 24 to move laterally. The telescopic rod 24 adopts a two-stage or multi-stage nested design, with a two-stage screw transmission mechanism inside, which can extend and retract significantly in the vertical direction, enabling the main robotic arm 25 to reach various height positions from the neck to the abdomen. The main robotic arm 25 is a multi-joint serial robotic arm, with each joint driven by a servo motor and a harmonic reducer, featuring high precision and high dynamic response, and can realize the posture adjustment of the probe at any position on the body surface. This design, which combines a gantry frame with multi-level guide rails, enables the main robotic arm 25 to have both a wide range of movement capabilities and the ability to perform precise operations in localized areas, thus meeting the travel requirements for whole-body inspections and the precision requirements for localized inspections.
[0062] like Figure 2 , Figure 3 and Figure 4As shown, the present invention also includes a quick-change disc 8, which includes a quick-change disc mother disc 8a and a quick-change disc plate 8b that can be locked and released by a locking mechanism 80. The quick-change disc mother disc 8a is fixedly connected to the end of the main robotic arm 25, and the ultrasonic probe 3 is fixedly connected to the quick-change disc plate 8b. The thin-film pressure sensor 11 is disposed on the assembly and pressing contact surface of the quick-change disc mother disc 8a and the quick-change disc plate 8b. There are multiple quick-change disc plates 8b, and the lower part of each quick-change disc plate 8b is respectively connected to the ultrasonic probe 3 of different frequencies. The lower part of the quick-change disc plate 8b is also provided with a hanging block 81 with a shape that is larger at the top and smaller at the bottom, such as... Figure 5 As shown, the inner wall of the base platform 1 is provided with a side wall support frame 12, and the quick-change tray 8b can be suspended on the side wall support frame 12 via a hanging block 81. The number of quick-change trays 8b is configured according to clinical needs, for example, four sub-traces are set, which are respectively connected to a low-frequency probe (for the heart and deep abdominal organs), a medium-frequency probe (for the gallbladder and pancreas), a high-frequency probe (for the thyroid and breast), and an ultra-high-frequency probe (for superficial skin examination). Each quick-change tray 8b is suspended on the side wall support frame 12 on the inner wall of the base platform 1 via a hanging block 81. The shape of the hanging block 81, which is larger at the top and smaller at the bottom, matches the slot on the side wall support frame 12, so that the sub-traces can be reliably suspended and the master tray at the end of the robotic arm can be docked and placed. When a probe change is needed, the main robotic arm 25 moves above the target sub-disk, aligning the spindle of the quick-change disc mother disc 8a with the center hole of the sub-disk. A guide pin is inserted into the pin hole for coarse positioning. Then, the mother disc descends to bring the assembly pressing contact surfaces into contact, and the locking mechanism 80 engages to lock the disc. The main robotic arm 25 can then lift the sub-disk away from the support frame for inspection. After inspection, the main robotic arm returns the sub-disk to the support frame, the locking mechanism releases, and the probe change is complete. This design allows the system to automatically select the most suitable probe type based on the inspection area, significantly improving the standardization of inspections and image quality.
[0063] It should be noted that the assembly extrusion contact surface is the annular end face where the quick-change disc mother plate 8a and the quick-change disc plate 8b are fitted together in the locked state. The thin-film pressure sensor 11 is embedded in the annular end face and is used to directly sense the pressure transmitted from the ultrasonic probe 3 to the quick-change disc plate 8b and then to the quick-change disc mother plate 8a.
[0064] like Figure 5As shown, the present invention also includes an automatic adhesive application unit 9, which comprises: a side-mounted robotic arm 91 fixed to the inner wall of the substrate platform 1; and a peristaltic pump 92 fixed to the substrate platform 1. One end of the feed tube 93 of the peristaltic pump 92 is connected to a coupling agent tank 94, and the other end is connected to the side-mounted robotic arm 91. The side-mounted robotic arm 91 is used to hold the feed tube 93 and automatically apply coupling agent to the surface of the ultrasonic probe 3 after the ultrasonic probe 3 is switched. When the system completes the probe switching, the central control module issues an adhesive application command. The side-mounted robotic arm 91 moves the discharge end of the feed tube 93 to above the surface of the ultrasonic probe 3. The peristaltic pump 92 starts and expels the coupling agent at a preset flow rate. At the same time, the side-mounted robotic arm 91 moves according to a preset trajectory to evenly apply the coupling agent to the working surface of the probe. After the adhesive application is completed, the side-mounted robotic arm 91 returns to its original position, and the ultrasonic probe 3 can begin detection. The process of applying coupling agent has been automated, avoiding problems such as uneven application and inaccurate dosage caused by manual application. The entire inspection process, from probe selection and coupling agent application to image acquisition, does not require manual intervention, thus improving the standardization and efficiency of the operation.
[0065] A method for intelligent ultrasonic detection control in a multimodal sensing and AI-driven intelligent ultrasonic detection system includes the following steps:
[0066] Step S1: Acquire 3D point cloud of the subject's torso using depth camera 4, acquire motion data of the subject using inertial measurement unit (IMU), fuse point cloud and IMU data using Kalman filter, estimate the subject's skeletal posture in real time, and infer the real-time position of internal organs based on pre-registered anatomical atlas.
[0067] More specifically, the Kalman filter fusion process described in step S1 includes: using IMU data as input to the prediction model to obtain the predicted attitude; performing iterative nearest point (ICP) registration between the point cloud acquired by the depth camera 4 and the point cloud at the previous moment to obtain the observed attitude; fusing the predicted attitude and the observed attitude using Kalman gain to obtain the corrected attitude estimate; and using the truncated symbolic distance field (TSDF) method to fuse the current point cloud into the global 3D model to generate a surface mesh.
[0068] Step S2: Obtain the coordinates of the human body surface using the infrared scanner 5, and combine them with the real-time position of the organs to plan a collision-free examination path for the ultrasound probe 3.
[0069] Step S3: As the ultrasound probe 3 moves along the examination path, the ultrasound image acquisition module acquires digital ultrasound images in real time, and the trained EfficientNet model is used to determine the image clarity and identify lesion features.
[0070] Step S4: Based on the preset target contact force range of the organ being examined and the image clarity judgment result in step S3, adjust the pressing force applied by the main robotic arm 25 through the PID control algorithm so that the image is clear and the contact force does not exceed the safety threshold.
[0071] More specifically, the PID control algorithm in step S4 incorporates image sharpness feedback, specifically: a preset target contact force F target Real-time acquisition of the filtered force value F from the diaphragm pressure sensor 11 filt (k), calculation error e(k) = F target -F filt (k); If the current ultrasound image clarity is lower than the preset threshold, then gradually increase F within the allowable force range. target The pressure is adjusted until the image is clear; the control quantity u(k) is calculated using the PID formula and mapped to the joint motor of the main robotic arm 25 to adjust the pressing force; at the same time, the pressing depth of the ultrasonic probe 3 is monitored by the infrared scanner 5, and if the pressing depth exceeds the set upper limit, the driving force is immediately withdrawn.
[0072] A storage medium storing an executable program, which, when executed by a processor, enables an intelligent ultrasonic detection control method for a multimodal sensing and AI-driven intelligent ultrasonic detection system.
[0073] To better understand the technical solution of this invention, it will be further elaborated below from three aspects: force feedback, human body scanning relocalization, and AI diagnosis, as follows:
[0074] I. Force Feedback
[0075] The pressure applied varies depending on the organ being examined, and is generally adjusted to ensure a clear view of the organ's structure. For example, in abdominal ultrasound: slightly stronger pressure is needed to push aside the ribs for liver examination; moderate pressure for gallbladder examination; lighter pressure for kidney examination; and moderate pressure for pancreas examination to dislodge intestinal gas. Therefore, the pressure should be dynamically adjusted according to the patient's body type and examination requirements. The pressure applied to the ultrasound probe 3 is determined by two factors: the data from the thin-film pressure sensor 11 and the clarity of the ultrasound image. The final implementation plan is as follows: aim for a clear image; if the image is unclear, the pressure applied to the probe is gradually increased, with the pressure sensor monitoring the pressure in real time. An upper limit of 20N is set for the pressure; if the probe pressure exceeds this value, the pressure will stop.
[0076] Considering the similar mapping function between image clarity and pressure applied during ultrasound scans, and that different patients require different pressure levels for clear images, a visual model was trained to determine the clarity of ultrasound images during automatic pressure application. First, the dataset needed to train this model was acquired. Collaboration was conducted with doctors from major tertiary hospitals. During ultrasound scans, doctors acquired ultrasound images, applying pressure from low to high, and determined the clarity of the images. The pressure applied and its corresponding image clarity were recorded and labeled. This method was used to acquire ultrasound images.
[0077] The large dataset collected was categorized according to the type of organs examined, and two important pieces of information were analyzed from it:
[0078] First, what constitutes a clear ultrasound image? This judgment model is trained using the same EfficientNet (efficient convolutional network) as described above. The entire training set is divided into four categories based on image clarity: clear, slightly clear, slightly blurry, and blurry (labeled by the doctor when acquiring the images). The training set is then input into EfficientNet for model training, and the resulting model can determine the clarity of ultrasound images.
[0079] Second, based on the force range corresponding to the organ with clear imaging, target contact forces are set, and different contact forces are set for different organs based on information obtained from dataset analysis. Specifically:
[0080] The liver (approximately 2-3 kg) is used to push aside the ribs; the gallbladder (1-2 kg); the kidneys (0.5-1 kg); the pancreas (1-2 kg) is used to push aside intestinal gas; and other organs weigh approximately 0.5-1 kg.
[0081] The trained model is embedded in an imaging computer, which monitors the clarity of the ultrasound image in real time. The force applied by the probe varies within the allowable range, and the model detects the image clarity in real time. If the image is unclear, the pressure applied by the probe is slowly increased until the image becomes clear.
[0082] The specific control method for probe pressure is as follows: Force data from the thin-film pressure sensor 11 is acquired using a sampling frequency of 150Hz. Since the force signal is frequently affected by jitter, noise, and friction, low-pass filtering is required. Therefore, a first-order low-pass filter is selected, as shown in the following formula:
[0083]
[0084] In the formula The original force data obtained from the kth sampling is as follows: The force value after the previous filtering is used to "memorize" the output during filtering, making the output smooth and without fluctuations. This is the filtered force value at the current moment. These are the filter coefficients.
[0085] When examining a specific organ, the real-time force applied by the probe is compared with the target contact force of that organ:
[0086]
[0087] In the formula: The error force at the current moment, For the target contact force, Output force at the current moment.
[0088] The control section uses a PID control algorithm, and the central control module controls the output, which can be mapped to the intensity of the vibration motor or the driving force of the force feedback actuator, thereby controlling the pressing force applied by the probe.
[0089]
[0090] In the formula: This is the output of the central control board at the current moment. The error force at the current moment, This is the error force from the previous calculation. Controlling the speed at which the applied force reaches the target force, Used to eliminate steady-state error Reduce rapid fluctuations / avoid oscillations. To improve stability and ensure patient safety, the control frequency must be high (≥100-200 Hz), the delay must be less than 10-20 ms, and a force cutoff limit should be set (e.g., forced stop >20 N).
[0091] Infrared feedback obtains the surface coordinates. Based on the obtained surface coordinates of different human bodies, the pressing depth is set (generally 2 cm). When the infrared scanner detects that the pressing depth exceeds the set depth, the motor will immediately cancel the driving force, and the probe will stop pressing down. The infrared scanner plays an auxiliary protection role; the normal pressing pressure is determined by the pressure sensor data.
[0092] Considering individual differences, the pressure applied to the same area may vary among different patients. Therefore, the detection device also includes an alarm system. When the patient feels discomfort and expresses a complaint while continuous pressure is applied to the probe, the probe stops applying pressure and retracts upwards.
[0093] II. Human Body Scan Repositioning
[0094] The method of fusion of depth camera 4 and inertial measurement unit (IMU) data is used to achieve repositioning of important parts of the human body.
[0095] Static localization (initialization): A high-precision 3D point cloud of the user's torso (from the jaw to the groin) is acquired using a depth camera 4. A standard 3D anatomical atlas (containing skin, bones, and target organs) is then aligned with the user's point cloud using a non-rigid registration algorithm, aligning its "skin" layer with the user's point cloud.
[0096] Dynamic tracking (sensor fusion): As the user moves, the IMU provides high-frequency motion data (attitude / acceleration), while the depth camera 4 provides low-frequency 3D shape data. A Kalman filter is used to fuse these two types of data to track the rigid motion of the user's torso smoothly in real time.
[0097] Location inference: The dynamically tracked motion posture (a transformation matrix) is applied to the already registered anatomical atlas. In this way, the internal organ models in the atlas will move in real time with the subject's movements, thereby achieving "localization".
[0098] Before starting, system calibration needs to be completed to find the rigid transformation matrix between the IMU coordinate system and the camera coordinate system, so as to know the relative position and attitude between the IMU and the depth camera.
[0099] The IMU was securely mounted on the subject's torso, primarily positioned at four points: the shoulders and hips, with the camera aimed from the chin to the groin. The subject maintained a standard posture (lying flat) and a 3D point cloud of their torso was acquired using the depth camera. The depth camera captured the point cloud data synchronously at 30Hz. For image coordinates... Depth value Convert to 3D point P using camera intrinsic parameters = ( , , ),in:
[0100]
[0101] in: , Focal length; , Center point.
[0102] The generated 3D point cloud is based on the camera coordinate system and needs to be converted to the world coordinate system. The transformation formula is as follows:
[0103]
[0104] in: These are 3D coordinates in the world coordinate system. These are the 3D coordinates in the camera coordinate system. For rotation matrix, It is a translation vector.
[0105] IMU data fusion: Using a Kalman filter, motion data from the IMU and 3D shape data from the depth camera are fused to estimate the global skeleton pose.
[0106]
[0107] in: For the final attitude estimation, Measurements for ICP registration of depth cameras, These are IMU predicted values.
[0108] This step aligns the drifting IMU model with the real world seen by the camera.
[0109] Next, the RMSE formula is used for evaluation, as follows:
[0110]
[0111] in: For the estimated posture, For accurate posture.
[0112] TSDF (Truncation Symbolic Range Field) uses this corrected pose to merge the current point cloud into the global 3D model, generating a surface mesh.
[0113] Based on the skeletal structure (shoulder-hip), organ locations are estimated. This is derived from the final state of the Kalman filter. In this process, we extract the final world transformation matrices for the "pelvis" and "thoracic cavity" skeletons: and .
[0114] For each vertex on the organ model (liver, lung, thyroid, etc.) It uses a linear hybrid skin formula to calculate its current world position. .
[0115] The position of each vertex is a weighted average of the transformed positions of all bones (here, the shoulder and hip) bound to the IMU.
[0116]
[0117] in, It is the personalized position of the vertex in its normal lying posture. It is the vertex Bind to bones Skin weights, It is bones The inverse of the transformation matrix in normal lying position. It is bones The current real-time transform matrix (derived from the Kalman filter). This represents the final position of the vertex in the world coordinate system at this moment.
[0118] Real-time organ grid is composed of all composition.
[0119] III. AI Diagnosis
[0120] Real-time identification is performed on the acquired ultrasound images, enabling real-time diagnoses of common diseases. Considering the need for deployment on edge devices and the requirement for efficient real-time diagnosis, EfficientNet (an efficient convolutional network) is selected for model training of the ultrasound images. The specific implementation steps are as follows:
[0121] 1. Data Acquisition and Preprocessing: Real-time image acquisition using ultrasound equipment ensures annotation of common diseases (such as tumors and cysts), with a target dataset size of at least 5000 images. Image pixel values are normalized, and image enhancement operations such as rotation, flipping, scaling, and noise addition are performed. Median filtering is used to remove relevant noise and enhance contrast. The acquired dataset is then labeled and supervised learning is implemented.
[0122] 2. The EfficientNet network architecture was chosen for model training. The core principle of EfficientNet is the compound scaling method. EfficientNet optimizes the network structure through systematic search to achieve the optimal balance between parameters and accuracy. The network framework design is as follows:
[0123] The MBConv module: a fundamental building block of EfficientNet, uses depthwise separable convolutions to reduce computation. The computational cost of a standard convolution is: The depth can be separated into:
[0124] Depthwise: Pointwise:
[0125] This method significantly reduces computation.
[0126] Composite scaling. By simultaneously scaling the network's depth, width, and resolution, EfficientNet maintains high computational efficiency as the model size increases. EfficientNet achieves this by combining coefficients. Uniform scaling depth ,width Resolution .
[0127]
[0128] in, , , It is a constant. Controlling the model size ensures efficient model scaling for B-type ultra-high resolution images.
[0129] The core MB-Conv module integrates the Squeeze-and-Excitation (SE) attention mechanism, enabling dynamic adjustment of channel weights. This allows the model to focus more on disease-related features (such as edge textures or density anomalies), thereby improving diagnostic accuracy. Algorithm flow: Squeeze (compress) global information → Excitation (excite) generate weights → Scale (scale) original features.
[0130] Squeeze operation (global average pooling):
[0131]
[0132] in, It is the input feature map of the c-th channel. This is the compressed global information vector. This formula compresses the spatial dimension H×W to 1×1, resulting in the channel descriptor.
[0133] Excitation operation (generating weights in a two-layer fully connected layer):
[0134]
[0135] in, It is the first fully connected layer. It is the second fully connected layer. It is an activation function. This is the channel weight vector. This operation uses a dimensionality reduction-increment structure to reduce the number of parameters.
[0136] Scale operation (weights multiplied by the original features):
[0137]
[0138] in, It involves recalibrating the output and improving the disease-focused channels.
[0139] Two infrared scanners 5 mounted on the gantry 22 emit infrared rays onto the human body surface and calculate the reflection time to obtain the surface coordinates of the human body. The specific steps are as follows:
[0140] Emitting light: The scanner emits invisible infrared light toward the human body.
[0141] Reflection reception: Infrared cameras (sensors) capture light reflected back from the human body surface.
[0142] Depth calculation: The system calculates the depth (Z coordinate) of each pixel based on the received signal (pattern deformation or time of flight of light).
[0143] Coordinate generation: Combining known camera internal parameters (such as focal length) and pixel positions (x, y), coordinates are generated for each point in the 2D image. The calculated depth Z is converted into three-dimensional coordinates in space. Ultimately, these coordinates converge to form a "point cloud," which is a collection of coordinates of the human body surface.
[0144] The obtained surface information is mainly used in the path planning and probe switching operation (i.e., the probe should not touch other areas of the human body when it moves).
[0145] The path planning is as follows:
[0146] For the organ 3D mesh defined above, calculate the shortest paths to all reachable points starting from the starting point. Obtain the coordinates of the organ.
[0147] The system is designed with two modes: one is to test only specific locations according to the patient's request; the other is to perform a full-body examination according to a pre-designed pathway for probe replacement and other procedures.
[0148] When examining a specific area using ultrasound, the depth camera pre-defines organ regions for each patient and designs a path from the corresponding ultrasound probe to the designated organ region. The overall path planning begins with the application of coupling gel and is completed by two robotic arms working together. Starting with the thyroid gland at the top of the body, the corresponding ultrasound probe is selected and operated. The same probe is then used for the heart, liver, and other organs in sequence. Finally, the testes are examined at the bottom. This sequence is followed to complete the ultrasound examination of the entire body.
[0149] In summary, the advantages of this invention are as follows:
[0150] 1) Through the collaborative work of the depth camera, infrared scanner, and surface-fixed IMU in the multimodal sensing unit, dual localization of the human body's external contours and internal organs is achieved. The Kalman filter fusion of the 3D point cloud acquired by the depth camera and the high-frequency motion data from the IMU can correct for organ position shifts caused by patient breathing or slight movements in real time. This ensures that the ultrasound probe driven by the robotic arm is always aligned with the target examination area, significantly improving the tracking accuracy of deep organs and minute lesions, and solving the bottleneck problem of traditional ultrasound relying on operator experience to judge organ position. The human body surface coordinates acquired by the infrared scanner provide the main robotic arm with precise obstacle avoidance path planning. Combined with preset whole-body or local examination paths, the examination process is standardized and reproducible, avoiding inconsistent diagnostic results due to differences in operator techniques.
[0151] 2) Regarding probe pressure control, a thin-film pressure sensor embedded in the quick-change plate directly senses the actual pressure of the probe in contact with the human body. The pressure signal is processed by a PID algorithm to adjust the downward pressure of the main robotic arm in real time. Simultaneously, the system incorporates ultrasound image clarity as a target parameter for feedback control. When the AI model determines the image is blurry, it automatically fine-tunes the pressure until the image is clear, and immediately stops applying pressure when the preset pressure limit is reached or the infrared scanner detects that the pressure depth exceeds the limit. This dual closed-loop control based on image quality and pressure thresholds ensures both the safety of the examination process and that the acquired images are always at the best diagnostic quality, overcoming the limitation of traditional examinations where pressure control relies entirely on the operator's feel.
[0152] 3) Through the cooperation of the quick-change disc and the automatic adhesive application unit, automatic switching of multi-frequency probes and quantitative application of coupling agent are achieved. This allows the system to automatically select the most suitable probe type for different examination sites, further improving the standardization and operational efficiency of the examination. The central control module's fusion processing of multimodal data and coordinated scheduling of various execution units enable the entire ultrasound examination process to automatically complete organ localization, path planning, pressure adjustment, and image diagnosis without manual intervention, significantly reducing reliance on the operator's professional skills.
[0153] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal sensing and AI-driven intelligent ultrasonic detection system, characterized in that: include: The base platform (1) serves as the foundation for the system and is equipped with a movable slide (10). A multi-degree-of-freedom robotic arm system (2) is mounted on the base platform (1) and has an ultrasonic probe (3) detachably connected to its end for driving the ultrasonic probe (3) to move on the surface of the human body; The multimodal sensing unit includes a depth camera (4) and an infrared scanner (5) fixed on a base platform (1), and an inertial measurement unit (IMU) that can be fixed to the surface of the subject's body; the depth camera (4) is used to acquire 3D point clouds of the human torso, the infrared scanner (5) is used to acquire the coordinates of the human body surface, and the inertial measurement unit (IMU) is used to capture the subject's motion posture. The force sensing unit includes a thin-film pressure sensor (11) disposed between the end of the multi-degree-of-freedom robotic arm system (2) and the ultrasonic probe (3), for real-time detection of the contact pressure between the ultrasonic probe (3) and the human body; The main control console (6) is electrically connected to the ultrasonic probe (3), the multi-degree-of-freedom robotic arm system (2), the multimodal sensing unit, and the force sensing unit; the main control console (6) integrates: An ultrasound image acquisition module, which receives the raw ultrasound signals acquired by the ultrasound probe (3) and generates digitized ultrasound image data; and The central control module is configured as follows: a) Fusion positioning: Receive the 3D point cloud data from the depth camera (4) and the motion data from the inertial measurement unit (IMU), fuse them through Kalman filtering, estimate the subject's skeletal posture in real time, and deduce the real-time position of internal organs based on the pre-registered anatomical atlas. b) Path planning: Receive the human body surface coordinates from the infrared scanner (5), and combine them with the real-time position of the organ to plan the collision-free inspection path of the ultrasound probe (3); c) Force-position coordinated control: control the multi-degree-of-freedom robotic arm system (2) to move along the inspection path with the ultrasonic probe (3), and at the same time receive the pressure data of the force sensing unit. According to the preset target contact force range, adjust the pressing force of the ultrasonic probe (3) through the PID control algorithm. d) Intelligent diagnostic feedback: Receives digitized ultrasound image data from the ultrasound image acquisition module, identifies lesion features and judges image clarity in real time through the EfficientNet model, and dynamically adjusts the target contact force based on clarity feedback until the image is clear.
2. The multimodal sensing and AI-driven intelligent ultrasonic detection system according to claim 1, characterized in that: It includes a slide drive mechanism (7) for driving the movable slide (10) to move longitudinally on the base platform (1).
3. The intelligent ultrasonic detection system with multimodal perception and AI drive according to claim 2, characterized in that: The multi-degree-of-freedom robotic arm system (2) includes: a longitudinal moving guide rail (21) set on the base platform (1), a gantry frame (22) set on the longitudinal moving guide rail (21), a transverse moving guide rail (23) set below the gantry frame (22), a telescopic rod (24) set on the transverse moving guide rail (23), and a main robotic arm (25) installed at the end of the telescopic rod (24); the telescopic rod (24) can extend and retract in the vertical direction.
4. The multimodal perception and AI-driven intelligent ultrasonic detection system according to claim 3, characterized in that: It also includes a quick-change disc (8), which includes a quick-change disc mother disc (8a) and a quick-change disc plate (8b) that can be locked and released by a locking mechanism (80). The quick-change disc mother disc (8a) is fixedly connected to the end of the main robotic arm (25), the ultrasonic probe (3) is fixedly connected to the quick-change disc plate (8b), and the thin-film pressure sensor (11) is set on the assembly extrusion contact surface of the quick-change disc mother disc (8a) and the quick-change disc plate (8b). There are multiple quick-change trays (8b), and the lower part of each quick-change tray (8b) is connected to the ultrasonic probe (3) of different frequencies; the lower part of each quick-change tray (8b) is also provided with a hanging block (81) with a larger upper part and a smaller lower part, and the inner side wall of the base platform (1) is provided with a side wall support frame (12), and the quick-change tray (8b) can be suspended on the side wall support frame (12) by the hanging block (81).
5. The intelligent ultrasonic detection system with multimodal perception and AI drive according to claim 4, characterized in that: The assembly extrusion contact surface is the annular end face of the quick-change disc mother plate (8a) and quick-change disc plate (8b) in the locked state. The thin film pressure sensor (11) is embedded in the annular end face and is used to directly sense the pressure transmitted by the ultrasonic probe (3) to the quick-change disc plate (8b) and then to the quick-change disc mother plate (8a).
6. The intelligent ultrasonic detection system with multimodal perception and AI drive according to claim 1, characterized in that: It also includes an automatic glue application unit (9), which includes: A side-mounted robotic arm (91) is fixed to the inner wall of the base platform (1). A peristaltic pump (92) is fixed on the base platform (1). One end of the feed pipe (93) of the peristaltic pump (92) is connected to the coupling agent tank (94), and the other end is connected to the side-mounted robotic arm (91). The side-mounted robotic arm (91) is used to hold the tube (93) and automatically applies coupling agent to the surface of the ultrasonic probe (3) after the ultrasonic probe (3) is switched.
7. The intelligent ultrasonic detection control method for a multimodal perception and AI-driven intelligent ultrasonic detection system according to claim 3, characterized in that: Includes the following steps: Step S1: Collect the 3D point cloud of the subject's torso using a depth camera (4), collect the subject's motion data using an inertial measurement unit (IMU), fuse the point cloud and IMU data using a Kalman filter, estimate the subject's skeletal posture in real time, and deduce the real-time position of the internal organs based on the pre-registered anatomical atlas. Step S2: Obtain the human body surface coordinates through the infrared scanner (5), and combine them with the real-time position of the organs to plan the collision-free inspection path of the ultrasound probe (3); Step S3: During the movement of the ultrasound probe (3) along the examination path, the ultrasound image acquisition module acquires digital ultrasound images in real time, and the trained EfficientNet model is used to judge the image clarity and identify lesion features. Step S4: Based on the preset target contact force range of the organ being examined and the image clarity judgment result in step S3, adjust the pressing force applied by the main robotic arm (25) through the PID control algorithm so that the image reaches a clear state and the contact force does not exceed the safety threshold.
8. The intelligent ultrasonic detection control method of a multimodal perception and AI-driven intelligent ultrasonic detection system according to claim 7, characterized in that: The specific process of Kalman filter fusion in step S1 includes: IMU data is used as input to the prediction model to obtain the predicted attitude; The point cloud acquired by the depth camera (4) is iteratively registered with the point cloud at the previous moment using the nearest point ICP to obtain the observation pose; The predicted attitude and the observed attitude are fused using Kalman gain to obtain the corrected attitude estimate. The truncated symbolic range field (TSDF) method is then used to fuse the current point cloud into the global 3D model to generate a surface mesh.
9. The intelligent ultrasonic detection control method of a multimodal perception and AI-driven intelligent ultrasonic detection system according to claim 7, characterized in that: The PID control algorithm described in step S4 incorporates image sharpness feedback, specifically as follows: Preset target contact force F target ; Real-time acquisition of the filtered force value F from the diaphragm pressure sensor (11) filt (k), calculation error e(k) = F target -F filt (k); If the current ultrasound image clarity is below the preset threshold, then gradually increase F within the allowable force range. target Continue until the image is clear; The control quantity u(k) is calculated using the PID formula and mapped to the joint motor of the main robotic arm (25) to adjust the pressing force. At the same time, the pressure depth of the ultrasonic probe (3) is monitored by an infrared scanner (5). If the pressure depth exceeds the set upper limit, the driving force is immediately withdrawn.
10. A storage medium, characterized in that: It contains an executable program, which, when executed by a processor, enables an intelligent ultrasonic detection control method for a multimodal sensing and AI-driven intelligent ultrasonic detection system as described in any one of claims 7 to 9.