Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-14
AI Technical Summary
[0007]本发明的目的在于提供一种基于多传感器融合的Freehand超声空间感知方法,旨在解决现有Freehand超声定位技术中存在的光学传感器遮挡、惯性传感器漂移等问题,提升Freehand超声的定位精度和稳定性,特别是在动态监测和多次扫描过程中,确保超声影像的空间一致性和可重复性
[0016]本发明取得的有益效果为:本发明通过力学传感信息与惯性测量信息的协同融合,引入三维力学传感器对操作者手部施力模式进行建模,有效弥补了单一IMU在长期运行过程中存在的漂移和误差累积问题,在光学定位模块被遮挡或失效情况下仍能够实现稳定、连续的位姿估计。具体如下:
Smart Images

Figure CN122350768B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to ultrasound imaging technology, and in particular to a freehand ultrasound spatial perception method based on multi-sensor fusion, which can be applied to the field of medical imaging, especially for ultrasound examination and dynamic monitoring of superficial organs such as the thyroid and breast. It aims to solve the problems of optical sensor occlusion and inertial measurement unit (IMU) drift in existing freehand ultrasound positioning technology, and provide a high-precision, long-term stable spatial positioning and attitude estimation method to improve the accuracy and robustness of freehand ultrasound technology in clinical applications. Background Technology
[0002] In ultrasound imaging technology, freehand ultrasound, as a flexible imaging method that does not require fixed equipment, is widely used in various clinical examination scenarios, such as real-time ultrasound scanning of organs like the thyroid, breast, and cardiovascular system. Its core advantages lie in its flexible and convenient operation, making it suitable for bedside examinations, emergency care, and interventional procedures. However, the accuracy and stability of freehand ultrasound positioning still face a series of technical bottlenecks in multiple scans and dynamic monitoring.
[0003] Currently, spatial positioning in freehand ultrasound primarily relies on several technologies, such as electromagnetic systems, optical tracking systems, and inertial measurement units. Electromagnetic systems offer high positioning accuracy, but their application is limited by electromagnetic interference and their operating range is relatively narrow, making them unsuitable for open environments or dynamic scenarios. Furthermore, the high cost of electromagnetic systems also restricts their widespread use in primary healthcare.
[0004] Optical tracking systems achieve spatial positioning by capturing marker points or the posture of the ultrasound probe. While providing relatively real-time positioning data, in practical clinical applications, optical systems are susceptible to external factors such as line-of-sight obstruction and changes in lighting, leading to decreased system accuracy or even tracking interruption. Furthermore, the frequency of optical positioning pose output depends on expensive computing equipment. IMU sensors, as a highly portable and robust positioning technology, are widely used in pose estimation for freehand ultrasound. However, IMU sensors suffer from high noise and drift issues; over long-term use, pose estimation errors gradually accumulate, affecting the system's stability and accuracy.
[0005] Therefore, existing freehand ultrasound spatial positioning technology still suffers from problems such as insufficient positioning accuracy, poor system stability, and poor repeatability in clinical applications. In particular, when performing multiple scans and long-term dynamic monitoring, accuracy and stability cannot be effectively guaranteed, which limits its application potential in complex clinical scenarios.
[0006] To address the aforementioned issues, this invention proposes a Freehand ultrasonic spatial sensing method based on multi-sensor fusion, aiming to overcome the shortcomings of existing technologies. By combining multi-sensor data fusion with pose estimation algorithms, it achieves higher precision and more stable ultrasonic positioning, and effectively improves the application performance of Freehand ultrasound in dynamic monitoring. Summary of the Invention
[0007] The purpose of this invention is to provide a freehand ultrasound spatial sensing method based on multi-sensor fusion, aiming to solve problems such as optical sensor occlusion and inertial sensor drift in existing freehand ultrasound positioning technologies, and to improve the positioning accuracy and stability of freehand ultrasound, especially ensuring the spatial consistency and repeatability of ultrasound images during dynamic monitoring and multiple scans. The specific technical solution adopted by this invention is as follows: A freehand ultrasonic spatial sensing method based on multi-sensor fusion includes the following steps: S1: Obtain the initial spatial pose information of the ultrasonic probe; S2: Simultaneously acquire optical pose information, mechanical sensing information and inertial measurement information to obtain multimodal raw observation data; S3: Perform time synchronization, preprocessing and coordinate system transformation on the multimodal raw observation data to obtain optical pose feature vector, mechanical feature vector and inertial feature vector, forming a mechanical-inertial measurement information fusion vector sequence; S4: When the optical positioning module is visible, the BiLSTM model is used to predict the pose based on the fusion vector sequence of mechanical-inertial measurement information. Then, KalmanNet is used to correct the result according to the optical pose, and the parameters of the BiLSTM model and KalmanNet network are updated online. S5: When the optical positioning module is blocked, the BiLSTM model is used to predict the pose based on the fusion vector sequence of mechanical-inertial measurement information, and then KalmanNet is used to infer and compensate the result based on the historical pose. S6: Outputs the spatial pose information of the ultrasound probe for use in Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
[0008] Furthermore, the method of S1 is as follows: S101: Install the three-dimensional mechanical sensor, six-axis IMU and optical positioning module at predetermined positions around the Freehand ultrasound probe to establish the initial working state of the ultrasound probe; S102: Start the camera and scan the optical positioning module under unobstructed conditions. Take its optical pose information during the initialization phase and calculate the average value as the initial spatial pose information of the ultrasonic probe in the world coordinate system.
[0009] Furthermore, the method of S2 is as follows: S201: Real-time acquisition of pressure distribution information between the operator's hand and the ultrasonic probe handle via a three-dimensional mechanical sensor located at the handle of the ultrasonic probe; S202: Real-time acquisition of angular velocity and linear acceleration information of the ultrasonic probe via a six-axis IMU; S203: Acquires real-time spatial pose information of the ultrasonic probe under unobstructed conditions through an optical positioning module and camera.
[0010] Furthermore, the S3 method is as follows: S301: Hardware-level time synchronization of optical pose information, mechanical sensing information and inertial measurement information is performed through FPGA to generate a unified timestamp and obtain synchronized multi-source sensing data. S302: Perform data preprocessing operations on the synchronized multi-source sensor data, including filtering and feature extraction, including: For inertial measurement information, the acceleration and angular velocity are low-pass filtered, the angular acceleration is calculated, and the inertial feature vector characterizing the instantaneous motion state of the ultrasonic probe is obtained, which is used to describe the instantaneous motion state of the ultrasonic probe. For mechanical sensing information, the pressure intensity vector is calculated based on the resistance change of the three-dimensional mechanical sensor, the pressure center is calculated, the pressure change rate and the overall force intensity are calculated, and a mechanical feature vector is obtained to describe the changes in the operator's hand force application pattern. For optical pose information, Kalman filtering is used to smooth the pose sequence to obtain optical pose feature vectors, and optical pose confidence is calculated for subsequent occlusion determination. S303: Map the three types of feature vectors extracted in step S302 to the ultrasonic probe coordinate system so that the three types of feature vectors describe the same spatial pose state; forming a continuous mechanical-inertial measurement information fusion vector sequence.
[0011] Furthermore, the S4 method is as follows: S401: Use the optical pose feature vector obtained in the preceding steps as the reference pose; S402: Input the fused vector sequence of mechanical-inertial measurement information obtained in the previous steps into BiLSTM, and output the predicted pose; S403: The residual between the predicted pose and the reference pose is learned using KalmanNet to correct the predicted pose, and the parameters of the BiLSTM model and KalmanNet network are updated online.
[0012] Furthermore, the method for determining whether the optical positioning module is in a visible state is as follows: The optical pose confidence is calculated. When the output of the optical positioning module is stable and the optical pose confidence of multiple consecutive frames is higher than the preset threshold, the optical positioning module is determined to be visible and in a usable state.
[0013] Furthermore, the method for determining whether the optical positioning module is in an occluded state is as follows: when the optical positioning module loses its pose or the confidence of the optical pose is lower than a preset threshold for multiple consecutive frames, the optical module is determined to be occluded and in an unusable state.
[0014] Furthermore, S5 includes: S501: Stop receiving optical pose, input the fusion vector sequence of mechanical-inertial measurement information into the BiLSTM model, and output the predicted pose; S502: KalmanNet is used for inference compensation based on historical poses to suppress IMU drift error.
[0015] Furthermore, the S6 method is as follows: S601: Outputs the spatial pose information of the ultrasonic probe at a frequency of not less than 120 Hz; S602: Time-align the spatial pose information of the ultrasound probe with the ultrasound image frame. S603: Used for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
[0016] The beneficial effects achieved by this invention are as follows: By synergistically fusing mechanical sensing information and inertial measurement information, this invention introduces a three-dimensional mechanical sensor to model the operator's hand force application pattern, effectively compensating for the drift and error accumulation problems that exist in a single IMU during long-term operation. Even when the optical positioning module is obstructed or malfunctions, stable and continuous pose estimation can still be achieved. Specifically: (1) This invention achieves hardware-level synchronization of sensor data by using a multi-source sensor data synchronous acquisition and preprocessing step based on a field programmable gate array (FPGA), establishes a unified time reference and coordinate mapping relationship, effectively reduces timing errors in the multimodal data fusion process, improves the overall spatial perception accuracy of the system, and can output high-frequency, sub-millimeter-level precision ultrasonic probe pose information. (2) The present invention constructs a multimodal fusion pose estimation algorithm. When the optical positioning module is visible, the optical pose is used as a reference to update the model parameters in real time. When the optical positioning module is occluded, unsupervised prediction compensation is achieved, thereby significantly improving the robustness and adaptability of the system in the face of complex clinical environments. (3) The present invention constructs a freehand ultrasound spatial sensing method that does not rely on a fixed support or electromagnetic field environment. It has the advantages of simple structure, wide applicability and flexible deployment, and can adapt to various clinical application scenarios such as bedside examination, intraoperative operation and multiple follow-up. Attached Figure Description
[0017] Figure 1 This is a flowchart of the Freehand ultrasonic spatial sensing method based on multi-sensor fusion provided by the present invention; Figure 2 This is a schematic diagram of the multi-source sensor data acquisition device provided by the present invention; Figure 3 This is a schematic diagram of the multimodal fusion pose estimation algorithm provided by the present invention; Figure 4 This is a diagram showing the results of a neck phantom scanning experiment according to the present invention.
[0018] The annotations in the attached figures are explained as follows: 1-Upper part of probe fixture; 2-Ultrasonic probe; 3-Three-dimensional mechanical sensor; 4-Optical calibration plate; 5-Lower part of probe fixture; 6-Six-axis IMU. Detailed Implementation
[0019] To make the core objectives, technical details, and application advantages of this technical solution embodiment clearer, the specific technical logic of this technical solution embodiment will be explained in detail and comprehensively below with the accompanying schematic diagrams. The embodiments presented here are only some typical embodiments of this technical solution, not all embodiments. It should be noted that this technical solution is not limited to the implementation forms detailed herein, and can also be implemented through other reasonable technical paths without departing from its core technical principles. Those skilled in the art can make adaptive adjustments based on the core connotation of this technical solution, therefore this technical solution is not bound by the specific embodiments disclosed below. Secondly, the "one embodiment" or "preferred embodiment" mentioned herein refers to the key technical features, implementation structures, or functional characteristics included in at least one practical application scenario of this technical solution; the phrase "in a feasible implementation" appearing in different paragraphs of this description does not all refer to the same embodiment, nor is it a specific embodiment that is independent of or exclusive to other embodiments.
[0020] Please see Figure 1 This invention provides a freehand ultrasonic spatial sensing method based on multi-sensor fusion, comprising the following steps: S1: Obtain the initial spatial pose information of the ultrasonic probe; S2: Simultaneously acquire optical pose information, mechanical sensing information and inertial measurement information to obtain multimodal raw observation data; S3: Perform time synchronization, preprocessing and coordinate system transformation on the multimodal raw observation data to obtain optical pose feature vector, mechanical feature vector and inertial feature vector, forming a mechanical-inertial measurement information fusion vector sequence; S4: When the optical positioning module is visible, pose prediction is performed based on the fusion vector sequence of mechanical-inertial measurement information using the BiLSTM (Bi-directional Long Short-Term Memory) model. Then, KalmanNet is used to correct the results according to the optical pose, and the parameters of the BiLSTM model and KalmanNet network are updated online. S5: When the optical positioning module is blocked, the BiLSTM model is used to predict the pose based on the fusion vector sequence of mechanical-inertial measurement information, and then KalmanNet is used to infer and compensate the result based on the historical pose. S6: Outputs the spatial pose information of the ultrasound probe for use in Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
[0021] As described in steps S1-S6 above, the present invention can achieve high-precision, long-term stable spatial perception of Freehand ultrasound in complex clinical operating environments.
[0022] The following description Figure 1 The execution method of each step is shown.
[0023] Regarding S1, in a preferred embodiment, the step of obtaining the initial spatial pose information of the Freehand ultrasound includes: S101: A multi-source sensor data acquisition device is used to install a three-dimensional mechanical sensor, a six-axis IMU, and an optical positioning module at predetermined positions around the Freehand ultrasound probe, establishing the initial working state of the ultrasound probe. The device structure is referenced below. Figure 2 It includes: the upper part of the probe holder 1; the ultrasonic probe 2; the three-dimensional mechanical sensor 3; the optical calibration plate 4; the lower part of the probe holder 5 and the six-axis IMU 6; S102: Activate the monocular high frame rate industrial camera to acquire the initial optical pose information of the ultrasonic probe in the world coordinate system, which serves as the reference pose for ultrasonic probe initialization. This information is used for subsequent multi-sensor coordinate alignment and model initialization.
[0024] In this embodiment, the Freehand ultrasound system is first initialized by fixing the three-dimensional mechanical sensor, six-axis IMU, and optical positioning module at predetermined positions on the ultrasound probe and connecting the power supply. The actual spatial pose of the ultrasound probe is assumed to be: in, This indicates the three-dimensional position of the ultrasonic probe in the world coordinate system. Represents a quaternion of attitude.
[0025] Start the monocular high frame rate industrial camera and scan the optical positioning module at a frequency of 90 Hz for no less than 2 seconds under unobstructed conditions to obtain the optical orientation information of the ultrasonic probe: And take the time window during the initialization phase. Calculate the mean of the optical pose samples within the range: This average value serves as the initial spatial pose of the ultrasound probe, and is used as the starting condition for subsequent multimodal spatial pose recursion. The optical positioning module is a square optical calibration plate with a side length of 5 cm, on which preset ArUco code marks are printed, containing a total of 6×6 coding grids to ensure that the marks occupy sufficient pixels in the image at typical working distances (0.5-1.5 meters).
[0026] Start the six-axis IMU and acquire inertial measurement information of the ultrasonic probe in a stationary state at a sampling frequency of 1000 Hz: in, Represents the three-axis angular velocity vector. This represents the three-axis acceleration vector. IMU bias calibration is performed using a multi-frame moving average method. Then it can be used later. This serves as the corrected inertial measurement information. The six-axis IMU uses a high-precision IMU module, which includes a three-axis gyroscope and a three-axis accelerometer, providing three-axis angular velocity and three-axis linear acceleration information.
[0027] At this point, the ultrasonic probe has completed the initial spatial pose state establishment and entered the continuous data acquisition stage.
[0028] Regarding S2, in a preferred embodiment, the step of simultaneously acquiring optical pose information, mechanical sensing information, and inertial measurement information to obtain multimodal raw observation data includes: S201: Real-time acquisition of pressure distribution information between the operator's hand and the ultrasonic probe handle via a three-dimensional mechanical sensor located at the handle of the ultrasonic probe; S202: Real-time acquisition of angular velocity and linear acceleration information of the ultrasonic probe via a six-axis IMU; S203: Acquires real-time spatial pose information of the ultrasonic probe under unobstructed conditions through an optical positioning module and a monocular high frame rate industrial camera.
[0029] In this embodiment, after the freehand ultrasonic scan begins, the system enters a continuous multimodal data acquisition state. The three-dimensional mechanical sensor operates at a sampling frequency of 500 Hz. Real-time acquisition of the pressure distribution vector between the operator's hand and the ultrasonic probe handle: This pressure distribution vector is used to characterize the direction, magnitude, and trend of the force applied by the operator. , and These represent the pressure distribution vector components along the x-axis, y-axis, and z-axis, respectively.
[0030] The six-axis IMU operates at a sampling frequency of 1000 Hz. Acquire inertial measurement information of the ultrasonic probe in a stationary state: in, Represents the three-axis angular velocity vector. This represents the three-axis acceleration vector. It is used to describe the instantaneous motion state of the ultrasonic probe.
[0031] When an optical positioning module is present within the field of view of a monocular high frame rate industrial camera, it continuously operates at a frequency of 90 Hz. Output the spatial pose information of the ultrasound probe: At this point, the system has formed a multimodal raw observation data set, including optical pose information, mechanical sensing information, and inertial measurement information: The observed characteristics provide a data foundation for subsequent steps.
[0032] Regarding S3, in a preferred embodiment, the steps of time synchronization, preprocessing, and coordinate system transformation of the multimodal raw observation data include: S301: Hardware-level time synchronization of optical pose information, mechanical sensing information and inertial measurement information is performed through FPGA to generate a unified timestamp and obtain synchronized multi-source sensing data. S302: Perform data preprocessing operations, including filtering and feature extraction, on the synchronized multi-source sensor data; S303: The three types of extracted feature vectors are uniformly mapped to the ultrasonic probe coordinate system so that the three types of feature vectors describe the same spatial pose state.
[0033] In this embodiment, since the multimodal raw observation data sets obtained in the preceding steps have different sampling frequencies and time bases, unified processing is required. First, hardware-level time synchronization is performed on the optical pose information, mechanical sensing information, and inertial measurement information to generate a unified timestamp. The FPGA module, as the system's main control clock source, uses a high-precision temperature-compensated crystal oscillator (TCXO, 50 MHz) to generate a global clock base, constructing a discrete unified time axis. in, Represents a time series index. To ensure a unified operating frequency for the system.
[0034] The FPGA simultaneously triggers a 3D motion sensor, a six-axis IMU, and a monocular high-frame-rate industrial camera to acquire data via a hardware trigger signal (TTL level), ensuring that all sensors begin sampling at the same time. Simultaneously, a 32-bit timestamp based on a global clock is added to each frame of data to record the data acquisition time of each sensor. , and Interpolation and resampling methods are used to map the original observations to a unified time axis: Thus, the synchronous observation sequence is obtained: Achieve sub-millisecond time alignment of multi-source sensor data with a maximum synchronization error of less than 10 μs.
[0035] Subsequently, the synchronized data was cleaned and its features were extracted.
[0036] For inertial measurement information, low-pass filtering is applied to acceleration and angular velocity: And further calculate the angular acceleration: The inertial eigenvector is then: This vector is used to describe the instantaneous motion state of the ultrasonic probe.
[0037] For mechanical sensing information, the pressure intensity vector is calculated based on the resistance change of a three-dimensional mechanical sensor: in, This indicates the resistance value of the piezoresistive unit before it is subjected to pressure. This indicates the resistance value of the piezoresistive unit after it is compressed. Calculate the pressure center: in, Indicates the first The spatial position vector of each piezoresistive unit. The rate of pressure change is then calculated. and overall force intensity : The final mechanical eigenvectors are obtained as follows: Used to describe changes in the operator's hand force application pattern.
[0038] For optical pose information, since the acquisition frequency of a monocular high frame rate industrial camera is relatively low, a Kalman filter is used to smooth the pose sequence, reducing pose fluctuations caused by marker detection jitter, thereby obtaining the optical pose feature vector: Simultaneously, calculate the optical pose confidence level. Used for subsequent occlusion determination.
[0039] Finally, the extracted feature vectors are uniformly mapped to the ultrasonic probe coordinate system. Since the six-axis IMU coordinate system, the three-dimensional mechanical sensor coordinate system, and the optical positioning module coordinate system all have a fixed, rigid transformation relationship with the ultrasonic probe center coordinate system, the system obtains the transformation matrix through pre-calibration. , and All feature vectors are uniformly mapped to the coordinate system of the ultrasound probe center: This allows all three types of feature vectors to describe the same spatial pose state. .
[0040] Constructing the mechanics-IMU fusion vector: The transformed fused observation vector was resampled to 500 Hz along a unified time axis. Finally, within a time window of length... At that time, a continuous sequence of fusion vectors of mechanical-inertial measurement information is formed: It is then transmitted to the host computer algorithm module via a high-speed serial interface to provide real-time input for subsequent algorithm models.
[0041] Regarding S4, in a preferred embodiment, the steps of using a BiLSTM model to predict pose while the optical positioning module is visible, then using KalmanNet to correct the result based on the optical pose, and updating the BiLSTM model and KalmanNet network parameters online include: S401: Use the optical pose feature vector obtained in the preceding steps as the reference pose; S402: Input the fused vector sequence of mechanical-inertial measurement information obtained in the previous steps into BiLSTM, and output the predicted pose; S403: The residual between the predicted pose and the reference pose is learned using KalmanNet to correct the predicted pose, and the parameters of the BiLSTM model and KalmanNet network are updated online.
[0042] In this embodiment, when the system detects that the output of the optical positioning module is stable and the optical pose confidence level is high for multiple consecutive frames, Higher than the preset threshold At this point, the optical positioning module is determined to be visible and usable, and the system enters the online network parameter update mode. Then, the mechanical-inertial measurement information fusion vector sequence obtained in the previous steps is used. Inputting into the BiLSTM pose prediction model, the output is the predicted pose of the ultrasound probe at the current moment: in, To predict the location, To predict the attitude. This prediction reflects a priori estimation of the ultrasonic probe's motion state under conditions relying solely on inertia and applied force observation. Simultaneously, the optical pose feature vector is: This serves as the reference pose for the current moment. Subsequently, KalmanNet is used to learn the residual between the predicted pose and the reference pose (using...). (represented) for Kalman network parameters, i.e., the dynamic Kalman gain matrix. Update and correct the predicted pose: in, This represents the pose correction mapping learned by KalmanNet during the process. Finally, based on the corrected predicted pose, the BiLSTM model parameters are updated online through a backpropagation strategy, allowing the model to gradually adapt to different operators' force application habits and scanning methods, effectively suppressing the cumulative drift caused by long-term IMU operation.
[0043] For S5, in a preferred embodiment, when the optical positioning module is occluded, the steps of using a BiLSTM model for pose prediction and then using KalmanNet to infer and compensate the results based on historical poses include: S501: Input the fused vector sequence of mechanical-inertial measurement information into BiLSTM and output the predicted pose; S502: KalmanNet is used for inference compensation based on historical poses to suppress IMU drift error.
[0044] In this embodiment, when the system detects that the optical positioning module has lost its pose, or that the confidence level of the optical pose has decreased for multiple consecutive frames, Below the preset threshold When the optical module is obstructed and unavailable, the system automatically switches to the collaborative prediction compensation mode. At this point, the system stops receiving optical pose data and first inputs the mechanical-IMU fusion vector into the BiLSTM to obtain the predicted pose of the ultrasonic probe at the current moment. Subsequently, KalmanNet performs inference compensation solely based on the current predicted pose and historical pose of the ultrasound probe: in, This represents the pose correction mapping obtained from the aforementioned steps. The symbol for empty set indicates that no usable optical pose information exists. This method can perform inference compensation for the current predicted pose based on historical poses, thereby maintaining the continuity and stability of pose output even when the optical module is occluded.
[0045] Regarding S6, in a preferred embodiment, the step of outputting continuous, high-frequency ultrasound probe spatial pose information for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications includes: S601: Outputs the spatial pose information of the ultrasonic probe at a frequency of not less than 120 Hz; S602: Time-align pose information with ultrasound image frames; S603: Used for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
[0046] In this embodiment, the system continuously outputs the ultrasonic probe pose information from the preceding steps at a frequency of not less than 120 Hz, and the spatial pose state of the ultrasonic probe output at any given time satisfies: in: Assume the ultrasonic probe is at time... Acquire two-dimensional image frames The system completes spatial binding through a unified timestamp: The system generates ultrasound image frame data with spatial location information for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications. Throughout the scanning process, the system continuously monitors the status of the optical module. When the optical signal recovers and meets the stability criteria, it automatically switches back to the online network parameter update mode, achieving a seamless transition between different operating states and ensuring the high accuracy and long-term stability of Freehand ultrasound spatial perception.
[0047] This invention refers to steps S4 and S5 as an algorithm called the multimodal fusion pose estimation algorithm, the process of which is as follows: Figure 3As shown. After the system completes the acquisition, calibration, and spatiotemporal alignment preprocessing of multi-source sensor data, it needs to estimate and reconstruct the continuous spatial motion state of the Freehand ultrasound during the scanning process. This multimodal fusion pose estimation algorithm aims to solve the problems of strong nonlinearity and randomness of handheld motion, accumulation of IMU error over time, and possible interruption of optical observation due to obstruction during actual ultrasound probe scanning.
[0048] Under a unified state space representation, the algorithm constructs a two-stage estimation mechanism of "motion prediction + correction or compensation": First, it uses inertial and mechanical multimodal feature sequences to model the motion of the ultrasonic probe through BiLSTM to achieve spatial pose prediction of the ultrasonic probe at the next moment; then, based on the availability of optical observation, it uses KalmanNet to adaptively correct or infer compensation the pose prediction, thereby obtaining the final spatial pose estimation result.
[0049] Within this algorithmic framework, inertial and mechanical information provides high-frequency continuous motion constraints, while optical observation information provides a low-frequency absolute spatial reference, thereby achieving robust and continuous estimation of the spatial pose of a freely handheld ultrasonic probe. Based on the above algorithmic idea, the multimodal fusion pose estimation algorithm will be described in detail below with specific embodiments.
[0050] This embodiment provides a detailed explanation of the multimodal fusion pose estimation algorithm based on the ultrasonic probe "motion prediction + correction or compensation" mechanism constructed in steps S4 and S5 above.
[0051] During motion prediction, the system constructs a two-layer bidirectional long short-term memory network (BiLSTM) to learn the nonlinear mapping relationship of the ultrasound probe's motion dynamics. Within each time window... Internally, the network simultaneously establishes forward motion state propagation and backward context state propagation, thereby comprehensively utilizing historical motion trends and short-term future information to improve the modeling ability for non-stationary handheld motions. The first layer of the network: This is used to learn the coupling relationship between IMU information and mechanical sensing information. This represents the dimension of the first layer of the network. The second layer of the network: This is used to model changes in the motion of the ultrasonic probe and changes in the mechanical scanning trajectory. This represents the dimension of the second layer of the network. The final predicted pose of the ultrasound probe is obtained: in Represents the output weight matrix. This represents the output bias. The prediction result will be used to update the model parameters in real time.
[0052] To address the issues of IMU drift and optical occlusion, KalmanNet is used to adaptively correct or infer compensation the pose prediction.
[0053] When optical observations are available, the predicted pose is first determined. Perform residual coding: And extract credibility features: in, The parameter is The residual feature encoding function is used for nonlinear mapping and feature extraction of the residuals. Subsequently, the network learns the dynamic Kalman gain matrix: This enables implicit modeling of noise covariance and adaptive allocation of multimodal weights. The parameter is The Kalman gain prediction function is then used. Finally, the predicted pose of the ultrasound probe is corrected using the dynamic Kalman gain matrix. in, This represents the pose correction mapping learned by KalmanNet during this process.
[0054] When optical observation is unavailable, pose correction mapping is used. A smooth recursion is performed to obtain the updated spatial pose state of the ultrasound probe: in, The empty set symbol indicates that no usable optical pose information exists. Finally, the algorithm outputs the spatial pose information of the ultrasound probe at an update frequency of no less than 120Hz and transmits it to the upper-level application module for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
[0055] In a specific embodiment of the present invention, the specific structure and parameter configuration of the two-layer bidirectional long short-term memory network (BiLSTM) in the multimodal fusion pose estimation algorithm are shown in Table 1.
[0056] Table 1. BiLSTM Model Structure and Parameters
[0057] Specifically, the execution logic of this BiLSTM network is as follows: The first step involves the input layer receiving a fusion vector sequence of mechanical-inertial measurement information after hardware-level time synchronization and coordinate transformation. The feature dimension of this input layer is represented as follows: ,in Batch Size. The length of the sliding time window (preferably 20 frames). The input feature number (preferably 9-dimensional, formed by concatenating a 3-dimensional pressure distribution vector and a 6-dimensional IMU inertial feature vector); In the second step, the input layer transmits the aforementioned feature sequence to the first bidirectional LSTM layer (BiLSTM_1). This layer utilizes forward and backward time propagation mechanisms to learn the temporal coupling features between mechanical sensing information and inertial measurement information. Its hidden layer has 128 neurons, and the output dimension after bidirectional feature concatenation is... The corresponding activation functions are the hyperbolic tangent function (Tanh) and the sigmoid function; In the third step, the output of the first bidirectional LSTM layer enters the fully connected coding layer (Dense_1), where spatial feature dimensionality reduction is performed through 128 neurons, resulting in an output dimension of... The modified linear unit (ReLU) is used as the activation function, and a regularization operation with a dropout rate of 0.2 is introduced to suppress overfitting. The fourth step involves inputting the dimensionality-reduced sequence features into a second bidirectional LSTM layer (BiLSTM_2) to further model the nonlinear trajectory variation trend of the handheld ultrasound probe movement. The output dimension after bidirectional splicing is... ; Fifth, the temporal output of the second bidirectional LSTM layer is truncated using a time-dimensional pooling layer, extracting only the last time step. The corresponding feature vectors are transformed to have the following dimensions: This enables the mapping of time series data to instantaneous spatial features. The sixth step involves mapping the instantaneous spatial features into a 6-dimensional continuous variable using a fully connected output layer (Dense_Out), with the output dimension... This corresponds to the 3D position prediction and 3D attitude prediction of the ultrasound probe in the world coordinate system. This layer uses a linear activation function to ensure the continuity of the regression calculation.
[0058] In a specific embodiment of the present invention, the specific structure and parameter configuration of the Kalman Network (KalmanNet) used to realize pose adaptive correction and inference compensation are shown in Table 2.
[0059] Table 2 KalmanNet Model Structure and Parameters
[0060] Specifically, the execution logic of this KalmanNet network is as follows: The first step is to receive a 12-dimensional input vector from the current state space equation at the input layer, where the dimension is represented as... The input vector is formed by nonlinearly concatenating and recombining the 6-dimensional observation residual (i.e., the difference between the predicted pose and the reference pose) and the 6-dimensional state prediction residual (i.e., the difference between the predicted pose of the current step and the predicted pose after correction of the previous step). In the second step, the 12-dimensional input vector enters a fully connected encoding layer (FC_Encode), which contains 64 neurons and uses a rectified linear unit (ReLU) as the activation function to map the low-dimensional residual to a high-dimensional sparse feature space. The output dimension is... ; Third, the encoded high-dimensional residual features are input into a standard gated recurrent unit (GRU_Cell) layer. The hidden layer dimension of this layer is set to 64. Its internal update and reset gate mechanisms store and transmit the temporal dependencies in historical states. The internal activation function is the hyperbolic tangent function (Tanh), and the output dimension remains constant. This is used to implicitly model the variation of nonlinear noise covariance; The fourth step involves the gain mapping layer (FC_Gain) receiving the output of the gated recurrent unit layer and performing a nonlinear transformation through a fully connected layer containing 36 neurons, resulting in an output dimension of... This layer employs a leaky modified linear unit (LeakyReLU, slope coefficient) () is used as an activation function to prevent gradient vanishing and preserve negative residual information; The fifth step involves using a matrix reconstruction layer to rearrange and reconstruct the 36-dimensional one-dimensional vector output from the aforementioned fully connected layer according to spatial geometric mapping relationships. The two-dimensional real matrix is the dynamic Kalman gain matrix required for pose correction.
[0061] In practical deployment, this invention is based on a deep learning framework built using Python 3.12 and PyTorch 2.8. The input features of the BiLSTM undergo Z-score normalization preprocessing to eliminate the influence of different sensor dimensions. During the training phase, the model uses the AdamW optimizer for gradient descent, with an initial learning rate set to [value missing]. A cosine annealing strategy is introduced for dynamic adjustment, with the minimum learning rate set to... This is to prevent the model from overfitting during training.
[0062] In a specific embodiment of the present invention, the method for updating the BiLSTM model and KalmanNet network parameters online in step S4 is as follows: the BiLSTM model and KalmanNet are not pre-trained on an offline dataset during the system initialization phase, but during the ultrasound scanning process, optical observation information that meets preset conditions is used as a reference ground value to update the network parameters online in real time. The specific steps are as follows.
[0063] Online update status determination and triggering: The system monitors in real time using a monocular high frame rate industrial camera. The confidence level of the optical pose returned by the time-of-flight optical positioning module The system sets a composite trigger condition to determine whether the optical positioning module is in a high-confidence available state. The composite trigger condition includes: (1) Confidence of optical pose at the current sampling time ; (2) The state of optical pose confidence exceeding the preset threshold is maintained for more than 15 consecutive frames.
[0064] When the system detects that both of the above conditions are met simultaneously, it determines that the current optical positioning module is in a visible and usable state, and the system automatically triggers the online real-time update mode of network parameters; if either of the above two conditions is not met, it determines that the optical positioning module is in an occluded or unusable state, and the system automatically switches to the inference compensation mode described in step S5.
[0065] Construction of the online dynamic loss function: When the online real-time update mode is triggered, the system directly uses the spatially registered 3D absolute pose data obtained by the optical positioning module within the current sliding time window as the instantaneous true label of the current operator's scanning state. To simultaneously optimize the motion prediction accuracy of the BiLSTM model and the Kalman gain output of Kalman Net, the system constructs a two-stage cascaded backpropagation loss function: in, This is the BiLSTM motion prediction loss function, which is used to constrain the accuracy of the BiLSTM model in predicting the current spatial pose of the ultrasonic probe by fusing vector sequences with mechanical-inertial measurement information. The loss function for KalmanNet is modified to constrain the accuracy of the dynamic Kalman gain matrix output by KalmanNet in correcting the predicted pose. In the formula, The total number of sample frames within the sliding time window that represents online updates, preferably in the range of 12 to 20 frames; The first one representing the output of the BiLSTM model The frame prediction pose vector includes a 3D position prediction component and a 3D attitude prediction component. Represents the corresponding number Frame optical instantaneous real-time labeling. Representing the The dynamic Kalman gain matrix output by KalmanNet for frames The corrected predicted pose vector.
[0066] Convergence and Hard Real-Time Termination Control for Online Fine-Tuning: In the online update mode of the network parameters, to ensure the real-time operation of the entire spatial perception system while optimizing network weights, the system employs a real-time gradient descent mechanism with a limited iteration step size. Specifically, during the processing interval of each frame, the backpropagation iteration of the online gradient forcibly stops training at the current time step when any of the following convergence termination conditions are met: (1) The difference in the composite loss function between two consecutive backpropagation iterations ; (2) The number of backpropagation iterations within the current sampling time step reaches the preset maximum limit. ,in Preferably, it is 3 to 5 times.
[0067] The system sets the maximum number of attempts. The computation time for updating online parameters in a single frame is limited to... This ensures that the entire system operates within a closed loop, simultaneously scanning, training, and correcting.
[0068] When the optical positioning module fails to meet the composite triggering condition due to occlusion, the system executes step S5. At this point, the system immediately stops receiving optical pose, cuts off the backpropagation fine-tuning path, and stops updating network parameters. The system inputs the fusion vector sequence of mechanical-inertial measurement information into the BiLSTM model that has completed online fine-tuning parameter updates, outputs the predicted pose, and uses KalmanNet to perform inference compensation on the predicted pose based on historical poses, thereby maintaining continuous output of ultrasonic probe spatial pose information during the absence of optical signals.
[0069] To further verify the positioning accuracy and stability of the proposed freehand ultrasound spatial perception method based on multi-sensor fusion in a real clinical complex scanning environment, a series of continuous neck phantom scanning experiments were designed. Senior clinicians held an ultrasound probe equipped with the device of this invention and performed a continuous S-shaped trajectory scan along the surface of the phantom, with a total scanning time of 60 seconds.
[0070] To comprehensively test the system's robustness, the optical positioning module was artificially occluded and restored during the scanning process, dividing the entire experiment into the following three continuous dynamic phases: (1) Stage I, 0~20 seconds, optically visible state: the camera is unobstructed, the system is in the online update mode of network parameters, the optical pose is used as the absolute spatial reference to correct the BiLSTM prediction results, and the network parameters are updated online; (2) Stage II, 20~40 seconds, optical occlusion state: the camera line of sight is completely blocked by human, the system triggers occlusion judgment, stops back propagation, automatically switches to collaborative prediction compensation mode, and relies entirely on the mechanical-IMU fusion vector sequence to perform pose unsupervised inference prediction. (3) Stage III, 40~60 seconds, restore optical visibility: remove the obstruction, the system detects that the optical confidence has been restored and maintained continuously, and re-triggers the online real-time update mode.
[0071] By using a high-precision electromagnetic tracking system additionally installed on the probe as a third-party truth reference, the true trajectory of the ultrasonic probe in the world coordinate system is recorded and compared with the pose output by the method of this invention to calculate the continuous linear displacement absolute positioning error and angular displacement absolute positioning error.
[0072] Depend on Figure 4 Analysis of the experimental results presented shows that: (1) Stage I: In the initial stage of system startup, since the BiLSTM model has not yet fully fitted the dynamic mechanical characteristics of the current scanning action, the initial linear displacement error is about 0.25 mm and the angular displacement error is about 0.20°. Subsequently, as the system triggers the online real-time update mode, the backpropagation mechanism begins to update the parameters online in real time. During the period from 0 to 12 seconds, the linear displacement and angular displacement errors show a significant fluctuating downward trend along with the dynamic update of the network parameters; after 12 seconds, the network weights approach the optimal point, and the system error enters a high-precision stable period. The linear displacement error finally stabilizes at around 0.08 mm and the angular displacement error stabilizes at around 0.06°, which fully demonstrates the rapid convergence and adaptive spatial perception capability of the online real-time update mode. (2) Stage II: During the 20-second gap when the optical signal is completely lost, the system automatically switches to the collaborative prediction compensation mode, at which point the inference compensation capability of KalmanNet begins to play its role. Despite the loss of absolute optical pose reference, the method of this invention does not degenerate into the traditional pure inertial navigation mode, but uses the dynamic Kalman gain matrix output in real time by KalmanNet to perform unsupervised inference constraints and real-time state correction on the pose prediction output by the BiLSTM model. At the end of the occlusion period (40 seconds), the cumulative linear displacement absolute error of the system is only 0.42 mm, and the angular displacement error is only 0.32°. This result proves that the system still has good positioning accuracy in extreme environments without absolute optical reference; (3) Stage III: After the occlusion is removed at 40 seconds, the system detects a high-confidence optical pose and instantly switches back to the online real-time update mode. Due to the high-precision compensation of KalmanNet in Stage II, the overall error is always kept at a low sub-millimeter level. After the system reactivates gradient backpropagation, the two-stage cascaded loss function does not experience gradient explosion, but completes a second fast convergence within 1.5 seconds and stabilizes again in a lower error range.
Claims
1. A freehand ultrasonic spatial sensing method based on multi-sensor fusion, characterized in that, Includes the following steps: S1: Obtain the initial spatial pose information of the ultrasonic probe; S2: Simultaneously acquire optical pose information, mechanical sensing information and inertial measurement information to obtain multimodal raw observation data; S3: Perform time synchronization, preprocessing and coordinate system transformation on the multimodal raw observation data to obtain optical pose feature vector, mechanical feature vector and inertial feature vector, forming a mechanical-inertial measurement information fusion vector sequence; S4: When the optical positioning module is visible, the BiLSTM model is used to predict the pose based on the fusion vector sequence of mechanical-inertial measurement information. Then, KalmanNet is used to correct the result according to the optical pose, and the parameters of the BiLSTM model and KalmanNet network are updated online. S5: When the optical positioning module is blocked, the BiLSTM model is used to predict the pose based on the fusion vector sequence of mechanical-inertial measurement information, and then KalmanNet is used to infer and compensate the result based on the historical pose. S6: Outputs the spatial pose information of the ultrasound probe for use in Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
2. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for S1 is: S101: Install the three-dimensional mechanical sensor, six-axis IMU and optical positioning module at predetermined positions around the Freehand ultrasound probe to establish the initial working state of the ultrasound probe; S102: Start the camera and scan the optical positioning module under unobstructed conditions. Take its optical pose information during the initialization phase and calculate the average value as the initial spatial pose information of the ultrasonic probe in the world coordinate system.
3. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for S2 is as follows: S201: Real-time acquisition of pressure distribution information between the operator's hand and the ultrasonic probe handle via a three-dimensional mechanical sensor located at the handle of the ultrasonic probe; S202: Real-time acquisition of angular velocity and linear acceleration information of the ultrasonic probe via a six-axis IMU; S203: Acquires real-time spatial pose information of the ultrasonic probe under unobstructed conditions through an optical positioning module and camera.
4. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for S3 is as follows: S301: Hardware-level time synchronization of optical pose information, mechanical sensing information and inertial measurement information is performed through FPGA to generate a unified timestamp and obtain synchronized multi-source sensing data. S302: Perform data preprocessing operations on the synchronized multi-source sensor data, including filtering and feature extraction, including: For inertial measurement information, the acceleration and angular velocity are low-pass filtered, the angular acceleration is calculated, and the inertial feature vector characterizing the instantaneous motion state of the ultrasonic probe is obtained, which is used to describe the instantaneous motion state of the ultrasonic probe. For mechanical sensing information, the pressure intensity vector is calculated based on the resistance change of the three-dimensional mechanical sensor, the pressure center is calculated, the pressure change rate and the overall force intensity are calculated, and a mechanical feature vector is obtained to describe the changes in the operator's hand force application pattern. For optical pose information, Kalman filtering is used to smooth the pose sequence to obtain optical pose feature vectors, and optical pose confidence is calculated for subsequent occlusion determination. S303: Map the three types of feature vectors extracted in step S302 to the ultrasonic probe coordinate system so that the three types of feature vectors describe the same spatial pose state; forming a continuous mechanical-inertial measurement information fusion vector sequence.
5. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for S4 is as follows: S401: Use the optical pose feature vector obtained in the preceding steps as the reference pose; S402: Input the fused vector sequence of mechanical-inertial measurement information obtained in the previous steps into BiLSTM, and output the predicted pose; S403: The residual between the predicted pose and the reference pose is learned using KalmanNet to correct the predicted pose, and the parameters of the BiLSTM model and KalmanNet network are updated online.
6. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for determining whether the optical positioning module is in a visible state is as follows: The optical pose confidence is calculated. When the output of the optical positioning module is stable and the optical pose confidence of multiple consecutive frames is higher than the preset threshold, the optical positioning module is determined to be visible and in a usable state.
7. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 6, characterized in that, The method for determining whether the optical positioning module is occluded is as follows: when the optical positioning module loses its pose or the confidence of the optical pose is lower than a preset threshold for multiple consecutive frames, the optical module is determined to be occluded and unusable.
8. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, S5 include: S501: Stop receiving optical pose, input the fusion vector sequence of mechanical-inertial measurement information into the BiLSTM model, and output the predicted pose; S502: KalmanNet is used for inference compensation based on historical poses to suppress IMU drift error.
9. The Freehand Ultrasonic Spatial Sensing Method Based on Multi-Sensor Fusion according to claim 1, characterized in that, The method for S6 is as follows: S601: Outputs the spatial pose information of the ultrasonic probe at a frequency of not less than 120 Hz; S602: Time-align the spatial pose information of the ultrasound probe with the ultrasound image frame. S603: Used for Freehand ultrasound 3D reconstruction, spatial registration, or intraoperative navigation applications.
Citation Information
Patent Citations
Multi-sensor fusion positioning system and positioning method
CN113124852A
Hand-held probe three-dimensional scanning imaging system and method
CN118161194A