An inter-image relative position prediction method based on inertial measurement data
By combining inertial measurement unit data with ultrasound images and using a deep learning model to predict the relative positions between ultrasound images, the drift error problem in deep learning technology is solved, and higher-precision three-dimensional ultrasound reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2022-06-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing deep learning techniques are easily affected by inter-image displacement and cumulative drift errors when estimating the relative positions of ultrasound images, leading to a decrease in the accuracy of 3D ultrasound reconstruction.
By combining inertial measurement unit (IMU) data with ultrasonic images, orientation and acceleration information are determined by acquiring IMU data from ultrasonic images. A deep learning model is used to predict relative position information, and the relative positions between images are predicted by combining IMU data and image sequences.
It improves the accuracy of relative position prediction between images, reduces drift error, and enhances the precision and stability of three-dimensional ultrasound reconstruction.
Smart Images

Figure CN115345933B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ultrasound imaging, and more particularly to a method for predicting the relative position between images based on inertial measurement data. Background Technology
[0002] Ultrasound imaging, with its advantages of safety, portability, and low cost, has become one of the main diagnostic tools in clinical practice. Three-dimensional ultrasound is widely used due to its intuitive display, ease of interaction, and rich clinical information. There are three acquisition methods for three-dimensional ultrasound: mechanical probes, electronic phased arrays, and free-form acquisition. Compared to the limitations of mechanical probes or electronic phased arrays in terms of field of view and operability, free-form acquisition offers advantages in flexibility and convenience. It mainly reconstructs the ultrasound volume by calculating the relative positions of a series of ultrasound images.
[0003] For free-form 3D ultrasound reconstruction, early reconstruction schemes primarily relied on external positioning systems, such as electromagnetic or optical positioning, using complex, expensive, and interference-prone external sensors to provide accurate estimates of ultrasound image positions. Schemes independent of external positioning mainly utilize speckle decorrelation, which estimates relative motion by leveraging the correlation of speckle patterns between adjacent ultrasound images and decomposes this relative motion into intra-image and extra-image components. However, reconstruction quality is easily affected by scan rate and angle. Current approaches primarily employ deep learning-based techniques, inputting ultrasound images into a deep learning model to estimate the relative positions of the ultrasound images for final 3D reconstruction. For example, analogies are used between deep learning models and traditional speckle decorrelation methods to estimate relative motion in ultrasound images; attention mechanisms are used to mine correlation information from multiple ultrasound images for 3D reconstruction; and consistency constraints and shape priors are used to uncover inherent clues in the ultrasound image sequence to improve reconstruction performance. However, current deep learning techniques rely solely on ultrasound images to estimate relative positions, making them susceptible to inter-image displacement and accumulated drift errors.
[0004] Therefore, existing technologies still need improvement and development. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for predicting the relative position between images based on inertial measurement data, which addresses the above-mentioned deficiencies of the prior art. This method aims to solve the problem that the existing deep learning technology relies solely on ultrasound images to estimate the relative position, and is easily affected by the displacement between images and the accumulated drift error.
[0006] The technical solution adopted by this invention to solve the problem is as follows:
[0007] In a first aspect, embodiments of the present invention provide a method for predicting the relative position between images based on inertial measurement data, wherein the method includes:
[0008] Acquire several ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence;
[0009] Based on the inertial measurement unit data corresponding to each of the ultrasonic images, determine the direction information and acceleration information corresponding to each of the ultrasonic images;
[0010] Based on each ultrasound image, and the direction and acceleration information corresponding to each ultrasound image, the relative position information corresponding to each ultrasound image is determined, wherein the relative position information is used to reflect the relative position change between two adjacent ultrasound images.
[0011] Secondly, embodiments of the present invention provide an image-to-image relative position prediction device based on inertial measurement data, wherein the device includes:
[0012] The data acquisition module is used to acquire several ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence;
[0013] The data conversion module is used to determine the direction information and acceleration information corresponding to each of the ultrasound images based on the inertial measurement unit data corresponding to each of the ultrasound images.
[0014] The information prediction module is used to determine the relative position information of each ultrasound image based on each ultrasound image, the direction information and acceleration information corresponding to each ultrasound image, wherein the relative position information is used to reflect the relative position change between two adjacent ultrasound images.
[0015] Thirdly, embodiments of the present invention provide a terminal, wherein the terminal includes a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing the image-to-image relative position prediction method based on inertial measurement data as described above; and the processor is used to execute the programs.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions, wherein the instructions are adapted to be loaded and executed by a processor to implement the steps of any of the above-described methods for predicting relative positions between images based on inertial measurement data.
[0017] The beneficial effects of this invention are as follows: This invention acquires several ultrasound images and corresponding inertial measurement unit (IMU) data for each ultrasound image, wherein each ultrasound image is located in the same image sequence. Based on the IMU data corresponding to each ultrasound image, the orientation and acceleration information corresponding to each ultrasound image are determined. Based on each ultrasound image and its corresponding orientation and acceleration information, the relative position information corresponding to each ultrasound image is determined, wherein the relative position information reflects the relative position change between two adjacent ultrasound images. This invention predicts the relative position information between images by combining IMU data and image sequences, solving the problem that existing deep learning techniques rely solely on ultrasound images to estimate relative positions, making them susceptible to the influence of image displacement and accumulated drift errors. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the relative position prediction method between images based on inertial measurement data provided in an embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of the process of three-dimensional reconstruction of ultrasound images provided in an embodiment of the present invention.
[0021] Figure 3 This is a schematic diagram of the combination of the inertial measurement unit and the ultrasonic probe provided in an embodiment of the present invention.
[0022] Figure 4 This is a schematic diagram of the target network structure provided in an embodiment of the present invention.
[0023] Figure 5 This is a schematic diagram of the adaptive optimization process of the target network provided in an embodiment of the present invention.
[0024] Figure 6 This is a schematic diagram of a module for a relative position prediction device between images based on inertial measurement data provided in an embodiment of the present invention.
[0025] Figure 7 This is a schematic diagram of the terminal provided in the embodiment of the present invention. Detailed Implementation
[0026] This invention discloses a method for predicting the relative position between images based on inertial measurement data. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.
[0027] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0028] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0029] To address the aforementioned shortcomings of existing technologies, this invention provides a method for predicting the relative position between images based on inertial measurement data. The method acquires several ultrasound images and corresponding inertial measurement unit (IMU) data for each ultrasound image, wherein the ultrasound images are located in the same image sequence. Based on the IMU data corresponding to each ultrasound image, orientation and acceleration information are determined for each ultrasound image. Based on the ultrasound images, their corresponding orientation and acceleration information, the relative position information for each ultrasound image is determined, whereby the relative position information reflects the relative position change between two adjacent ultrasound images. This invention, by combining IMU data and the image sequence to predict the relative position information between images, solves the problem that existing deep learning techniques, which rely solely on ultrasound images to estimate relative positions, are easily affected by inter-image displacement and accumulated drift errors.
[0030] like Figure 1 As shown, the method includes the following steps:
[0031] Step S100: Acquire a number of ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence.
[0032] Specifically, in this embodiment, each ultrasound image is an image from the same image sequence. Since predicting the relative position changes between images solely based on ultrasound images is easily affected by image displacement and accumulated drift errors, this embodiment needs to obtain inertial measurement unit data corresponding to each ultrasound image. The acceleration and orientation information provided by the inertial measurement unit data can effectively improve the accuracy of predicting the relative position between images.
[0033] In one implementation, each of the ultrasound images is acquired based on a preset ultrasound imaging device, and the inertial measurement unit data corresponding to each of the ultrasound images is acquired based on a preset inertial measurement unit located on the ultrasound imaging device.
[0034] Specifically, an inertial measurement unit (IMU) is a sensor that integrates a three-axis accelerometer, gyroscope, and magnetometer, capable of measuring an object's three-axis attitude angles and acceleration. Compared to electromagnetic or optical positioning systems, IMUs are low-cost and small in size; therefore, integrating an IMU into an ultrasonic imaging device does not increase the complexity of the ultrasonic imaging scan. This embodiment integrates the IMU and the ultrasonic imaging device, which can effectively improve the prediction accuracy of relative position changes between images.
[0035] For example, assuming the ultrasound imaging device is an ultrasound probe, integrating an inertial measurement unit into the ultrasound probe does not increase the complexity of the ultrasound scan. Figure 3 As shown, an inertial measurement unit (IMU) is integrated with an ultrasound probe. Multiple ultrasound images are acquired through the probe, and the IMU provides acceleration and orientation information for each image to predict relative position changes between them. In a clinical setting, the addition of the IMU does not alter the ultrasound physician's usage habits and avoids the susceptibility to interference and system complexity of traditional external positioning systems. It also improves upon the displacement and drift errors inherent in estimating relative position solely based on ultrasound images, thus enhancing the accuracy of predicting relative position changes between images.
[0036] like Figure 1 As shown, the method further includes the following steps:
[0037] Step S200: Determine the direction information and acceleration information corresponding to each ultrasonic image based on the inertial measurement unit data corresponding to each ultrasonic image.
[0038] Specifically, since the inertial measurement unit and the ultrasonic imaging device are integrated together, for each ultrasonic image, the inertial measurement unit data corresponding to that ultrasonic image can be used to indicate the direction and acceleration of that ultrasonic image.
[0039] In one implementation, the direction information is Euler angle information, and the data of each inertial measurement unit includes measured Euler angle information and measured acceleration information. Step S200 specifically includes the following steps:
[0040] Step S201: For each ultrasound image, calculate the rotation matrix based on the measured Euler angle information corresponding to the ultrasound image, and determine the Euler angle information corresponding to the ultrasound image, wherein the measured Euler angle information corresponds to the northeast-northeast coordinate system, and the Euler angle information corresponds to the coordinate system of the ultrasound image;
[0041] Step S202: Determine the gravity direction based on the measured Euler angle information corresponding to the ultrasound image, and determine the acceleration information corresponding to the ultrasound image based on the measured acceleration information and the gravity direction, wherein the acceleration information is determined based on the difference between the measured acceleration information and the component of the measured acceleration information in the gravity direction.
[0042] Specifically, since the inertial measurement unit (IMU) and the ultrasonic imaging device correspond to different coordinate systems, the measured Euler angle information output by the IMU needs to be transformed to obtain the Euler angle information corresponding to each ultrasonic image. Similarly, the measured acceleration information output by the IMU needs to have the gravitational component subtracted to obtain the acceleration information corresponding to each ultrasonic image. In one implementation, the average value of the acceleration in each direction can be adjusted to 0 to reduce the influence of noise.
[0043] For example, an ultrasound image sequence of length N acquired from an ultrasound probe is defined as I = {I...} i |i=1,2,…,N}, the inertial measurement unit acquires spatial position and motion information for each frame of ultrasound image, including Euler angles O in the northeast-northeast coordinate system. i =(O x O y O z ) i and acceleration A in three spatial directions i =(A x A y A z ) i For Euler angles, it is necessary to convert them into relative position transformations Φ between adjacent ultrasound images. i =(Φ x ,Φy ,Φ z ) i As shown in formula (1):
[0044] Φ i =M -1 (M(O i ) -1 *M(O i+1 ),i=1,2,…,N-1#(1)
[0045] Where M(·) transforms the direction vector into a 3×3 rotation matrix, M -1 (·) represents the inverse operation of M(·). For acceleration, it is necessary to subtract the value from O. i The calculated component of gravity in the direction of g is g. i Meanwhile, the average value of the acceleration sequence is adjusted to 0 to reduce the influence of noise, as shown in formula (2):
[0046]
[0047] like Figure 1 As shown, the method further includes the following steps:
[0048] Step S300: Based on each ultrasound image and the direction and acceleration information corresponding to each ultrasound image, determine the relative position information corresponding to each ultrasound image, wherein the relative position information is used to reflect the relative position change between two adjacent ultrasound images.
[0049] Specifically, an inertial measurement unit (IMU) can measure the three-axis attitude angles and acceleration of an object. IMU-assisted positioning can maintain the correct trend over a large range. Therefore, this embodiment uses the ultrasonic images acquired by the ultrasonic imaging device and the direction and acceleration information of each ultrasonic image provided by the IMU to comprehensively predict the relative position change between two adjacent ultrasonic images.
[0050] In one implementation, the direction information is Euler angle information, and step S300 specifically includes the following steps:
[0051] Step S301: Input each ultrasound image, the Euler angle information and the acceleration information corresponding to each ultrasound image into a pre-trained target network, and obtain the relative position information through the target network.
[0052] In simple terms, this embodiment requires training a deep learning model, namely the target network, on the training dataset to predict the relative position changes between two adjacent ultrasound images in an image sequence. Since the target network has already learned the image features, Euler angle features, and acceleration features corresponding to different relative position information based on the training dataset, inputting each ultrasound image, its corresponding acceleration information, and Euler angle information into the target network will yield the relative position information for each ultrasound image.
[0053] In one implementation, the relative position information includes several relative transformation parameters corresponding to two adjacent frames in each of the ultrasound images. The several relative transformation parameters include several translation variables and several rotation angles. The several translation variables correspond to different directions, and the several rotation angles correspond to different directions.
[0054] For example, adjacent ultrasound images {(I i ,I i+1 The relative transformation parameter θ = {θ | i = 1, 2, ..., N-1} between these parameters is θ. i |1,2,…,N-1}, where θ i Represents three translation variables t i =(t x ,t y ,t z ) i and 3 rotation angles φ i =(φ x ,φ y ,φ z ) i .
[0055] In one implementation, the target network includes a feature extraction module, a fusion module, and a prediction network, and step S301 specifically includes the following steps:
[0056] Step S3011: Input each ultrasound image, the acceleration information and the direction information corresponding to each ultrasound image into the feature extraction module to obtain the image features, acceleration features and Euler angle features corresponding to each ultrasound image. The image features corresponding to two adjacent ultrasound images are used to reflect the relative distance information between the two ultrasound images.
[0057] Step S3012: Input the image features and acceleration features corresponding to each ultrasound image into the fusion module to obtain the velocity features corresponding to each ultrasound image.
[0058] Step S3013: Input the image features, velocity features and Euler angle features corresponding to each ultrasound image into the prediction network to obtain the relative position information.
[0059] Specifically, the target network in this embodiment mainly includes three modules. The first module is a feature extraction module, which takes each ultrasound image and its Euler angle and acceleration information as input, and outputs the image features, acceleration features, and Euler angle features of each ultrasound image. The second module is a feature fusion module. Since the image features of each ultrasound image implicitly contain relative distance information between images, when the sampling time between ultrasound images is constant, the relative distance information can reflect the scanning speed. Therefore, the image features and acceleration features of each ultrasound image are correlated and can be complementary. So, in this embodiment, the image features and acceleration features of each ultrasound image are used as input to the fusion module, which fuses the two features and outputs the velocity features of each ultrasound image. The third module is a prediction network for predicting the relative position changes between images. This prediction network takes the image features, velocity features, and Euler angle features of each ultrasound image as input, and outputs the relative position information corresponding to each ultrasound image. This relative position information can reflect the relative position changes between two adjacent ultrasound images.
[0060] In one implementation, such as Figure 4 As shown, the feature extraction module includes a residual network, a first fully connected layer, and a second fully connected layer. The input of the residual network is each ultrasound image, and the output of the residual network is the image feature corresponding to each ultrasound image. The input of the first fully connected layer is the acceleration information corresponding to each ultrasound image, and the output of the first fully connected layer is the acceleration feature corresponding to each ultrasound image. The input of the second fully connected layer is the Euler angle information corresponding to each ultrasound image, and the output of the second fully connected layer is the Euler angle feature corresponding to each ultrasound image.
[0061] Specifically, residual networks are powerful feature extractors, therefore this embodiment uses residual networks to extract image features from each ultrasound image. The first fully connected layer maps the input acceleration information to a high-dimensional space and reshapes it into two-dimensional acceleration features, thus facilitating the combination of acceleration features and image features. The second fully connected layer extracts features from the input Euler angle information and feeds it into the prediction network in the target network to enhance the Euler angle estimation in the relative position information output by the prediction network.
[0062] In one implementation, the fusion module includes a fusion unit and a feature enhancement network, and step S3012 specifically includes the following steps:
[0063] Step S30121: Input the image features and acceleration features corresponding to each ultrasound image into the fusion unit to obtain the fusion velocity features corresponding to each ultrasound image.
[0064] Step S30122: Input the fusion velocity features corresponding to each of the ultrasound images into the feature enhancement network to obtain the velocity features corresponding to each of the ultrasound images.
[0065] Simply put, such as Figure 4 As shown, to enhance the velocity features, this embodiment adds an additional feature enhancement network. Specifically, the image features and acceleration features of each ultrasound image are input into the fusion unit of the fusion module, so that the acceleration features are added to the image features, thereby constructing fused velocity features. Then, the fused velocity features of each ultrasound image are input into the feature enhancement network, which reduces invalid or interfering information in the fused velocity features, and finally outputs the final velocity features of each ultrasound image.
[0066] For example, consider the acceleration characteristic f a Image features f added to the latent velocity features to construct the fused velocity features f v Specifically, as shown in formula (3):
[0067]
[0068] In one implementation, such as Figure 4 As shown, both the feature enhancement network and the prediction network are constructed using long short-term memory networks.
[0069] Specifically, such as Figure 4As shown, the target network in this embodiment is actually a temporal and multi-branch fusion network. It predicts the relative position information between ultrasound images by fusing each ultrasound image with its corresponding acceleration and Euler angle information. Since Long Short-Term Memory (LSTM) networks can process temporal information and memorize the knowledge between all historical image frames as contextual information, this embodiment uses an LTM network to generate a feature enhancement network and a prediction network to help estimate the relative position information between future image frames. For the LTM network used for feature enhancement, to minimize the impact of acceleration noise on the target network performance within short sequence intervals, the fused velocity features output by the fusion unit are input into the LTM network to enhance the velocity features using temporal context information, resulting in the final velocity features. The output of the LTM network is merged into the main branch and combined with the output of the residual network. By fusing the acceleration information from the inertial measurement unit, the multi-branch structure can better estimate the displacement between images. Furthermore, for the long short-term memory network used to predict relative position information, the Euler angle information determined by the inertial measurement unit is fed into multiple fully connected layers and connected to the corresponding long short-term memory network in the main branch as one of the input data to enhance the Euler angle estimation in the relative position of the ultrasound image.
[0070] In one implementation, the target network is a pre-trained network, and the training process of the target network includes:
[0071] Step S10: Obtain the training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence respectively; input the training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence into the initial network to obtain the predicted relative position information corresponding to the training image; wherein, the initial network is the target network that has not been fully trained.
[0072] Step S11: Obtain the standard relative position information corresponding to the training image sequence, and determine the first loss value corresponding to the initial network based on the predicted relative position information and the standard relative position information;
[0073] Step S12: Update the network parameters of the initial network according to the first loss value to obtain an updated network. Determine whether the updated network has converged to the target value. If not, continue to execute the steps of obtaining the training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence, until the obtained updated network converges to the target value, and obtain the trained target network.
[0074] Specifically, since the target network uses a deep learning model, it requires a large amount of training and testing data to train. This embodiment pre-constructs a data acquisition system to acquire multiple training image sequences, orientation and acceleration information provided by the inertial measurement units corresponding to each training image frame in each training image sequence, and the true standard relative position information corresponding to each training image frame in each training image sequence. During the training phase, the untrained initial network outputs the predicted relative position information corresponding to each training image frame based on the input training image sequence and the orientation and acceleration information of each training image frame in that sequence. Since the initial network also obtains the true standard relative position information, it continuously optimizes the error between the predicted relative position information and the standard relative position information, making the prediction results increasingly accurate until it converges to the target value, at which point training stops, resulting in the trained target network.
[0075] In one implementation, the true standard relative position information can be obtained using an electromagnetic locator.
[0076] In one implementation, the loss function of the initial network consists of two terms. The first term minimizes the predicted relative position information of the output. The first term is the mean absolute error between the relative position information θ and the standard. The second term is the Pearson correlation loss, used to understand the overall trend of the ultrasound scan.
[0077] In one implementation, the initial network can also be subjected to deep supervision to compare the orientation and acceleration information provided by the inertial measurement unit at the feature level, so as to improve the drift error accumulated in the initial network and thus obtain the trained target network.
[0078] In one implementation, the method further includes the following steps:
[0079] Step S400: Based on the relative position information, determine the predicted acceleration information and predicted Euler angle information corresponding to each of the ultrasound images;
[0080] Step S401: Determine the acceleration loss value corresponding to the target network based on the acceleration information and the predicted acceleration information corresponding to each of the ultrasound images;
[0081] Step S402: Determine the Euler angle loss value corresponding to the target network based on the Euler angle information and the predicted Euler angle information corresponding to each ultrasound image;
[0082] Step S403: Determine the second loss value corresponding to the target network based on the acceleration loss value and the Euler angle loss value;
[0083] Step S404: Update the network parameters of the target network according to the second loss value.
[0084] In short, most deep learning networks rely primarily on offline training and direct online inference. However, this strategy struggles to handle data with different distributions than the training data during online inference. Therefore, this embodiment introduces a multimodal learning strategy, such as... Figure 5 As shown, the target network can perform online self-supervised learning during online inference. Specifically, the target network can adaptively optimize based on the Euler angle and acceleration information corresponding to each ultrasonic image provided by the inertial measurement unit as weak labels, thereby reducing drift error and mitigating the impact of acceleration noise.
[0085] In one implementation, the second loss value is calculated as follows: the predicted acceleration information corresponding to the centroid of each ultrasound image. It is based on the relative position transformation of the target network output. Translation in The calculated acceleration, scaled to match the average zeroing of the inertial measurement unit, is shown in Equation (4):
[0086]
[0087] in, express The translation in the inverse transform. The second loss value corresponding to the adaptive optimization of the target network is determined based on two parts: the acceleration loss value and the Euler angle loss value. Specifically, Pearson correlation loss is used to measure and predict acceleration information. The difference between the acceleration information A provided by the inertial measurement unit and the acceleration information A is used to obtain the acceleration loss value. This acceleration loss value allows for the accurate measurement of acceleration trends over a large range during ultrasonic scanning using the inertial measurement unit, thereby improving the estimation of the fusion network and mitigating the effects of noise. Simultaneously, this embodiment also requires measuring the predicted Euler angle information. The Euler angle loss value is obtained by calculating the average absolute error between the Euler angle information Φ provided by the inertial measurement unit and the Euler angle information Φ. The adaptive optimization method for the target network provided in this embodiment allows the target network to automatically optimize its network parameters even during the online inference phase, improving the performance of the target network in processing out-of-distribution data.
[0088] In one implementation method, the method further includes:
[0089] Step S500: Perform three-dimensional reconstruction on each of the ultrasound images based on the relative position information.
[0090] Specifically, since relative position information can reflect the relative transformation relationship between two adjacent ultrasound images, the position of each non-first frame ultrasound image relative to the first frame ultrasound image can be calculated based on the relative position information, thereby completing the three-dimensional reconstruction of each ultrasound image (e.g., Figure 2 (As shown). In one implementation, unstructured interpolation is used to reconstruct three dimensions from each ultrasound image, and the images are then rendered on a graphics processor to improve data processing efficiency and reduce waiting time.
[0091] The advantages of this invention are:
[0092] 1) This invention proposes a novel and effective free-form 3D ultrasound reconstruction scheme. Specifically, this scheme combines an inertial measurement unit (IMU) with ultrasound images, fully exploits the complementary information between the IMU and the ultrasound images, uses deep learning technology to estimate the relative positions of adjacent ultrasound images, and performs 3D ultrasound reconstruction and rendering.
[0093] 2) The online learning strategy included in the scheme of the present invention optimizes the deep learning model to improve the accumulated drift error by comparing the difference between the information provided by the inertial measurement unit and the information estimated by the deep learning model during the inference stage of the deep learning model.
[0094] Based on the above embodiments, the present invention also provides an image-to-image relative position prediction device based on inertial measurement data, such as... Figure 6 As shown, the device includes:
[0095] The data acquisition module 01 is used to acquire several ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence;
[0096] The data conversion module 02 is used to determine the direction information and acceleration information corresponding to each of the ultrasound images based on the inertial measurement unit data corresponding to each of the ultrasound images.
[0097] The information prediction module 03 is used to determine the relative position information of each ultrasound image based on each ultrasound image and the direction information and acceleration information corresponding to each ultrasound image, wherein the relative position information is used to reflect the relative position change between two adjacent ultrasound images.
[0098] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 7As shown, the terminal includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for predicting relative positions between images based on inertial measurement data. The display screen can be a liquid crystal display (LCD) or an e-ink display.
[0099] Those skilled in the art will understand that Figure 7 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0100] In one implementation, the terminal's memory stores one or more programs, configured to be executed by one or more processors, the programs containing instructions for performing an image-to-image relative position prediction method based on inertial measurement data.
[0101] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink, DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0102] In summary, this invention discloses a method for predicting the relative position between images based on inertial measurement data. The method acquires several ultrasound images and corresponding inertial measurement unit (IMU) data for each ultrasound image, wherein the ultrasound images are located in the same image sequence. Based on the IMU data corresponding to each ultrasound image, the method determines the orientation and acceleration information corresponding to each ultrasound image. Based on the ultrasound images, their corresponding orientation and acceleration information, the method determines the relative position information corresponding to each ultrasound image, wherein the relative position information reflects the relative position change between two adjacent ultrasound images. This invention, by combining IMU data and the image sequence to predict the relative position information between images, solves the problem that existing deep learning techniques, which rely solely on ultrasound images to estimate relative positions, are easily affected by inter-image displacement and accumulated drift errors.
[0103] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for predicting the relative position between images based on inertial measurement data, characterized in that, The method includes: Acquire several ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence; Based on the inertial measurement unit data corresponding to each of the ultrasonic images, the direction information and acceleration information corresponding to each ultrasonic image are determined, including: the direction information is Euler angle information, and each inertial measurement unit data includes measured Euler angle information and measured acceleration information; for each ultrasonic image, a rotation matrix is calculated based on the measured Euler angle information corresponding to the ultrasonic image to determine the Euler angle information corresponding to the ultrasonic image, wherein the measured Euler angle information corresponds to the celestial-northeast coordinate system, and the Euler angle information corresponds to the coordinate system of the ultrasonic image; the direction of gravity is determined based on the measured Euler angle information corresponding to the ultrasonic image, and the acceleration information corresponding to the ultrasonic image is determined based on the measured acceleration information and the direction of gravity, wherein the acceleration information is determined based on the difference between the measured acceleration information and the component of the measured acceleration information in the direction of gravity; Based on each ultrasound image, and the direction and acceleration information corresponding to each ultrasound image, the relative position information corresponding to each ultrasound image is determined. The relative position information reflects the relative position change between two adjacent ultrasound images. This includes: inputting each ultrasound image, the Euler angle information, and the acceleration information corresponding to each ultrasound image into a pre-trained target network to obtain the relative position information; the target network includes a feature extraction module, a fusion module, and a prediction network. The feature extraction module inputs each ultrasound image, the acceleration information, and the direction information corresponding to each ultrasound image to obtain image features, acceleration features, and Euler angle features corresponding to each ultrasound image. The image features corresponding to two adjacent ultrasound images reflect the relative distance information between the two ultrasound images. The image features and acceleration features corresponding to each ultrasound image are input into the fusion module to obtain velocity features corresponding to each ultrasound image. The image features and acceleration features corresponding to each ultrasound image are then input into the fusion module to obtain velocity features corresponding to each ultrasound image. The velocity features and the Euler angle features are input into the prediction network to obtain the relative position information. The relative position information includes several relative transformation parameters corresponding to two adjacent frames in each ultrasound image. The several relative transformation parameters include several translation variables and several rotation angles. The several translation variables correspond to different directions, and the several rotation angles correspond to different directions. The feature extraction module includes a residual network, a first fully connected layer, and a second fully connected layer. The input of the residual network is each ultrasound image, and the output of the residual network is the image features corresponding to each ultrasound image. The input of the first fully connected layer is the acceleration information corresponding to each ultrasound image, and the output of the first fully connected layer is the acceleration feature corresponding to each ultrasound image. The input of the second fully connected layer is the Euler angle information corresponding to each ultrasound image, and the output of the second fully connected layer is the Euler angle feature corresponding to each ultrasound image. The fusion module includes a fusion unit and a feature enhancement network. Both the feature enhancement network and the prediction network are constructed using long short-term memory networks.
2. The image-to-image relative position prediction method based on inertial measurement data according to claim 1, characterized in that, Each of the ultrasound images is acquired based on a preset ultrasound imaging device, and the inertial measurement unit data corresponding to each of the ultrasound images is acquired based on a preset inertial measurement unit located on the ultrasound imaging device.
3. The image-to-image relative position prediction method based on inertial measurement data according to claim 1, characterized in that, The fusion module includes a fusion unit and a feature enhancement network. The step of inputting the image features and acceleration features corresponding to each ultrasound image into the fusion module to obtain the velocity features corresponding to each ultrasound image includes: The image features and acceleration features corresponding to each ultrasound image are input into the fusion unit to obtain the fusion velocity features corresponding to each ultrasound image. The fusion velocity features corresponding to each of the ultrasound images are input into the feature enhancement network to obtain the velocity features corresponding to each of the ultrasound images.
4. The image-to-image relative position prediction method based on inertial measurement data according to claim 1, characterized in that, The training process of the target network includes: The training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence are obtained respectively. The training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence are input into an initial network to obtain the predicted relative position information corresponding to the training image. The initial network is the target network that has not been fully trained. Obtain the standard relative position information corresponding to the training image sequence, and determine the first loss value corresponding to the initial network based on the predicted relative position information and the standard relative position information; The network parameters of the initial network are updated according to the first loss value to obtain an updated network. It is then determined whether the updated network has converged to the target value. If not, the steps of obtaining the training image sequence and the orientation information and acceleration information corresponding to each training image frame in the training image sequence are continued until the obtained updated network converges to the target value, thus obtaining the trained target network.
5. The image-to-image relative position prediction method based on inertial measurement data according to claim 1, characterized in that, The method further includes: Based on the relative position information, the predicted acceleration information and predicted Euler angle information corresponding to each of the ultrasound images are determined respectively; Based on the acceleration information and predicted acceleration information corresponding to each of the ultrasound images, the acceleration loss value corresponding to the target network is determined. Based on the Euler angle information and the predicted Euler angle information corresponding to each of the ultrasound images, the Euler angle loss value corresponding to the target network is determined; Based on the acceleration loss value and the Euler angle loss value, determine the second loss value corresponding to the target network; The target network parameters are updated based on the second loss value.
6. A device for predicting the relative position between images based on inertial measurement data, characterized in that, The device includes: The data acquisition module is used to acquire several ultrasound images and inertial measurement unit data corresponding to each ultrasound image, wherein each ultrasound image is located in the same image sequence; The data conversion module is used to determine the direction information and acceleration information corresponding to each of the ultrasound images based on the inertial measurement unit data corresponding to each of the ultrasound images, including: the direction information is Euler angle information, and each inertial measurement unit data includes measured Euler angle information and measured acceleration information; for each ultrasound image, a rotation matrix is calculated based on the measured Euler angle information corresponding to the ultrasound image to determine the Euler angle information corresponding to the ultrasound image, wherein the measured Euler angle information corresponds to the northeast-northeast coordinate system, and the Euler angle information corresponds to the coordinate system of the ultrasound image; the gravity direction is determined based on the measured Euler angle information corresponding to the ultrasound image, and the acceleration information corresponding to the ultrasound image is determined based on the measured acceleration information and the gravity direction, wherein the acceleration information is determined based on the difference between the measured acceleration information and the component of the measured acceleration information in the gravity direction; An information prediction module is used to determine the relative position information of each ultrasound image based on each ultrasound image, the direction information and acceleration information corresponding to each ultrasound image, wherein the relative position information is used to reflect the relative position change between two adjacent ultrasound images. This includes: inputting each ultrasound image, the Euler angle information and acceleration information corresponding to each ultrasound image into a pre-trained target network, and obtaining the relative position information through the target network; the target network includes a feature extraction module, a fusion module, and a prediction network; inputting each ultrasound image, the acceleration information and direction information corresponding to each ultrasound image into the feature extraction module to obtain image features, acceleration features, and Euler angle features corresponding to each ultrasound image, wherein the image features corresponding to two adjacent ultrasound images are used to reflect the relative distance information between the two ultrasound images; inputting the image features and acceleration features corresponding to each ultrasound image into the fusion module to obtain velocity features corresponding to each ultrasound image; and inputting the image features and acceleration features corresponding to each ultrasound image into the fusion module to obtain velocity features corresponding to each ultrasound image; and inputting the image features corresponding to each ultrasound image into the fusion module. The features, velocity features, and Euler angle features are input into the prediction network to obtain the relative position information. The relative position information includes several relative transformation parameters corresponding to two adjacent frames in each ultrasound image. The several relative transformation parameters include several translation variables and several rotation angles. The several translation variables correspond to different directions, and the several rotation angles correspond to different directions. The feature extraction module includes a residual network, a first fully connected layer, and a second fully connected layer. The input of the residual network is each ultrasound image, and the output of the residual network is the image features corresponding to each ultrasound image. The input of the first fully connected layer is the acceleration information corresponding to each ultrasound image, and the output of the first fully connected layer is the acceleration feature corresponding to each ultrasound image. The input of the second fully connected layer is the Euler angle information corresponding to each ultrasound image, and the output of the second fully connected layer is the Euler angle feature corresponding to each ultrasound image. The fusion module includes a fusion unit and a feature enhancement network. Both the feature enhancement network and the prediction network are constructed using long short-term memory networks.
7. A terminal, characterized in that, The terminal includes a memory and one or more processors; the memory stores one or more programs; the programs contain instructions for executing the image-to-image relative position prediction method based on inertial measurement data as described in any one of claims 1-5; the processor is used to execute the programs.
8. A computer-readable storage medium storing a plurality of instructions, characterized in that, The instructions are applicable to being loaded and executed by a processor to implement the steps of the image relative position prediction method based on inertial measurement data as described in any one of claims 1-5.
Citation Information
Patent Citations
Ultrasonic imaging guiding method, ultrasonic equipment and storage medium
CN113116386A