Fetal craniocerebral ultrasonic section unsupervised spatial registration method and device based on comparative learning
Through an unsupervised learning method based on contrast learning, self-supervised spatial sampling and LBP feature selection are designed to achieve unsupervised registration of two-dimensional ultrasound sections in three-dimensional space, solving the problem of time-consuming and labor-consuming registration of fetal cranial brain sections in the prior art, and improving diagnostic efficiency and accuracy.
Patent Information
- Application Number
- CN202510562236.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The existing technology is difficult to effectively combine two-dimensional ultrasound and three-dimensional ultrasound information to achieve accurate registration of fetal cranial brain sections, which makes it time-consuming and difficult for doctors to obtain standard sections. The existing methods require manual annotation and three-dimensional model alignment, and cannot effectively utilize three-dimensional spatial information.
Unsupervised learning method based on contrast learning is adopted, positive sample pairs are constructed through self-supervised spatial sampling and LBP feature selection, and a contrast learning model is designed, combining spatial geometric knowledge and search methods to realize unsupervised registration of two-dimensional sections in three-dimensional space. The search library is constructed using the anatomical similarity of adjacent sections to realize spatial registration of two-dimensional sections at any position.
Unsupervised two-dimensional ultrasonic section to three-dimensional space registration is achieved, reducing manual labeling work, improving diagnostic efficiency, assisting doctors in quickly adjusting the probe position and accurately obtaining standard sections.
Smart Images

Figure CN120471965A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning technology, specifically to the field of machine learning supporting neural network models, and more particularly to a method and device for unsupervised spatial registration of fetal cranial ultrasound sections based on contrast learning. Background Art
[0002] Central nervous system (CNS) malformations are the most common defects among all neonatal (congenital) malformations and are the main cause of perinatal mortality. The International Society of Ultrasound in Obstetrics and Gynecology states that pregnant women should undergo ultrasound examinations during the second trimester to detect whether the CNS is developing normally. Currently, fetal brain screening during pregnancy is mainly performed through three standard diagnostic sections: two-dimensional ultrasound scanning through the thalamus, through the lateral ventricles, and through the cerebellum. Level III screening is performed through these sections. Through these sections, physiological parameters such as head circumference (HC), biparietal diameter (BPD), cerebellar diameter (CD), and posterior cranial fossa (Cisterna magna, CM) width can be qualitatively or quantitatively analyzed. In order to obtain more biological information, such as the developmental status of structures such as the lateral ventricle (LV), cavum septum pellucidi (CSP) and choroid plexus (CP), and to ensure the accuracy of the diagnosis, some doctors will also perform level IV screening by scanning four standard diagnostic sections: the coronal section of the body of the lateral ventricle, the coronal section of the cerebellum, the coronal section of the trigone of the lateral ventricle, and the midsagittal plane.
[0003] However, due to differences in fetal posture and physician experience, manually locating these standard sections using 2D ultrasound is extremely difficult and time-consuming. On the one hand, physicians need to mentally establish the position of the 2D sections within the 3D cranial anatomical space to continuously adjust the probe angle and position. On the other hand, due to the limitations of the 2D probe scanning angle and the influence of the fetal posture, some standard sections, especially those used in Level IV screening, are difficult to scan. These factors increase the difficulty of obtaining standard sections and the physician's workload. In clinical practice, even for physicians with a certain level of work experience, an ultrasound examination can take more than two hours to obtain all seven standard sections for diagnosis. If 2D and 3D ultrasound information could be combined, and by predicting the position of fetal cranial sections in 3D space and achieving registration of 2D ultrasound sections to 3D ultrasound space, this could provide physicians with excellent dynamic observation in clinical practice, assisting them in quickly adjusting the 2D probe to the standard sections, and improving diagnostic efficiency.
[0004] Numerous researchers have investigated the registration of cranial ultrasound sections. Common approaches employ CNN or Transformer architectures to achieve registration by predicting the coordinates of anchor points or spatial parameters such as rotation and translation for 2D sections in 3D space. However, these approaches require manual annotation, manual debugging, or pre-alignment of 3D volume data using methods such as 3D alignment networks. While reinforcement learning can rapidly generate standard sections, it cannot cope with more complex clinical scenarios and requires prior annotation of standard sections. Consequently, due to challenges with manual sample annotation and 3D model alignment, researchers have begun using unsupervised learning to investigate fetal cranial sections. These techniques employ traditional contrastive learning losses or anatomically-aware contrastive learning losses for feature extraction. Compared to the high sample annotation requirements of supervised regression models based on CNN or Transformer models, unsupervised learning based on contrastive learning significantly reduces the manual annotation and data preprocessing required for registration. Furthermore, they excel in fetal cranial ultrasound recognition tasks with low inter-class variability and individual differences. However, the current application of contrastive learning in cranial sections is based on the classification of two-dimensional standard sections. This method cannot establish a connection with three-dimensional space, nor can it effectively utilize the spatial information of three-dimensional ultrasound, and cannot solve the problem of guiding doctors to locate standard sections. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the present invention proposes a method and device for unsupervised spatial registration of fetal cranial ultrasound sections based on contrastive learning, which aims to predict the position of two-dimensional sections in three-dimensional space. On the one hand, utilizing the ability of contrastive learning in classifying and anatomical perception of sections, a contrastive learning model is first designed, the purpose of which is to utilize the anatomical similarity and consistency of adjacent sections in the cranial space, hoping that the features finally extracted from adjacent sections are similar, so as to realize the classification and recognition of sections at different positions. On this basis, the spatial coordinates of the sections are regressed by combining spatial geometry knowledge and retrieval methods to realize an unsupervised ultrasound two-dimensional to three-dimensional registration; on the other hand, the present invention utilizes the anatomical similarity and consistency of adjacent sections in the cranial space to design a new proxy task, so that the contrastive learning model can learn the features of two-dimensional sections at different spatial positions, and then constructs a retrieval library through an automated spatial sampling method, which contains anatomical information and spatial information of different section positions, for realizing spatial registration of two-dimensional sections at arbitrary positions.
[0006] In a first aspect, the present invention proposes an unsupervised spatial registration method for fetal cranial ultrasound sections based on contrastive learning, the method comprising the steps of:
[0007] S1, the self-supervised spatial sampling method samples uniformly distributed two-dimensional slices from the three-dimensional volume data of the training set;
[0008] S2. The pretext task module based on LBP feature selection constructs the adjacent two-dimensional sections on the same normal vector or the adjacent two-dimensional sections on adjacent normal vectors into positive sample pairs;
[0009] S3, inputting all the positive samples of the three-dimensional volume data in the training set into a self-supervised contrastive learning module for training;
[0010] S4. Migrate the contrast learning module to the retrieval module, input the two-dimensional ultrasound image to be registered into the retrieval module, and determine the spatial position of the two-dimensional section to be registered in the three-dimensional volume data.
[0011] Furthermore, specifically, the function of the self-supervised spatial sampling method is to sample a two-dimensional ultrasound section from a three-dimensional volume data, which is specifically obtained by sampling using the following formula:
[0012]
[0013] where φ i and θ i They are the azimuth and elevation of the i-th normal vector in space, and along each normal vector, equally spaced two-dimensional sections can be generated. Where C is the number of channels, H and W are the length and width of the image; the spatial coordinates of the two-dimensional slice corresponding to the three-dimensional volume data are automatically generated according to the polar coordinates of the sampled two-dimensional slice.
[0014] Furthermore, the step S2 specifically includes the following steps:
[0015] S210, constructing a preliminary positive sample pair from adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors;
[0016] S220, calculating the LBP features of each of the two adjacent slices:
[0017]
[0018] Where n is the number of neighborhood pixels, p c is the gray value of the center pixel, p i is the gray value of the ith neighboring pixel of the center pixel, and s(x) is the sign function;
[0019] S230 , selecting the best matching two-dimensional slice according to the correlation and minimum Euclidean distance of the LBP histogram to form a final positive sample pair.
[0020] Specifically, the self-supervised contrastive learning module has two identical feature encoders with shared weights. The training model in step S3 specifically includes:
[0021] The positive sample pair is input into the feature encoder to obtain two representation vectors; and the model is trained by minimizing the similarity between the two representation vectors.
[0022] Specifically, a section retrieval library is provided in the retrieval module, and the two-dimensional ultrasound image to be registered is input into the retrieval module. Determining the spatial position coordinates of the two-dimensional section to be registered specifically includes: inputting the two-dimensional ultrasound image to be registered into the encoder of the retrieval module, calculating the distance similarity between the feature vector extracted by the encoder and the feature vector of the two-dimensional section in the retrieval library, selecting the coordinates corresponding to the two-dimensional section with the closest distance to establish a mapping relationship with the two-dimensional section to be retrieved, and determining the spatial position coordinates of the two-dimensional section to be retrieved.
[0023] In a second aspect, the present invention proposes an unsupervised spatial registration device for fetal cranial ultrasound sections based on contrastive learning, comprising:
[0024] The slicing module uses a self-supervised spatial sampling method to sample uniformly distributed 2D slices from the 3D volume data of the training set;
[0025] A pretext task module constructs positive sample pairs from adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors based on LBP features;
[0026] A contrastive learning module, inputting all the positive samples of the three-dimensional volume data in the training set into the self-supervised contrastive learning module for training;
[0027] The retrieval module migrates the contrast learning module to the retrieval module, the two-dimensional ultrasound image to be registered is input into the retrieval module, and the spatial position of the two-dimensional section to be registered in the three-dimensional volume data is determined.
[0028] In a third aspect, the present invention provides an ultrasound imaging display system, comprising a probe, a display, and the unsupervised spatial registration device for fetal cranial ultrasound sections based on contrast learning as described in claim 6, wherein a retrieval module of the device is electrically connected to the probe, and the display is electrically connected to the retrieval module of the device, and is characterized in that it comprises:
[0029] The probe includes a first acquisition mode and a second acquisition mode, wherein the first acquisition mode is used to acquire three-dimensional volume data of the subject, and the second acquisition mode is used to acquire two-dimensional cross-sectional images of the subject in real time;
[0030] The device is configured to receive the three-dimensional volume data in the first acquisition mode; and to receive the two-dimensional slice image and obtain its spatial position coordinates in the three-dimensional volume data in the second acquisition mode;
[0031] A display comprising a first display area, a second display area, and a third display area, wherein the first display area is used to display a standard model of the three-dimensional volume data, and a display mark is set on the model to distinguish and mark the section position of the standard section;
[0032] The second display area is used to display the two-dimensional section image in real time;
[0033] The third display area is used to display the section image corresponding to the spatial position coordinates of the two-dimensional section image in the three-dimensional volume data model.
[0034] Specifically, in one embodiment, the first display area is provided with a human-computer interaction interface or interface control; the display mark includes reference lines or reference planes of different colors;
[0035] There is a linkage relationship between the display mark and the standard body model, and the standard model and the display mark are translated or rotated by using the human-computer interaction device or the interface control.
[0036] Specifically, in one embodiment, a display mark corresponding to the two-dimensional slice is set at a corresponding position in the model in the first display area according to the spatial position coordinates.
[0037] Specifically, in one embodiment, the second display area further includes a setting module for setting a target section, characterized in that:
[0038] The movement strategy of the probe is determined according to the two-dimensional section image acquired by the probe in the second acquisition mode and the coordinate position of the target section in the three-dimensional model; the movement strategy includes the direction and angle of movement of the probe.
[0039] In a fourth aspect, the present invention proposes an electronic device, characterized in that it includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, including for executing any one of the above-mentioned methods.
[0040] In a fifth aspect, the present invention proposes a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any one of the methods described are implemented.
[0041] In summary, compared with the prior art, the unsupervised spatial registration method, device, and ultrasound imaging display system for fetal cranial ultrasound sections based on contrastive learning of the present invention achieve the following beneficial technical effects:
[0042] (1) The present invention proposes a method and device for unsupervised spatial registration of fetal brain ultrasound sections based on contrastive learning, which aims to predict the position of a two-dimensional section in three-dimensional space. On the one hand, by utilizing the ability of contrastive learning in classifying and anatomically perceiving sections, a contrastive learning model is first designed. The purpose is to utilize the anatomical similarity and consistency of adjacent sections in the brain space, hoping that the features finally extracted from adjacent sections are similar, so as to realize the classification and recognition of sections at different positions. On this basis, the spatial coordinates of the sections are regressed by combining spatial geometry knowledge with retrieval methods, so as to realize an unsupervised ultrasound two-dimensional to three-dimensional registration. On the other hand, the present invention utilizes the anatomical similarity and consistency of adjacent sections in the brain space, designs a new proxy task, so that the contrastive learning model can learn the features of two-dimensional sections at different spatial positions, and then constructs a retrieval library through an automated spatial sampling method, which contains anatomical information and spatial information of different section positions, so as to realize spatial registration of two-dimensional sections at arbitrary positions.
[0043] (2) The ultrasound imaging display system of the present invention can help doctors dynamically know the current ultrasound probe acquisition position corresponds to the exact position in the fetal brain by dynamically displaying the standard model of the three-dimensional body of the fetus and the two-dimensional section image acquired by the ultrasound probe in the current mode, and displaying the position between the two in real time. The system can also set the target section to be acquired through the setting module, determine the movement strategy of the probe based on the two-dimensional section acquired by the current probe and the coordinate position of the target section in the three-dimensional body model, and assist in guiding doctors to adjust the direction and movement angle of the probe, which helps doctors to accurately approach the acquisition of the two-dimensional image of the target section. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 It is a schematic diagram of the method framework provided by an embodiment of the present invention.
[0046] in, Figure 1 (a) represents the self-supervised sampling process, Figure 1(b) represents the contrastive learning encoder training process, Figure 1 (c) represents the retrieval module, Figure 1 (d) represents the LBP feature selection module.
[0047] Figure 2 Schematic diagram of visualization of the predicted plane and actual position provided by an embodiment of the present invention.
[0048] in, Figure 2 (a) represents the real plane, Figure 2 (b) represents the prediction plane, Figure 2 (c) shows the spatial position visualization of two planes. DETAILED DESCRIPTION
[0049] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the various embodiments and features of the following methods, devices, and systems can be combined with each other unless they conflict.
[0050] In view of the shortcomings of existing technologies, limited by a series of factors such as the small scale of public medical image datasets and time-consuming manual annotation, and with the rise of unsupervised learning, this paper will apply the comparative learning framework to the research direction of fetal cranial sections to explore the versatility and effectiveness of unsupervised learning methods in different medical image downstream tasks. The present invention proposes a method and device for unsupervised spatial registration of fetal cranial ultrasound sections based on contrastive learning, aiming to predict the position of two-dimensional sections in three-dimensional space. On the one hand, by utilizing the ability of contrastive learning in section classification and anatomical perception, a contrastive learning model is first designed, the purpose of which is to utilize the anatomical similarity and consistency of adjacent sections in the cranial space, hoping that the features finally extracted from adjacent sections are similar, so as to realize the classification and recognition of sections at different positions. On this basis, the spatial coordinates of the sections are regressed by combining spatial geometry knowledge and retrieval methods to realize an unsupervised ultrasound two-dimensional to three-dimensional registration; on the other hand, the present invention utilizes the anatomical similarity and consistency of adjacent sections in the cranial space to design a new proxy task, so that the contrastive learning model can learn the features of two-dimensional sections at different spatial positions, and then constructs a retrieval library through an automated spatial sampling method, which contains anatomical information and spatial information of different section positions, for realizing spatial registration of two-dimensional sections at arbitrary positions.
[0051] In a specific embodiment, the fetal 3D volume data used in the present invention is 3D volume data of the fetal brain acquired using a 3D ultrasound probe (Ultrasound Department of Western Theater Command General Hospital), stored in DCM format. A total of 39 cases were obtained, 31 of which were used for training, and the remaining data was used to verify the accuracy of the method. All data acquisition was handled with privacy protection and does not involve any infringement or ethical issues. Specifically, all data in the present invention were acquired by seven ultrasound physicians using a GE Voluson E8 ultrasound probe. The data was acquired with an ultrasound image automatic acquisition device (Tele-IAQ, Chengdu Tianfu Jincheng Frontier Medical Equipment Research Institute, China). 39 cases of 3D volume data were obtained through 3D reconstruction in 3D mode, with a size of 128×128×128. Eight of these cases were selected as the test set, and the remaining 31 cases were used as the training set. The present invention selected different amounts of volume data for training based on the experimental settings. During training, 10,000 2D slices were randomly sampled from each training set, with 1000 sampled normal vectors and 10 planes in each normal vector direction. For each test set, 100 two-dimensional slices were randomly sampled during testing to form a query set for indicator evaluation. The number of sampled normal vectors was 25, and the number of planes per normal vector was 4. All the above data were anonymized.
[0052] In the following embodiments of the present invention, the specific exemplary technical solutions are for illustrative purposes only, and the implementation methods of the technical solutions are not limited thereto.
[0053] First embodiment
[0054] In one embodiment, the present invention proposes an unsupervised spatial registration method for fetal brain ultrasound sections based on contrast learning, and the method specifically includes the following steps.
[0055] Step S100: The self-supervised spatial sampling method samples uniformly distributed two-dimensional slices from the three-dimensional volume data of the training set.
[0056] In one embodiment, illustratively, the three-dimensional volume data of the present invention is three-dimensional volume data of a fetal skull acquired using a three-dimensional ultrasound probe. The three-dimensional volume data used here has the universal characteristics of three-dimensional data, and the three-dimensional volume data has vertices, vertex coordinates, spatial coordinates of each volume data point, and is provided with characteristic contents such as sampling normal vectors. In one embodiment, the training set of data is 31 cases of three-dimensional volume data. Before inputting into the comparison model, the three-dimensional volume data is subjected to two-dimensional section sampling through spatial sampling. The number of sampled normal vectors for each volume data is 1000, and the number of sampled planes for each normal vector direction is 10. Therefore, the number of sampled two-dimensional planes for each case of volume data is 10,000. The 31 cases of volume data generate a total of 310,000 sections of different angles and positions and corresponding spatial coordinates, such as Figure 1 As shown in (a), the three-dimensional volume data is sampled to obtain a two-dimensional cross-sectional image.
[0057] Preferably, in one embodiment, the function of the self-supervised spatial sampling method of the present invention is to sample a large number of two-dimensional ultrasound slices from a three-dimensional volume data. The sampling is obtained by the following formula:
[0058]
[0059] where φ i and θ i They are the azimuth and elevation of the i-th normal vector in space, and along each normal vector, equally spaced two-dimensional sections can be generated. Here C is the number of channels, H and W are the length and width of the image. The spatial coordinates of the slice corresponding to the three-dimensional volume data can be automatically generated based on the polar coordinates of the sampled slice. Preferably, in one embodiment, the present invention selects the center point P in the plane i1 , point P in the lower left corner i2 and point P in the lower right corner i3 As the anchor point P for the section plane registration in space i =[P i1 , P i2 , P i3 ].like Figure 1 As shown in (a), the three-dimensional volume data is sampled to obtain two-dimensional cross-sectional images, and the anchor point coordinates of each two-dimensional cross-sectional image in space are obtained.
[0060] In one embodiment, the subscript i depends on the number m of sampled normal vectors set in formula (3). When training, m is set to 1000, and the range of i is 1 to 1000. The above formula will obtain 1000 pairs of different φ i and θ i, i.e., direction angle and elevation angle. Each pair of direction angle and elevation angle with the same subscript determines the direction of a sampling normal vector. The above determines the spatial information of all the slices that need to be sampled. A more uniformly distributed two-dimensional section is obtained. Through this spatial sampling method, the present invention can generate φ by setting the number of normal vectors. i and θ i Theoretically, an infinite number of positive sample pairs can be obtained, which helps the contrastive learning model learn a wider distribution of brain anatomical features, so that the model can obtain information input from a sufficiently wide range of spatial locations.
[0061] It is understood that the method of segmenting to obtain two-dimensional slice images of the present invention can also generate equally spaced two-dimensional slices along each normal vector, and the number of normal vectors and the interval size can be customized according to actual needs. In other embodiments, any other implementation method capable of segmenting three-dimensional volume data to obtain standard two-dimensional slices can also be used.
[0062] Step S200: The pretext task module based on LBP feature selection constructs the adjacent two-dimensional slices on the same normal vector or the adjacent two-dimensional slices on adjacent normal vectors into positive sample pairs.
[0063] In one embodiment, specifically, the spatial information in step S100 is represented, and in one embodiment, optionally, the spatial information includes information such as the normal vector corresponding to each slice during the slicing process, the spatial anchor point coordinates corresponding to each slice, and in another embodiment, only the anchor point coordinate information may be included. The input data of the pretexttask module based on LBP feature selection is a two-dimensional slice and two adjacent two-dimensional slices on the same normal vector or two adjacent two-dimensional slices on adjacent normal vectors. The LBP feature screening output obtains the one of these adjacent slices that is most similar to the two-dimensional slice, forming a positive sample pair.
[0064] In one embodiment, step S200 specifically includes the following steps:
[0065] S210, constructing a preliminary positive sample pair from adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors. S220, calculating the LBP features of each of the two adjacent slices:
[0066]
[0067] Where n is the number of neighborhood pixels, p c is the gray value of the center pixel, p i is the gray value of the ith neighborhood pixel of the center pixel, and s(x) is the sign function.
[0068] S230: Filter and retain the most matching slices based on the LBP features and the minimum Euclidean distance to form the final positive sample pair.
[0069] The positive sample pairs obtained in step S200 are input into the contrastive learning module for training, such as Figure 1 (d) shown.
[0070] In one embodiment, after obtaining the LBP features of each slice, the Euclidean distance between the LBP features of the two slices is calculated, and the two slices with the smallest Euclidean distance are selected to form the final sample pair. In one embodiment, the formula for the Euclidean distance is a conventional formula in the art.
[0071] The existing classic proxy tasks of contrastive learning usually follow the instance discrimination principle, and construct positive sample pairs by rotating, cropping, and other image enhancement methods of a sample (such as an ultrasound section image), and the other samples constitute negative samples. However, this method may ignore the inherent anatomical consistency characteristics in fetal cranial ultrasound, resulting in the adjacent sections belonging to the same anatomical category as the sample being pushed away. Therefore, through the proxy task pretexttask module designed by the present invention, the present invention designs a new pretext task based on the anatomical information of fetal cranial brain images: constructing two adjacent sections on the same spatial normal vector or on adjacent normal vectors as positive sample pairs at the same position. At the same time, taking into account the anatomical variability of cranial brain sections, even sections in the same normal vector direction sometimes have anatomical structural variations, and because LBP performs well in capturing image texture and is invariant to grayscale changes and rotations, we hope to use the LBP feature selection module to screen out samples with large structural variations, so as to improve the anatomical similarity of positive sample pairs and improve the sample quality of representation learning.
[0072] Step S300: All the positive sample pairs of the three-dimensional volume data in the training set are input into a self-supervised contrastive learning module for training.
[0073] In one embodiment, specifically, the contrastive learning module is designed based on Simsiam, and the model has two identical feature encoders f with shared weights. In other embodiments, the contrastive learning module can also be designed with other modules as long as it can achieve its functions.
[0074] Specifically, the contrastive learning module inputs the two slices corresponding to the positive sample pair. Features are extracted separately through two identical encoders. The loss is designed based on calculating the cosine similarity of the two features, and then the model is trained by backpropagation. The purpose of the contrastive learning module is to train an encoder. The trained contrastive learning encoder should classify images with the closest anatomical similarity into the same category. This entire process does not require labeling and is unsupervised. This encoder serves the retrieval module and can bring slices with similar spatial positions and anatomical structures closer together (the extracted features are very similar) and push slices with different spatial positions further apart (the extracted features are very different), thereby minimizing intra-class differences and maximizing inter-class differences.
[0075] In one embodiment, specifically, the training model in step S300 includes:
[0076] Step S310: Input the positive sample pair (x1, x2) into two identical feature encoders f with shared weights to obtain two representation vectors with large receptive fields and small dimensions:
[0077] z1=f(x1), (6)
[0078] z2=f(x2), (7)
[0079] Step S320: After a multi-layer perceptron (MLP) h mapping, p1=h(x1) is obtained, and the model is trained by minimizing the similarity between p1 and z2:
[0080]
[0081] In one embodiment, specifically, on the one hand, by training a self-supervised contrastive learning module, a correspondence between ultrasound fetal section features (including the spatial coordinates of the section and the feature vector corresponding to the section) and the coordinates of the corresponding spatial anchor points is established. On the other hand, for each three-dimensional volume data in the training set, a section retrieval library is established through a spatial sampling method. Each two-dimensional section in the library stores the position information P of the three anchor points of each section. i , and extract features through the encoder trained in the present invention to obtain the feature F of each image i Therefore, each slice in the retrieval library has a corresponding (P i , F i ), the size of the retrieval library is determined by the number of sampled normal vectors and the plane interval on each normal vector.
[0082] Step S400: Migrate the contrast learning module to the retrieval module, input the two-dimensional ultrasound image to be registered into the retrieval module, and determine the spatial position coordinates of the two-dimensional section to be registered in the three-dimensional volume data.
[0083] In one embodiment, in the retrieval module, a doctor uses a handheld 3D probe in 3D ultrasound mode to scan a 3D volume data and construct a slice retrieval library through spatial sampling and feature extraction. Specifically, in one embodiment, after the doctor obtains an offline 3D volume data, the 3D volume data is used to construct a slice retrieval library through spatial sampling and feature extraction. The following steps can be performed: a slice retrieval library is established using the spatial sampling method in step S100. Each 2D slice in the library stores the position information P of its three anchor points. i , and extract features through the encoder of the training contrast learning module in step S300 to obtain the feature F of each image i Similarly, according to the two-dimensional section method mentioned above, it is understandable that each three-dimensional volume data can obtain infinite two-dimensional sections. At the same time, each section in each search library has a corresponding (P i , F i ).
[0084] Specifically, in one embodiment, an encoder trained by the contrastive learning module may be provided in the retrieval module; in another embodiment, the encoder of the contrastive learning module may be migrated to the retrieval module.
[0085] Specifically, the two-dimensional ultrasound image to be registered may be an offline fetal ultrasound section image collected by a doctor. In another embodiment, it may be a two-dimensional section obtained by switching the two-dimensional mode of the ultrasound instrument to locate the fetal skull section.
[0086] In one embodiment, the two-dimensional ultrasound image to be registered is input to the retrieval module, and determining the spatial coordinates of the two-dimensional slice to be registered may include: inputting the two-dimensional ultrasound image to be registered into an encoder of the retrieval module, calculating distance similarity between feature vectors extracted by the encoder and feature vectors of two-dimensional slices in a retrieval library, establishing a mapping relationship between the coordinates corresponding to the closest two-dimensional slice in the retrieval library and the two-dimensional slice to be retrieved, and determining the spatial coordinates of the two-dimensional slice to be retrieved. The output spatial coordinates are the spatial coordinates of the two-dimensional ultrasound image to be registered, thereby achieving a spatial registration effect.
[0087] In another embodiment, the scanned two-dimensional section to be searched is searched by the search module to find the most similar section in the search library, and is visualized in three-dimensional space based on the matching position information provided by the search library, which can help doctors adjust the direction and position of the probe and quickly locate the standard section to be scanned. The clinical scenario is demonstrated as follows Figure 2 shown.
[0088] The present invention proposes an unsupervised spatial registration method for fetal cranial ultrasound sections based on contrastive learning, which aims to predict the position of two-dimensional sections in three-dimensional space. On the one hand, by utilizing the ability of contrastive learning in section classification and anatomical perception, a contrastive learning model is first designed. The purpose is to utilize the anatomical similarity and consistency of adjacent sections in the cranial space, hoping that the features finally extracted from adjacent sections are similar, so as to realize the classification and recognition of sections at different positions. On this basis, the spatial coordinates of the sections are regressed by combining spatial geometry knowledge and retrieval methods to realize an unsupervised ultrasound two-dimensional to three-dimensional registration; on the other hand, the present invention utilizes the anatomical similarity and consistency of adjacent sections in the cranial space to design a new proxy task, so that the contrastive learning model can learn the features of two-dimensional sections at different spatial positions, and then constructs a retrieval library through an automated spatial sampling method, which contains anatomical information and spatial information of different section positions, for realizing spatial registration of two-dimensional sections at arbitrary positions.
[0089] Example 2
[0090] The present invention also provides another embodiment, which proposes an unsupervised spatial registration device for fetal cranial ultrasound sections based on contrast learning, comprising:
[0091] The slicing module uses a self-supervised spatial sampling method to sample uniformly distributed 2D slices from the 3D volume data of the training set;
[0092] A pretext task module constructs positive sample pairs from adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors based on LBP features;
[0093] A contrastive learning module, inputting all the positive samples of the three-dimensional volume data in the training set into the self-supervised contrastive learning module for training;
[0094] The retrieval module migrates the contrast learning module to the retrieval module, the two-dimensional ultrasound image to be registered is input into the retrieval module, and the spatial position of the two-dimensional section to be registered in the three-dimensional volume data is determined.
[0095] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the steps, processing methods, devices, modules, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0096] Third embodiment
[0097] The present invention also provides another embodiment. The present invention proposes an ultrasound imaging display system, which includes a probe, a display, and an unsupervised spatial registration device for fetal cranial ultrasound sections based on contrastive learning as described in Example 2. The retrieval module of the unsupervised spatial registration device for fetal cranial ultrasound sections based on contrastive learning is electrically connected to the probe, and the display is electrically connected to the retrieval module of the device for generating standard fetal cranial sections. The system is characterized in that it includes:
[0098] The probe includes a first acquisition mode and a second acquisition mode, wherein the first acquisition mode is used to acquire three-dimensional volume data of the subject, and the second acquisition mode is used to acquire two-dimensional cross-sectional images of the subject in real time;
[0099] The unsupervised spatial registration device for fetal cranial ultrasound sections based on contrastive learning is configured to receive the three-dimensional volume data in the first acquisition mode; and to receive the two-dimensional section image and obtain its spatial position in the three-dimensional volume data in the second acquisition mode;
[0100] A display comprising a first display area, a second display area, and a third display area, wherein the first display area is used to display a standard model of the three-dimensional volume data, and a display mark is set on the model to distinguish and mark the section position of the standard section;
[0101] The second display area is used to display the two-dimensional section image in real time;
[0102] The third display area is used to display the section image corresponding to the spatial position coordinates of the two-dimensional section image in the three-dimensional volume data model.
[0103] In an embodiment of the present invention, exemplarily, three standard planes are displayed in the standard model of the three-dimensional volume data, corresponding to the three standard planes of the x, y, and z axes of the three-dimensional data, namely the transverse plane, the coronal plane, and the sagittal plane. The first display area is provided with a human-computer interaction interface, and the display mark includes reference lines and reference planes of different colors, or other presentation methods that can distinguish different standard sections. There is a linkage relationship between the display mark and the standard model of the three-dimensional volume data. By using a human-computer interaction device or interface control such as a trackball, a knob, or a menu button to translate or rotate the standard model of the three-dimensional volume data and the display mark (reference line or reference plane), dynamic observation of various angles of the standard model of the three-dimensional volume data can be achieved.
[0104] In an embodiment of the present invention, exemplarily, the third display area and the second display area have a linkage relationship, and based on the two-dimensional section image acquired in real time by the probe in the second mode, its coordinate position in the standard model of the three-dimensional volume data is obtained in real time, and a display mark corresponding to the section is set based on the coordinate position of the three-dimensional coordinate position in the standard model of the three-dimensional volume data in the third display area, that is, the reference line or reference plane corresponding to the real-time two-dimensional section image is dynamically displayed in the standard model of the three-dimensional volume data. By dynamically displaying the position of the two-dimensional section image acquired by the current probe position, it can help doctors determine in real time the accurate position of the currently acquired section position corresponding to the three-dimensional volume data of the fetal brain.
[0105] In an embodiment of the present invention, exemplarily, further, the second display area of the ultrasound imaging display system further includes a setting module for setting a target section, the target section being one or more of the three standard planes of the cross-section that the doctor wants to collect, namely, one or more of the thalamic horizontal cross-section (Transthalamic, TT), the lateral ventricle horizontal cross-section (Transventricular, TV), and the cerebellum horizontal cross-section (Transcerebellar, TC), among the seven or more standard diagnostic sections, including the thalamic horizontal cross-section (Transthalamic, TT), the lateral ventricle horizontal cross-section (Transventricular, TV), the cerebellum horizontal cross-section (Transcerebellar, TC), the midsagittal (Midsagittal, MS) section, the coronal to the body of the lateral ventricles (CLV) section, the coronal cerebellar (Cerebellar coronal, CC) section, the coronal to the trigone of the lateral ventricles (Coronal to the trigone of the lateral ventricles) There are multiple standard sections, including the three standard planes of the cross section, including the three standard planes of the transverse section, which can be obtained by two-dimensional ultrasound, but it is not easy to obtain other section images through two-dimensional ultrasound. In the second mode of the probe, the doctor can set the target section to be collected through the setting module, and determine the movement strategy of the probe according to the coordinate position of the two-dimensional section currently collected by the probe and the target section in the standard model of the three-dimensional volume data. Exemplarily, the movement strategy includes the direction and movement angle of the probe movement, and is dynamically displayed in the standard model of the three-dimensional volume data in the first display area. Through the implementation of the above embodiment, the doctor can be assisted and guided to adjust the direction and movement angle of the probe, which helps the doctor to accurately approach the acquisition of the image of the target section.
[0106] Example 4
[0107] The present invention also provides another embodiment. The present invention proposes an electronic device, including: a processor 1 and a memory 2.
[0108] The memory 2 is used to store computer programs.
[0109] The memory 2 includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk.
[0110] The processor 1 is connected to the memory 2 and is used to execute the computer program stored in the memory 2 so that the unsupervised spatial registration device of fetal cranial ultrasound sections based on contrast learning executes the above-mentioned unsupervised spatial registration method of fetal cranial ultrasound sections based on contrast learning.
[0111] Preferably, the processor 1 may be a central processing unit (CPU) or an application specific integrated circuit (ASIC).
[0112] Example 5
[0113] The present invention also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program, and the computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the unsupervised spatial registration method of fetal cranial ultrasound sections based on contrast learning as described above.
[0114] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0115] Those skilled in the art will appreciate that the modules, units, and / or method steps of the various embodiments described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0116] In the several embodiments provided herein, it should be understood that the disclosed devices, apparatuses, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules or components may be combined or integrated into another device or system, or some features may be omitted or not implemented.
[0117] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0118] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised spatial registration method for fetal cranial ultrasound sections based on contrastive learning, characterized in that: The method comprises the steps of: S1, the self-supervised spatial sampling method samples uniformly distributed two-dimensional slices from the three-dimensional volume data of the training set; S2. The pretext task module based on LBP feature selection constructs the adjacent two-dimensional sections on the same normal vector or the adjacent two-dimensional sections on adjacent normal vectors into positive sample pairs; S3, inputting all the positive samples of the three-dimensional volume data in the training set into a self-supervised contrastive learning module for training; S4. Migrate the contrast learning module to the retrieval module, input the two-dimensional ultrasound image to be registered into the retrieval module, and determine the spatial position coordinates of the two-dimensional section to be registered in the three-dimensional volume data.
2. The unsupervised spatial registration method for fetal cranial ultrasound sections based on contrast learning according to claim 1 is characterized in that: The function of the self-supervised spatial sampling method is to sample a two-dimensional ultrasound section from a three-dimensional volume data, which is specifically obtained by sampling using the following formula: where φ i and θ i They are the azimuth and elevation of the i-th normal vector in space, and along each normal vector, equally spaced two-dimensional sections can be generated. Where C is the number of channels, H and W are the length and width of the image; the spatial coordinates of the two-dimensional slice corresponding to the three-dimensional volume data are automatically generated according to the polar coordinates of the sampled two-dimensional slice.
3. The unsupervised spatial registration method for fetal cranial ultrasound sections based on contrast learning according to claim 1 or 2, characterized in that: The step S2 specifically includes the following steps: S210, constructing a preliminary positive sample pair from adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors; S220, calculating the LBP features of each of the two adjacent slices: Where n is the number of neighborhood pixels, p c is the gray value of the center pixel, p i is the gray value of the i-th neighboring pixel of the center pixel; S230 , selecting the best-matching two-dimensional slice according to the LBP feature and the minimum Euclidean distance to form a final positive sample pair.
4. The unsupervised spatial registration method for fetal cranial ultrasound sections based on contrast learning according to claim 3, characterized in that: The self-supervised contrastive learning module has two identical feature encoders with shared weights. The training model in step S3 specifically includes: The positive sample pair is input into the feature encoder to obtain two representation vectors; and the model is trained by minimizing the similarity between the two representation vectors.
5. The unsupervised spatial registration method for fetal cranial ultrasound sections based on contrast learning according to claim 4, characterized in that: The retrieval module is provided with a section retrieval library. The two-dimensional ultrasound image to be registered is input to the retrieval module, and the spatial position coordinates of the two-dimensional section to be registered are determined specifically including: the two-dimensional ultrasound image to be registered is input into the encoder of the retrieval module, the feature vector extracted by the encoder is calculated with the feature vector of the two-dimensional section in the retrieval library for distance similarity, the coordinates corresponding to the two-dimensional section with the closest distance are selected to establish a mapping relationship with the two-dimensional section to be retrieved, and the spatial position coordinates of the two-dimensional section to be retrieved are determined.
6. A device for unsupervised spatial registration of fetal cranial ultrasound sections based on contrastive learning, comprising: The slicing module uses a self-supervised spatial sampling method to sample uniformly distributed two-dimensional slices from the three-dimensional volume data of the training set. The pretext task module constructs positive sample pairs based on the LBP feature by combining adjacent two-dimensional slices on the same normal vector or adjacent two-dimensional slices on adjacent normal vectors. A contrastive learning module, inputting all the positive samples of the three-dimensional volume data in the training set into the self-supervised contrastive learning module for training; The retrieval module migrates the contrast learning module to the retrieval module, the two-dimensional ultrasound image to be registered is input into the retrieval module, and the spatial position coordinates of the two-dimensional section to be registered in the three-dimensional volume data are determined.
7. An ultrasound imaging display system comprising a probe, a display, and the unsupervised spatial registration device for fetal cranial ultrasound sections based on contrast learning according to claim 6, wherein a retrieval module of the device is electrically connected to the probe, and the display is electrically connected to the retrieval module of the device, characterized in that: include: The probe includes a first acquisition mode and a second acquisition mode, wherein the first acquisition mode is used to acquire three-dimensional volume data of the subject, and the second acquisition mode is used to acquire two-dimensional cross-sectional images of the subject in real time; The device is configured to receive the three-dimensional volume data in the first acquisition mode; and to receive the two-dimensional slice image and obtain its spatial position coordinates in the three-dimensional volume data in the second acquisition mode; A display comprising a first display area, a second display area, and a third display area, wherein the first display area is used to display a standard model of the three-dimensional volume data, and a display mark is set on the model to distinguish and mark the section position of the standard section; The second display area is used to display the two-dimensional section image in real time; The third display area is used to display the section image corresponding to the spatial position coordinates of the two-dimensional section image in the three-dimensional volume data model.
8. The ultrasonic imaging display system according to claim 7, characterized in that: The first display area is provided with a human-computer interaction interface or interface control; the display mark includes reference lines or reference planes of different colors; the display mark and the standard body model are linked, and the standard model and the display mark are translated or rotated using the human-computer interaction device or the interface control.
9. The ultrasonic imaging display system according to claim 7 or 8, characterized in that: According to the spatial position coordinates, a display mark corresponding to the two-dimensional slice is set at a corresponding position in the model in the first display area.
10. The ultrasound imaging display system according to claim 9, wherein the second display area further comprises a setting module for setting a target section, wherein: The movement strategy of the probe is determined according to the two-dimensional section image acquired by the probe in the second acquisition mode and the coordinate position of the target section in the three-dimensional model; the movement strategy includes the direction and angle of movement of the probe.
11. An electronic device, characterized in that: It includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include steps for executing the unsupervised spatial registration method of fetal cranial ultrasound sections based on contrast learning as described in any one of claims 1-5.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the unsupervised spatial registration method of fetal cranial ultrasound sections based on contrast learning as described in any one of claims 1 to 5.