Ultrasonic image prediction method and device based on multi-level spatial-temporal characteristic recurrent neural network
A multi-level spatiotemporal feature recurrent neural network addresses the lack of skilled personnel in ultrasound screening by predicting standard image slices with enhanced accuracy, facilitating broader screening and training in underserved areas.
Patent Information
- Application Number
- CN202510789772.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Traditional artificial intelligence algorithms cannot effectively assist ultrasound scanning screening in underdeveloped areas, resulting in unbalanced medical resources, insufficient coverage of ultrasound scanning screening, and medical staff lack professional knowledge, making it difficult for them to complete scanning and screening of specific human parts.
A multi-level spatiotemporal feature recurrent neural network is adopted, combined with long and short-term memory networks and improved Transformer, and the local and global features of ultrasound images are extracted through a variable window sampling mechanism, and the prediction accuracy is improved using the spatial mean loss function to build an ultrasound image prediction model.
It reduces the knowledge barriers for medical staff to perform human medical ultrasound scans, improves the coverage and training efficiency of ultrasound screening, alleviates the problem of imbalance in medical resources, and provides auxiliary tools for ultrasound imaging and training.
Smart Images

Figure CN120318228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medical image processing and recurrent neural networks, and particularly to an ultrasonic image prediction method and device based on a multi-level spatio-temporal feature recurrent neural network. Background Art
[0002] Medical ultrasound scanning and screening require medical professionals with a great deal of professional knowledge. The uneven distribution of medical resources across the country has led to insufficient coverage of ultrasound scanning and screening in underdeveloped areas. Traditional artificial intelligence algorithms and technologies cannot effectively assist medical staff lacking professional knowledge in completing the scanning and screening of specific human body parts. The automatic prediction technology of diagnostic image sections based on the learning of the change pattern of ultrasonic scan image features can be used as an auxiliary tool for human ultrasound scanning training and examination, reducing the medical knowledge threshold required for reading ultrasound scans and obtaining diagnostic images, increasing the coverage rate of ultrasound screening in rural areas, and thus alleviating the problem of uneven distribution of medical resources. During the scanning process of the human organs by ultrasound doctors, the standard sectional images required for diagnosis are collected by confirming the approximate scanning area, observing the changes in the medical image features, and judging and saving the diagnostic images. Among them, the medical feature changes of ultrasonic images can be learned by an image recurrent network for the image feature changes in continuous time series. The long short-term memory neural network LSTM is usually used to process natural language generation tasks with temporal attributes, learn and understand the semantic changes in the above text, and infer and generate the following text. An image is composed of continuous images in the same time series. After feature extraction of the image by a convolutional neural network, combined with a recurrent network for training, the natural image at a later time in the time series is predicted and generated, which is the earliest neural network model for image prediction.
[0003] In terms of image feature extraction, the convolutional neural network CNN is more sensitive to local features in the image, has higher computational efficiency, and is better at capturing subtle changes in local features; Transformer captures global dependencies based on the attention mechanism, that is, the mutual relationship between local image features and the whole image and other local features, and has stronger modeling ability for long-distance dependencies. In order to better extract the medical features in ultrasonic images, Transformer integrates the hierarchical construction method of CNN, and the downsampling multiple of the feature map changes with the increase in the number of layers. Subsequent research improves the ability of the Transformer module to extract medical image features with a small increase in calculation by using a windowed attention mechanism to reduce the calculation amount of a single Transformer module. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention proposes an ultrasonic image prediction method and device based on a multi-level spatio-temporal feature recurrent neural network. This method researches and develops a data processing method for the prediction task of using standard section diagnostic images in retrospective ultrasonic image data, mainly locates the ultrasonic standard section images required for diagnosis in the image data, so as to obtain the temporal position of the gold standard image in the image; the prediction of the ultrasonic standard section image is completed by combining image patch embedding, multi-level spatio-temporal feature recurrent network, image patch downsampling, image patch upsampling and image reconstruction layer as needed; among them, the multi-level feature recurrent network is based on the long short-term memory network, uses a Transformer with variable window sampling to replace the traditional CNN for extracting spatio-temporal features of medical images, and uses a spatial mean loss function to calculate the temporal mean of several frames of images before and after the target prediction image, and calculates the result with the gold standard to calculate and train the loss function with spatial structural constraints.
[0005] The object of the present invention is achieved by the following technical solutions: an ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network, the method comprising the following steps:
[0006] S1, establish a retrospective human ultrasonic scan image database, and pair the gold standard of the diagnostic image with the corresponding image;
[0007] S2, use an ultrasonic image label positioning algorithm to determine the temporal position of the label in the corresponding image, and at the same time perform noise reduction and cropping processing on the image;
[0008] S3, construct an ultrasonic image prediction model, including: image embedding patch, multi-level spatio-temporal feature recurrent network, image patch up and down sampling layer and image reconstruction layer, and the final output result is a human ultrasonic scan image that meets the doctor's diagnostic criteria; the multi-level spatio-temporal feature recurrent network is composed of a long short-term recurrent network nested with a Transformer including variable window sampling and a spatial mean loss function;
[0009] S4, input the processed image data set into the ultrasonic image prediction model for training, and perform ultrasonic image prediction based on the trained model.
[0010] Further, in step S2, each image file corresponds to a gold standard image for diagnosis, and the positioning algorithm combines the RMSE index and the structural similarity index to calculate the frame position of the gold standard image in the time series of the image to which it belongs.
[0011] Further, in step S3, the image embedding patch, the image patch upsampling layer, and the downsampling layer are respectively based on the image vectorization processing, image patch expansion, and image patch stitching techniques used by the Transformer, and the image reconstruction layer is a two-dimensional deconvolution operation.
[0012] Further, in step S3, the variable window sampling changes the original global self-attention mechanism operation into a calculation within a given window size range, and enables the transfer of image feature information between adjacent windows by moving the window.
[0013] Further, in step S3, the spatial mean loss function, based on the positioning algorithm, selects the image output of the current frame and the image prediction results of several previous frames for mean processing as the prediction result of the current frame, which is used to simulate the characteristics of ultrasound images during the actual scanning process and is a loss function that includes image spatial features and the law of image temporal changes.
[0014] Further, in step S3, in the ultrasound image prediction model described above, it can include multiple multi-level spatio-temporal feature recurrent networks connected in series.
[0015] Further, in step S3, the multi-level spatio-temporal feature recurrent network can include multiple Transformers with variable window sampling for image feature extraction, and the number thereof varies according to specific tasks and computing power resources.
[0016] Further, in step S4, the ultrasound image prediction model uses the Adam optimizer to complete network training, autonomously selects the initial learning rate, and continuously changes the learning rate according to the gradient change during training.
[0017] In a second aspect, the present invention also provides an ultrasound image prediction device based on a multi-level spatio-temporal feature recurrent neural network, including a memory and one or more processors. An executable code is stored in the memory, and when the processor executes the executable code, it implements the above-mentioned ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network.
[0018] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network.
[0019] The beneficial effects of the present invention are as follows: The present invention pairs retrospective medical ultrasound image data through an ultrasound image label positioning algorithm to construct a dynamic image data database; uses a multi-level spatio-temporal feature recurrent network composed of a long short-term memory network and a modified Transformer to predict and generate the images required by doctors during the scanning process. A variable window sampling mechanism is developed during Transformer feature extraction to extract both local and global features of the image with less computational effort; further simulates the thinking of doctors in selecting diagnostic images, and improves the accuracy of the target prediction image through a spatial mean loss function. The invention can reduce the knowledge barrier for medical staff during human medical ultrasound scanning, alleviate the problem of lack of talents in grass-roots ultrasound departments, and enable the inspection and screening of ultrasound images to be more widely applied; at the same time, the results of this method can be used as auxiliary materials for ultrasound scanning guidance training to improve training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is the overall flowchart of the present invention.
[0022] Figure 2 is the architecture diagram of the multi-level spatio-temporal feature recurrent network in the present invention, which includes the main innovative algorithms of variable window sampling and spatial mean loss function.
[0023] Figure 3 is the structural diagram of an ultrasound image prediction device based on a multi-level spatio-temporal feature recurrent neural network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the drawings.
[0025] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0026] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0027] The present invention provides an ultrasound image prediction method based on a multi-level spatiotemporal feature recurrent neural network, such as Figure 1 As shown, the method comprises the following steps:
[0028] S1, establish a retrospective human ultrasound scan image database;
[0029] In this database, the gold standard of diagnostic images and the corresponding images need to maintain a one-to-one correspondence. The paired image data are read frame by frame according to the acquired time series and saved as image files;
[0030] S2, obtaining the temporal position of the gold standard in the image, performing image preprocessing, and establishing an ultrasound image prediction data set;
[0031] The temporal position of the gold standard in the image is obtained, and image preprocessing is performed to establish an ultrasound image prediction data set, including: using an ultrasound image label positioning algorithm to determine the temporal position of the gold standard in the corresponding image, and using the video frame corresponding to the temporal position and its 24 forward video frames as the ultrasound image prediction data for this scan, of which the first 15 frames are input frames and the last 10 frames are prediction frames; using Gaussian filtering to perform denoising on the gold standard image and the ultrasound image prediction data, and then scaling the high-quality image to a resolution of 320×240. The data enhancement method is the same as the general data enhancement method for deep learning; dividing the ultrasound image prediction data into a training set, a validation set, and a test set according to a 6:2:2 data partitioning method to establish an ultrasound image prediction data set.
[0032] The ultrasound image label positioning algorithm is used to calculate the frame number position of the gold standard image in the time series of the image to which it belongs. The positioning algorithm formula is:
[0033]
[0034] in, Represents the similarity index of mixed RMSE and SSIM images, represents a custom hyperparameter between 0 and 1. represents the RMSE indicator, Indicates the structural similarity index. In the specific implementation case is 0.1.
[0035] S3, predicting the standard section images required for doctors’ diagnosis through ultrasound images;
[0036] As shown Figure 2 in the figure, the ultrasonic image prediction model includes: an image embedding patch, a multi-level spatio-temporal feature recurrent network, an image patch upsampling layer, and an image reconstruction layer, which are combined as needed to complete the task.
[0037] The image embedding patch, the image patch upsampling layer, and the downsampling layer are respectively the image vectorization processing, image patch expansion, and image patch splicing techniques used by a general Transformer. The image reconstruction layer uses two-dimensional deconvolution operations. In a specific implementation case, two multi-level spatio-temporal feature recurrent networks are embedded before and after the image patch downsampling layer.
[0038] S4. Construction of a multi-level spatio-temporal feature recurrent network;
[0039] The multi-level spatio-temporal feature recurrent network includes: Transformer image feature extraction with variable window sampling, a long short-term memory neural network, and a spatial mean loss function;
[0040] The multi-level spatio-temporal feature recurrent network is an improvement based on the long short-term memory neural network. It mainly uses a Transformer with variable window sampling to replace the traditional CNN for spatio-temporal feature extraction. By capturing the relationship between local image features and global features, the model has stronger analysis ability in long-distance dependence relationships. In a specific implementation case, each multi-level spatio-temporal feature recurrent network embeds two Transformer modules with variable window sampling.
[0041] The variable window sampling changes the original global self-attention mechanism operation into a calculation within a given window size range, and allows image feature information to be transmitted between adjacent windows by moving the window. In a specific implementation case, the window size is 2.
[0042] Furthermore, based on the ultrasonic image label localization algorithm, the image output of the current frame and the output results of its previous two frames are respectively averaged, and finally the average sum is used as the prediction result of the current frame. The formula is:
[0043]
[0044] Among them, represents the index of the video frame sequence, is the image output of the th frame, represents the mean output result of the th frame.
[0045] The spatial mean loss function is different from calculating the loss by comparing the prediction result of a single image with the gold standard. Instead, it concatenates the prediction results from the 1st frame to the penultimate frame to obtain a prediction vector, and concatenates the ground truth frames from the 2nd frame to the last frame to obtain a ground truth vector. Calculating the loss using the prediction vector and the ground truth vector can incorporate the spatial features and temporal variation patterns of the images. The formula for the loss function is as follows:
[0046]
[0047] In a specific implementation case, represents the concatenated vector of the prediction results from the 1st to 24th frames, represents the concatenated vector of the ground truth frames from the 2nd to 25th frames, that is:
[0048]
[0049]
[0050] S5. Training of the ultrasound image prediction model;
[0051] The ultrasound image prediction model uses the Adam optimizer in torch. This optimizer adopts an optimization algorithm with an adaptive learning rate, and adjusts the learning rate of each parameter by calculating the exponential moving averages of the first moment estimate and the second moment estimate of the gradient during the training process. In a specific implementation case, the learning rate is 0.0001 and the batch size is 2.
[0052] In one embodiment, after both training and validation converge, the network model structure and its hyperparameters are saved. When calling the neural network algorithm, the network model can be read and data can be input, and predictions can be directly made according to the saved hyperparameters.
[0053] Corresponding to the foregoing embodiments of an ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network, the present invention also provides an embodiment of an ultrasound image prediction system based on a multi-level spatio-temporal feature recurrent neural network. The system includes a data pairing module, an image processing module, a model construction module, and an image prediction module; the implementation processes of the modules in the system can refer to the specific implementation steps of the foregoing ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network.
[0054] The data pairing module is used to establish a retrospective human ultrasound scan image database and pair the gold standard of the diagnostic image with the corresponding image;
[0055] The image processing module is used to determine the timing position of the label in the corresponding image using the ultrasonic image label positioning algorithm, and at the same time perform noise reduction and cropping processing on the image; the positioning algorithm combines the RMSE index and the structural similarity index to calculate the frame number position of the gold standard image in the time series of the image.
[0056] The model construction module is used to construct an ultrasonic image prediction model, including: image embedding patches, multi-level spatio-temporal feature recurrent network, image patch upsampling layer and image reconstruction layer. The final output result is a human ultrasonic scan image that meets the doctor's diagnosis standard. Among them, multiple multi-level spatio-temporal feature recurrent networks can be connected in series; the multi-level spatio-temporal feature recurrent network is composed of a long short-term recurrent network nested with a Transformer containing variable window sampling and spatial mean loss function for image feature extraction, and the number of them varies according to specific tasks and computing resources; the image embedding patches, image patch upsampling layer and downsampling layer are respectively based on the image vectorization processing, image patch expansion and image patch splicing techniques used by the Transformer, and the image reconstruction layer is a two-dimensional deconvolution operation. The variable window sampling changes the original global self-attention mechanism operation into a calculation within a given window size range, and allows image feature information to be transmitted in adjacent windows by moving the window. The spatial mean loss function, based on the positioning algorithm, selects the mean of the current frame's image output and the image prediction results of several previous frames as the prediction result of the current frame to simulate the characteristics of ultrasonic images in the real scanning process, and is a loss function that includes image spatial features and image temporal variation rules.
[0057] The image prediction module is used to input the processed image data set into the ultrasonic image prediction model for training, and perform ultrasonic image prediction based on the trained model. The ultrasonic image prediction model uses the Adam optimizer to complete network training, autonomously selects the initial learning rate, and continuously changes the learning rate according to the gradient change during training.
[0058] Corresponding to the foregoing embodiment of an ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network, the present invention also provides an embodiment of an ultrasonic image prediction device based on a multi-level spatio-temporal feature recurrent neural network.
[0059] See Figure 3 , an ultrasonic image prediction device based on a multi-level spatio-temporal feature recurrent neural network provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement an ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network in the above embodiment.
[0060] An embodiment of the ultrasonic image prediction device based on a multi-level spatio-temporal feature recurrent neural network provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the ultrasonic image prediction device based on a multi-level spatio-temporal feature recurrent neural network provided by the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, the device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0061] The implementation processes of the functions and roles of each unit in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.
[0062] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0063] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for ultrasonic image prediction based on a multi-level spatio-temporal feature recurrent neural network in the above embodiment.
[0064] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0065] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the described ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network.
[0066] The foregoing is only a preferred embodiment of one or more embodiments of this specification, and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. An ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network, characterized in that, The method includes the following steps: S1. Establish a retrospective human ultrasound scan image database, and pair the gold standard of diagnostic images with the corresponding images for data; S2. Use an ultrasound image tagging localization algorithm to determine the temporal position of the tag in the corresponding image, and at the same time perform noise reduction and cropping processing on the image; S3. Construct an ultrasound image prediction model, including: image embedding patches, multi-level spatio-temporal feature recurrent networks, image patch upsampling and downsampling layers, and an image reconstruction layer. The final output result is a human ultrasound scan image that meets the doctor's diagnostic criteria; the multi-level spatio-temporal feature recurrent network is composed of a long short-term memory network nested with a Transformer including variable window sampling and a spatial mean loss function; S4. Input the processed image dataset into the ultrasound image prediction model for training, and perform ultrasound image prediction based on the trained model.
2. The ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein In step S2, each image file corresponds to a gold standard image for diagnosis. The localization algorithm combines the RMSE index and the structural similarity index to calculate the frame position of the gold standard image in the time series of the image to which it belongs.
3. The ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein, In step S3, the image embedding patches, the image patch upsampling layer, and the downsampling layer are respectively based on the image vectorization processing, image patch expansion, and image patch splicing techniques used by the Transformer. The image reconstruction layer is a two-dimensional deconvolution operation.
4. The ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein In step S3, the variable window sampling changes the original global self-attention mechanism operation into a calculation within a given window size range, and enables the image feature information to be transmitted between adjacent windows by moving the window.
5. A method for predicting ultrasonic images based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein In step S3, the spatial mean loss function, based on the localization algorithm, selects to perform mean processing on the image output of the current frame and the image prediction results of several previous frames as the prediction result of the current frame, and is used to simulate the characteristics of ultrasound images in the real scanning process. It is a loss function that includes image spatial features and the law of image temporal changes.
6. The ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein In step S3, in the ultrasound image prediction model, multiple multi-level spatio-temporal feature recurrent networks connected in series can be included.
7. A method for predicting ultrasonic images based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein, In step S3, the multi-level spatio-temporal feature recurrent network can include multiple Transformers with variable window sampling for image feature extraction, and the number thereof varies according to specific tasks and computing power resources.
8. The ultrasonic image prediction method based on a multi-level spatio-temporal feature recurrent neural network according to claim 1, wherein In step S4, the ultrasound image prediction model uses the Adam optimizer to complete network training, autonomously selects the initial learning rate, and continuously changes the learning rate according to the gradient change during training.
9. An ultrasonic image prediction device based on a multi-level spatio-temporal feature recurrent neural network, comprising a memory and one or more processors, wherein executable code is stored in the memory, and is characterized in that, When the processor executes the executable code, it implements an ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network as described in any one of claims 1-8.
10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements an ultrasound image prediction method based on a multi-level spatio-temporal feature recurrent neural network as described in any one of claims 1-8.
Citation Information
Patent Citations
Medical image report generation method and system based on convolution and circulation network
CN115690038A
COVID-19 focus prediction system based on smart contract and self-attention
CN117058088A
Thyroid ultrasound contrast nodule benign and malignant classification method based on 3D ConvFormer
CN117237739A
Ultrasonic video identification method and system based on distracted attention
CN118447292A
Systems and methods for image reconstruction
US20210166446A1