Method of processing an ultrasound image, ultrasound device and storage medium
By analyzing ultrasound image sequences using a trained target detection model and combining static and dynamic features, lesion areas are automatically identified. This solves the problems of inconsistent ultrasound image diagnostic results and low efficiency in existing technologies, achieving higher accuracy and faster lesion detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONOSCAPE MEDICAL CORP
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-24
AI Technical Summary
Current ultrasound examination techniques rely on manual visual interpretation of ultrasound images, which makes it difficult to guarantee the consistency and accuracy of diagnostic results and fails to meet the needs of rapid clinical diagnosis, especially in scenarios where the location and morphology of lesions are dynamically changing.
The trained target detection model is used to analyze ultrasound image sequences. Through multiple feature extraction and feature fusion modules, combined with the static visual features and dynamic change information of the images, the target area, especially the lesion area, is automatically identified.
It improves the accuracy and consistency of ultrasound image analysis, enabling continuous tracking and precise detection in scenarios where lesions are dynamically changing, reducing missed detections and decreased positioning accuracy, and improving diagnostic efficiency.
Smart Images

Figure CN122004924B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology, and more specifically to a method for processing ultrasound images, an ultrasound device, a storage medium, and a computer program product. Background Technology
[0002] Ultrasound examination, as a non-invasive diagnostic technique, has been widely applied in various clinical medical diagnostic scenarios due to its advantages of being non-invasive, real-time, and repeatable. However, current technologies primarily rely on manual visual inspection or experience-based judgment to interpret ultrasound images of the target object and determine the examination results. Because this method is highly dependent on the user's experience and expertise, differences in subjective judgment among different users can lead to inconsistent and inaccurate results, thus affecting the reliability of diagnosis and the effectiveness of clinical decisions. Furthermore, manual analysis is inefficient when dealing with large amounts of ultrasound image data, failing to meet the needs of rapid clinical diagnosis. Therefore, there is an urgent need for an objective, efficient, and repeatable automated analysis method to overcome the shortcomings of existing technologies. Summary of the Invention
[0003] The present invention was proposed in view of the above-mentioned problems. The present invention provides a method for processing ultrasound images, an ultrasound device, a storage medium, and a computer program product.
[0004] According to one aspect of the present invention, a method for processing ultrasound images is provided. The method includes: acquiring a target ultrasound image sequence of a target object, wherein the target ultrasound image sequence includes target ultrasound images of the target object acquired sequentially; inputting the target ultrasound image sequence into a trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence, wherein the trained target detection model is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence, wherein, for each training ultrasound image sequence, the training dynamic change degree sequence corresponding to the training ultrasound image sequence is used to represent the degree of dynamic change of the labeled images corresponding to each of the multiple training ultrasound images in the training ultrasound image sequence, the target region is used to represent the image region in the target ultrasound image where a tissue of interest exists, and the training region is used to represent the image region in the training ultrasound image where a tissue of interest exists.
[0005] For example, the trained target detection model includes: multiple first feature extraction modules, a first feature fusion module connected to all of the multiple first feature extraction modules, a target dynamic change degree determination module connected to the first feature fusion module, a second feature fusion module connected to both the first feature fusion module and the target dynamic change degree determination module, and a region detection module connected to the second feature fusion module. The model inputs a target ultrasound image sequence into the trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence, including:
[0006] For each of the multiple first feature extraction modules, the target ultrasound image input to the first feature extraction module is processed by the first feature extraction module to obtain the target ultrasound image features corresponding to the target ultrasound image;
[0007] The first feature fusion module performs feature splicing on the target ultrasound image features corresponding to each target ultrasound image to obtain the first fused feature corresponding to the target ultrasound image sequence.
[0008] The target dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the first fusion feature. For each target ultrasound image, the target dynamic change degree corresponding to the target ultrasound image is used to represent the image dynamic change degree corresponding to the target ultrasound image.
[0009] The second feature fusion module determines the second fusion feature corresponding to the target ultrasound image sequence based on the target dynamic change degree and the first fusion feature corresponding to each target ultrasound image.
[0010] Using the region detection module, based on the second fusion feature, at least one target region in the last target ultrasound image in the target ultrasound image sequence is identified and displayed.
[0011] For example, the target dynamic change degree determination module includes a floating point displacement feature extraction module, a cyst septum deformation feature extraction module, a lesion deformation feature extraction module, a third feature fusion module connected to the floating point displacement feature extraction module, the cyst septum deformation feature extraction module, and the lesion deformation feature extraction module, and a first dynamic change degree determination module connected to the third feature fusion module. Based on the first fused features, it determines the target dynamic change degree corresponding to each target ultrasound image, including:
[0012] The floating point displacement feature extraction module extracts features from the first fused features to obtain floating point displacement features, which are used to represent the image features of moving floating points in the target ultrasound image sequence.
[0013] The cyst septum deformation feature extraction module extracts features from the first fused features to obtain cyst septum deformation features. The cyst septum deformation features are used to represent the image features of deformed cyst septum in the target ultrasound image sequence. The receptive field of the cyst septum deformation feature extraction module is larger than that of the floating point displacement feature extraction module.
[0014] The lesion deformation feature extraction module extracts features from the first fusion feature to obtain lesion deformation features. The lesion deformation features are used to represent the image features of deformed lesions in the target ultrasound image sequence. The receptive field of the lesion deformation feature extraction module is larger than that of the cyst septum deformation feature extraction module.
[0015] The third feature fusion module is used to fuse the displacement features of floating point objects, the deformation features of cyst septa, and the deformation features of lesions to obtain the third fused feature.
[0016] The first dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the third fusion feature.
[0017] For example, determining and displaying at least one target region in the last target ultrasound image in a target ultrasound image sequence based on a second fusion feature includes:
[0018] The second fusion feature is compressed and fused in the channel dimension, and based on the compressed and fused second fusion feature, at least one target region in the last target ultrasound image in the target ultrasound image sequence is determined and displayed.
[0019] For example, the processing method further includes:
[0020] For each training ultrasound image sequence, the predicted dynamic change degree sequence corresponding to the training ultrasound image sequence is determined through the region detection module to be trained. Based on the predicted dynamic change degree sequence and the training dynamic change degree sequence corresponding to the training ultrasound image sequence, the dynamic change degree loss is determined. The predicted dynamic change degree sequence corresponding to the training ultrasound image sequence is used to represent the predicted image dynamic change degree corresponding to each training ultrasound image in the training ultrasound image sequence. The dynamic change degree loss is part of the overall loss used to adjust the model parameters of the target detection model to be trained.
[0021] For example, the trained target detection model is also trained based on the ultrasound gain level corresponding to each training ultrasound image sequence. For each training ultrasound image sequence, the ultrasound gain level corresponding to the training ultrasound image sequence is used to represent the full-image gain level of the training ultrasound image sequence annotation.
[0022] For example, the trained region detection module includes multiple sequentially connected trained second feature extraction modules, multiple trained third feature extraction modules, a trained ultrasound gain feature fusion module connected to all of the multiple trained third feature extraction modules, multiple trained fourth feature fusion modules, and multiple trained first detection head modules. For each of the multiple trained third feature extraction modules, the trained third feature extraction module is connected to a trained second feature extraction module, and this trained second feature extraction module is different from the trained second feature extraction modules connected to the other trained third feature extraction modules. Similarly, for each of the multiple trained fourth feature fusion modules... The trained fourth feature fusion module is connected to both the trained ultrasound gain feature fusion module and a trained second feature extraction module. The trained second feature extraction module connected to the trained fourth feature fusion module is different from the trained second feature extraction modules connected to the other trained fourth feature fusion modules. For each of the multiple trained first detection head modules, the trained first detection head module is connected to a trained fourth feature fusion module, and the trained fourth feature fusion module connected to the trained first detection head module is different from the trained fourth feature fusion module connected to the other trained first detection head modules.
[0023] For example, before training is complete, the trained region detection module further includes a second detection head module to be trained, which is connected to the ultrasound gain feature fusion module to be trained.
[0024] The processing methods also include:
[0025] The gain prediction level of the training ultrasound image sequence is determined by the second detection head module to be trained, based on the gain training features output by the ultrasound gain feature fusion module to be trained. The gain training features are obtained by sequentially passing through the second feature extraction module, the third feature extraction module, and the ultrasound gain feature fusion module to be trained, based on the second fusion features corresponding to the training ultrasound image sequence. The gain prediction level is used to represent the full-image gain degree predicted by the training ultrasound image sequence.
[0026] The processing methods also include:
[0027] Based on the gain prediction level corresponding to the training ultrasound image sequence and the ultrasound gain level corresponding to the training ultrasound image sequence, the gain level loss is determined, wherein the gain level loss is part of the overall loss used to adjust the model parameters of the target detection model to be trained.
[0028] For example, the trained target detection model is also trained based on the region category corresponding to each training region, where the region category is used to represent the disease category or organization name of the organization of interest in that training region.
[0029] According to one aspect of the present invention, an ultrasound device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program to implement the above-described ultrasound image processing method.
[0030] According to one aspect of the present invention, a storage medium is provided storing computer program instructions, which, when executed, are used to perform the above-described ultrasound image processing method.
[0031] According to one aspect of the present invention, a computer program product is provided, comprising computer program instructions which, when executed by a processor, are used to perform the above-described ultrasound image processing method.
[0032] According to the above-described scheme of the present invention, a target ultrasound image sequence of a target object can be obtained. Then, the target ultrasound image sequence is input into a trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence. The trained target detection model is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence. The trained target detection model in the above scheme is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence. The training dynamic change degree sequence can be used to represent the degree of dynamic change of the labeled images corresponding to each of the multiple training ultrasound images. Therefore, the trained target detection model can combine the dynamic information of the target ultrasound image changing over time and the static visual features of the target ultrasound image to determine the target region with higher accuracy, especially in scenarios where some diseased tissues exhibit dynamic changes.
[0033] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0034] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0035] Figure 1 A schematic flowchart of a method for processing ultrasound images according to an embodiment of the present invention is shown;
[0036] Figure 2 A schematic diagram of a target detection model according to an embodiment of the present invention is shown;
[0037] Figure 3 A schematic diagram of a target dynamic change degree determination module according to an embodiment of the present invention is shown;
[0038] Figure 4 A schematic diagram of a region detection module according to an embodiment of the present invention is shown;
[0039] Figure 5 A schematic diagram of a region detection module according to an embodiment of the present invention is shown;
[0040] Figure 6 A schematic block diagram of an ultrasonic device according to an embodiment of the present invention is shown. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a subset of embodiments of the present invention.
[0042] Among related technologies, ultrasound examination technology, with its advantages of being non-invasive and capable of real-time imaging, has been widely applied in medical diagnostic scenarios for various clinical diseases. For example, in the screening and diagnosis of ovarian adnexal lesions, ultrasound images can serve as core imaging evidence, typically used to assess the benign or malignant nature of ovarian adnexal lesions, the degree of lesion development, and the level of malignancy risk. However, since this assessment process usually relies on manual visual interpretation by the user, the interpretation results are easily influenced by differences in the user's clinical experience, subjective judgment bias, and the complexity of the lesion's pathological characteristics. This leads to a lack of consistency and reliability in diagnostic results, ultimately limiting the accuracy and efficiency of lesion assessment. Furthermore, although detection models exist that can locate lesion areas in ultrasound images, these traditional detection models only analyze and process static single-frame ultrasound images. They are not applicable to real-time dynamic changes in lesion location and morphology, making it difficult to continuously track and accurately detect dynamically changing lesions. This results in defects such as missed detection of dynamic lesions, decreased detection and localization accuracy, and insufficient robustness of detection results.
[0043] To at least partially solve the above problems, embodiments of the present invention provide a method for processing ultrasound images. Figure 1 A schematic flowchart illustrating a method for processing ultrasound images according to an embodiment of the present invention is shown. (In conjunction with...) Figure 1 As shown, the ultrasound image processing method includes steps S110 and S120.
[0044] In step S110, the target ultrasound image sequence of the target object is obtained.
[0045] The target object can be any object that requires internal observation by the user using ultrasound equipment. It is understood that the aforementioned ultrasound equipment can be various types of ultrasound imaging devices, such as ultrasound diagnostic instruments or ultrasound imaging workstations. Specifically, ultrasound diagnostic instruments can have different probe types to examine different areas, such as blood vessels, abdomen, gynecology, and breasts. Ultrasound imaging workstations can be devices integrating functional modules such as patient registration, image acquisition, diagnostic editing, report printing, image post-processing, medical record retrieval, and statistical analysis. Ultrasound imaging workstations can communicate with ultrasound diagnostic instruments, for example, through any wired or wireless communication method. After the ultrasound diagnostic instrument completes beamforming and grayscale mapping processing of the ultrasound echo signal, it transmits the ultrasound image data to the ultrasound imaging workstation for subsequent processing, analysis, and storage.
[0046] The target ultrasound image sequence comprises sequentially acquired ultrasound images of the target object. For example, continuous acquisition of ultrasound images of the area to be examined on the target object yields a temporally continuous ultrasound image sequence. A base ultrasound image sequence is generated from this temporally continuous sequence by selecting a preset number of frames. Alternatively, medical record data of the target object can be retrieved to extract a temporally continuous ultrasound image sequence corresponding to the area to be examined. A base ultrasound image sequence is then generated from this temporally continuous sequence by selecting a preset number of frames. In one example, the base ultrasound image sequence can be directly used as the target ultrasound image sequence. In another example, each frame of the base ultrasound image sequence can undergo standardized preprocessing to unify each frame to a preset size. The preprocessed base ultrasound image sequence is then used as the target ultrasound image sequence. It is understood that the preset size can be 640×640×C, where C is the number of ultrasound image channels. For example, when the ultrasound image is a three-channel color image, C is 3, and the preset size is 640×640×3. When the ultrasound image is a single-channel grayscale image, C is set to 1, and the preset size is 640×640×1. This preset size can be adjusted adaptively according to actual needs. The preset number of frames can be determined based on the actual scenario; for example, a preset frame count of 8 is possible. It is understood that the above standardized preprocessing can also achieve denoising, grayscale normalization, and other processing on each frame of the basic ultrasound image sequence.
[0047] In step S120, the target ultrasound image sequence is input into the trained target detection model to obtain and display at least one target region in the last target ultrasound image in the target ultrasound image sequence.
[0048] A target region is used to represent the image region in a target ultrasound image where a tissue of interest exists. For example, the tissue of interest can be at least one of a target lesion, a target anatomical structure, or a target organ. In this case, the target region can be the bounding box annotated with at least one of the target lesion, target anatomical structure, or target organ in the last target ultrasound image in the target ultrasound image sequence. The aforementioned tissue of interest can be determined according to the training objective of the target detection model. Specifically, for example, if the training objective of the target detection model is to identify the lesion region located in the ovary in a target ultrasound image, then the tissue of interest is the tissue in the ovary containing the lesion (i.e., the aforementioned target lesion), and the target region is the region corresponding to the tissue in the ovary containing the lesion in the last target ultrasound image in the target ultrasound image sequence. As another example, if the training objective of the target detection model is to identify the region corresponding to the heart and the lesion region in the heart in a target ultrasound image, then the tissue of interest is the heart (i.e., the aforementioned target organ) and the tissue in the heart containing the lesion (i.e., the aforementioned target lesion), and the target region is the region corresponding to the heart in the last target ultrasound image in the target ultrasound image sequence and the region corresponding to the tissue in the heart containing the lesion.
[0049] The trained target detection model is trained based on multiple training ultrasound image sequences, the training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence.
[0050] Each training ultrasound image sequence may include multiple training ultrasound images. For example, training ultrasound images may be ultrasound images used to train an object detection model. A training ultrasound image sequence may include a predetermined number of training ultrasound images acquired sequentially in chronological order, and all training ultrasound images in the same sequence may be ultrasound images of the same object at the same location. Specifically, for example, all training ultrasound images in a training ultrasound image sequence may be ultrasound images of the object's heart. It is understood that the size of each training ultrasound image in each training ultrasound image sequence is the same as the size of the target ultrasound image in the target ultrasound image sequence.
[0051] For example, for each training ultrasound image sequence, the training dynamic change degree sequence corresponding to the training ultrasound image sequence is used to represent the degree of dynamic change of the labeled images corresponding to each of the multiple training ultrasound images in the training ultrasound image sequence. For example, the training dynamic change degree sequence corresponding to each training ultrasound image sequence can include multiple training dynamic change degrees, and the multiple training dynamic change degrees correspond one-to-one with the multiple training ultrasound images in the training ultrasound image sequence corresponding to the training dynamic change degree sequence. The training dynamic change degree corresponding to each frame of training ultrasound image can be used to represent the degree of image change between it and the previous frame of ultrasound image (i.e., the ultrasound image acquired at the time before the acquisition time of the current training ultrasound image). Among them, the magnitude of the training dynamic change degree is positively correlated with the difference between the current training ultrasound image and the previous frame of ultrasound image. The larger the value of the training dynamic change degree, the more obvious the difference in tissue morphology between the tissue region in the current training ultrasound image and the tissue region in the previous frame of ultrasound image. Correspondingly, the smaller the value of the training dynamic change degree, the more stable the tissue morphology between the tissue region in the current training ultrasound image and the tissue region in the previous frame of ultrasound image.
[0052] For each training ultrasound image sequence corresponding to a training dynamic change degree sequence, except for the earliest acquired training ultrasound image in the sequence (hereinafter referred to as the first training ultrasound image), the training dynamic change degree corresponding to the remaining training ultrasound images can be manually determined based on the image differences between the training ultrasound image and the previous frame. It is understood that when manually annotating the training dynamic change degree sequences corresponding to multiple training ultrasound image sequences, for each training ultrasound image in each training ultrasound image sequence, except for the first training ultrasound image in the sequence, if, compared to its previous frame, the training ultrasound image indicates significant changes within the tissue (e.g., significant movement of floating objects within the tissue, appearance of atrial septa in multilocular cysts, appearance of irregular morphology in hydrosalpinx, clear presentation of various features of teratomas, clear presentation of the outer membrane of the luteum, etc.), then a larger training dynamic change degree (e.g., 0.9, 1.1, etc.) is assigned to that training ultrasound image. Conversely, if the training ultrasound image shows almost no change within the tissue compared to its previous frame, then a small training dynamic change level (e.g., 0.7, 0.8, etc.) is assigned to this training ultrasound image. It is understood that for the first training ultrasound image in each training ultrasound image sequence, the training dynamic change level corresponding to this first training ultrasound image can be a preset training dynamic change level (the specific value of the preset training dynamic change level can be determined according to the actual situation). The training dynamic change level corresponding to this first training ultrasound image can also be determined by obtaining the previous frame of the first training ultrasound image and based on the image dynamic change level between the previous frame and the first training ultrasound image. It is understood that the value range of the above-mentioned training dynamic change level can be determined according to the actual situation; here, a value range of 0.7, 0.8, 0.9, and 1.1 is used as an example.
[0053] For example, for each training ultrasound image sequence, the training region is used to represent the image region in the last training ultrasound image of the sequence where the tissue of interest exists. For instance, if the tissue of interest is tissue with a lesion in the ovary, the training region corresponding to the last training ultrasound image in the training ultrasound image sequence could be the region in the training ultrasound image corresponding to the lesion in the ovary. As another example, if the tissue of interest is the heart, the training region corresponding to the last training ultrasound image in the training ultrasound image sequence could be the region in that last training ultrasound image corresponding to the heart. More specifically, if the tissue of interest is the heart and tissue with a lesion in the heart, the training region corresponding to the last training ultrasound image in the training ultrasound image sequence could be the region in the last training ultrasound image corresponding to the heart and the region corresponding to the tissue with a lesion in the heart.
[0054] This invention provides a training method for an object detection model for reference. A first training dataset is obtained. The object detection model is iteratively trained using the first training dataset, and the model parameters are adjusted using a loss function until training is complete. The first training dataset includes multiple training ultrasound image sequences. The labels for each training ultrasound image sequence may include: the training dynamic change degree sequence corresponding to the training ultrasound image sequence, and at least one training region in the last training ultrasound image in the training ultrasound image sequence. The conditions for training completion may include the loss value calculated by the loss function stabilizing, the number of iterations reaching a preset number, etc.
[0055] This invention provides an example for reference. Taking the use of a trained target detection model to determine the lesion region located in the ovary in the last target ultrasound image of a target ultrasound image sequence as an example, an ultrasound image sequence of the ovary region of the target object can be acquired through the probe of an ultrasound device, serving as the aforementioned target ultrasound image sequence. Then, this target ultrasound image sequence is input into the trained target detection model. The trained target detection model can output and display, on a display device, the last frame of the target ultrasound image sequence, and at least one ovarian lesion region (i.e., the target region) within that frame.
[0056] According to the above-described scheme of the present invention, a target ultrasound image sequence of a target object can be obtained. Then, the target ultrasound image sequence is input into a trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence. The trained target detection model is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence. In the above scheme, the trained target detection model is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence. The training dynamic change degree sequence can be used to represent the degree of dynamic change of the labeled images corresponding to each of the multiple training ultrasound images. Therefore, the trained target detection model can combine the dynamic information of the target ultrasound image changing over time and the static visual features of the target ultrasound image to determine the target region with higher accuracy, especially in scenarios where some diseased tissues exhibit dynamic changes.
[0057] For example, the trained target detection model includes: multiple first feature extraction modules, a first feature fusion module connected to all of the multiple first feature extraction modules, a target dynamic change degree determination module connected to the first feature fusion module, a second feature fusion module connected to both the first feature fusion module and the target dynamic change degree determination module, and a region detection module connected to the second feature fusion module.
[0058] Figure 2 A schematic diagram of a target detection model according to an embodiment of the present invention is shown. Figure 2 The first feature extraction module M1 to the first feature extraction module M n These are the aforementioned multiple first feature extraction modules.
[0059] like Figure 2 As shown, the first feature extraction module M1 to the first feature extraction module M n The output of the first feature fusion module is connected to the input of the target dynamic change degree determination module and the input of the second feature fusion module. The output of the target dynamic change degree determination module is connected to the input of the second feature fusion module. The output of the second feature fusion module is connected to the input of the region detection module.
[0060] Step S120 above, which inputs the target ultrasound image sequence into the trained target detection model to obtain and display at least one target region in the last target ultrasound image in the target ultrasound image sequence, includes steps S121 to S125.
[0061] Step S121: For each of the multiple first feature extraction modules, feature extraction processing is performed on the target ultrasound image input to the first feature extraction module to obtain the target ultrasound image features corresponding to the target ultrasound image.
[0062] This invention provides an example for reference, in which different first feature extraction modules can be used to extract features from different target ultrasound images in a target ultrasound image sequence to obtain the target ultrasound image features of each frame of the target ultrasound image. For example... Figure 2 As shown, with a preset frame count of 3 (i.e., the target ultrasound image sequence contains 3 frames of target ultrasound images) and the size of each frame of target ultrasound image being 640×640×1, the first feature extraction module M1 to the first feature extraction module M... nTaking an example where n equals 3 (i.e., the trained target detection model includes three modules: M1, M2, and M3), we will use the extraction of target ultrasound image features from the first frame of the target ultrasound image sequence (i.e., the earliest acquired target ultrasound image in the target ultrasound image sequence) by the first feature extraction module M1 as an example: First, the first feature extraction module M1 uses 64 3×3 convolution kernels to perform convolution operations on the input first frame of the target ultrasound image with a stride of 1 and a padding value of 1, resulting in 64-dimensional image features. Then, batch normalization (BN) is performed on these 64-dimensional image features to obtain the normalized result. Finally, a nonlinear transformation is performed on the normalized result using the SiLU (Sigmoid-Weighted Linear Unit) activation function, and the nonlinear transformation result is used as the target ultrasound image feature corresponding to the first frame of the target ultrasound image (the size of this target ultrasound image feature is 640×640×64). It is understood that the first feature extraction modules M2 and M3 can employ the same principle and processing flow to output the target ultrasound image features corresponding to the second frame and the third frame of the target ultrasound image sequence, respectively. It is also understood that the size, stride, and padding of the convolution kernel used in each first feature extraction module can be determined based on actual conditions. Furthermore, it is understood that the number of first feature extraction modules does not need to match the number of target ultrasound images in the target ultrasound image sequence; a suitable number can be selected based on actual conditions.
[0063] Step S122: The first feature fusion module performs feature splicing processing on the target ultrasound image features corresponding to each target ultrasound image to obtain the first fused feature corresponding to the target ultrasound image sequence.
[0064] This invention provides an example for reference. Referring again to the example in step S121 above, such as... Figure 2As shown, the first feature fusion module can receive the target ultrasound image features of the first frame of the target ultrasound image output by the first feature extraction module M1, the target ultrasound image features of the second frame of the target ultrasound image output by the first feature extraction module M2, and the target ultrasound image features of the third frame of the target ultrasound image output by the first feature extraction module M3. The first feature fusion module can perform feature stitching processing on the target ultrasound image features of the first frame, the second frame, and the third frame along the channel dimension and according to the order of image acquisition time to obtain the stitched image features. The stitched image features are used as the first fusion feature corresponding to the target ultrasound image sequence. The size of the first fusion feature is 640×640×(64×preset number of frames). In the above example, the preset number of frames is 3, so the size of the first fusion feature is 640×640×(64×3), that is, 640×640×192.
[0065] Step S123: The target dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the first fusion feature.
[0066] For each target ultrasound image, the target dynamic change level is used to represent the degree of dynamic change in that target ultrasound image. For example, the target dynamic change level can be used to predict the image difference between the target ultrasound image and its previous frame. Understandably, the larger the value of the target dynamic change level, the more significant the changes within the tissue represented by that target ultrasound image compared to its previous frame (e.g., significant movement of floating objects within the tissue, appearance of septa in multilocular cysts, irregular morphology of hydrosalpinx, clear presentation of various features of teratomas, clear presentation of the outer membrane of the luteum, etc.). In other words, a larger value of the target dynamic change level indicates a more significant difference in tissue morphology between the tissue region in that frame and the tissue region in the previous frame. A smaller value of the target dynamic change level indicates a smaller change within the tissue represented by that target ultrasound image compared to its previous frame. In other words, the smaller the value of the degree of dynamic change of the target, the more stable the tissue morphology of the tissue region in the current frame of the ultrasound image is compared with that of the tissue region in the previous frame of the ultrasound image.
[0067] This invention provides an example for reference. Referring again to the example in step S122 above, such as... Figure 2As shown, the target dynamic change degree determination module can receive the first fused feature output by the first feature fusion module. Then, the target dynamic change degree determination module can perform multi-scale feature extraction on the first fused feature to capture inter-frame dynamic change information at different levels, thereby obtaining multiple feature extraction results at different scales. These multiple feature extraction results at different scales are then subjected to feature fusion processing (e.g., feature concatenation, weighted feature addition, etc.) to obtain fused image features. Finally, based on these fused image features, the target dynamic change degree corresponding to each target ultrasound image is determined. Specifically, for example, the target dynamic change degree determination module uses a structure with three sets of convolutional layers and pooling layers at different scales to process the first fused feature. For the first scale: 16 3×3 convolutional kernels are used to perform convolution operations on the first fused feature with a stride of 1 and a padding value of 1, obtaining the first convolution result. Then, max pooling is performed on the first convolution result to obtain the first scale image features. For the second scale: 32 5×5 convolutional kernels are used to perform convolution operations on the first fused features with a stride of 1 and a padding value of 2, resulting in the second convolution result. Max pooling is then performed on the second convolution result to obtain the image features for the second scale. For the third scale: 64 7×7 convolutional kernels are used to perform convolution operations on the first fused features with a stride of 1 and a padding value of 3, resulting in the third convolution result. Max pooling is then performed on the third convolution result to obtain the image features for the third scale. It is understandable that the number, size, stride, and padding value of the convolutional kernels at each scale can be adaptively adjusted according to the actual scene to capture features of changes in different sizes of tissues. The target dynamic change degree determination module can assign weights to image features at different scales; these weights can be optimized through iterative training of the target detection model. The number of channels from the first-scale image features to the third-scale image features can be unified through operations such as channel-dimensional feature concatenation or convolution processing. Next, based on the weights corresponding to the image features at the first to third scales, the image features at the first to third scales after unifying the number of channels are weighted and summed to obtain the fused image features. Dimensionality reduction processing is then performed on the fused image features to obtain the dimensionality-reduced feature vector. Finally, two fully connected layers map the dimensionality-reduced feature vector into a numerical sequence corresponding to the number of frames in the target ultrasound image sequence. The values in this numerical sequence correspond one-to-one with each frame of the target ultrasound image sequence according to the image acquisition time sequence, and each value represents the degree of dynamic change of the target in its corresponding target ultrasound image.
[0068] Step S124: Through the second feature fusion module, based on the target dynamic change degree and the first fusion feature corresponding to each target ultrasound image, the second fusion feature corresponding to the target ultrasound image sequence is determined.
[0069] This invention provides an example for reference. The second feature fusion module can use the degree of dynamic change of the target corresponding to each frame of the target ultrasound image as a weight to perform overall weighted enhancement on the channel segments in the first fusion feature corresponding to each frame of the target ultrasound image. For specific examples, continue to refer to the example in step S123 above, such as... Figure 2 As shown, the second feature fusion module can receive the first fusion feature output by the first feature fusion module, and the target dynamic change degree output by the target dynamic change degree determination module for a preset number of frames (the preset number of frames is the number of target ultrasound images in the target ultrasound image sequence; in this example, the preset number of frames equals 3) of target dynamic change degree. The size of the first fusion feature is 640×640×(64×preset number of frames), i.e., 640×640×(64×3), or 640×640×192. The second feature fusion module can use the target dynamic change degree corresponding to each target ultrasound image as the weight of the corresponding channel segment in the first fusion feature for that frame of target ultrasound image. Then, based on this weight, a channel-by-channel weighting operation is performed on the corresponding channel segment of the target ultrasound image in the first fusion feature. Finally, after weighting, the second feature fusion module can use the obtained overall feature as the second fusion feature corresponding to the target ultrasound image sequence. The size of the second fusion feature is 640×640×(64×preset number of frames), i.e., when the preset number of frames equals 3, the size of the second fusion feature is 640×640×192.
[0070] Step S125: Using the region detection module, based on the second fusion feature, determine and display at least one target region in the last target ultrasound image in the target ultrasound image sequence.
[0071] This invention provides an example for reference. Continuing to refer to the example in step S124 above, such as... Figure 2 As shown, the region detection module can receive the second fused feature output by the second feature fusion module. Based on the input second fused feature, the region detection module can determine at least one candidate region in the last target ultrasound image of the target ultrasound image sequence, and the confidence level of each candidate region (the confidence level represents the probability that the candidate region is a region of interest, and its value ranges from greater than or equal to 0 to less than or equal to 1). Then, all candidate regions can be aggregated, filtered by a preset confidence threshold, and subjected to non-maximum suppression to determine at least one target region in the last target ultrasound image of the target ultrasound image sequence. Finally, the last target ultrasound image of the target ultrasound image sequence marked with at least one target region can be displayed on a display device. It is understood that the target region can be marked using a preset style (e.g., using a colored rectangular detection box to mark the target region, with a different color than the background ultrasound image).
[0072] According to the above-described scheme of the present invention, for each of the plurality of first feature extraction modules, feature extraction processing is performed on the target ultrasound image input to the first feature extraction module to obtain the target ultrasound image features corresponding to the target ultrasound image. Then, through the first feature fusion module, feature stitching processing is performed on the target ultrasound image features corresponding to each target ultrasound image to obtain the first fused features corresponding to the target ultrasound image sequence. Then, through the target dynamic change degree determination module, the target dynamic change degree corresponding to each target ultrasound image is determined based on the first fused features. Then, through the second feature fusion module, the second fused features corresponding to the target ultrasound image sequence are determined based on the target dynamic change degree corresponding to each target ultrasound image and the first fused features. Finally, through the region detection module, at least one target region in the last target ultrasound image in the target ultrasound image sequence is determined and displayed based on the second fused features. The above scheme can extract features from the target ultrasound image through multiple first feature extraction modules so that the obtained target ultrasound image features can be used to represent static visual features. Furthermore, the first feature fusion module in the above scheme can perform feature stitching processing on the target ultrasound image so that the obtained first fused feature can contain the temporal correlation information of the target ultrasound image, which can be used to represent the dynamic information that changes over time. Therefore, the target detection model trained in the above scheme can combine the dynamic information that changes over time in the target ultrasound image and the static visual features of the target ultrasound image to determine the target region with higher accuracy, especially in scenarios where some lesion tissues exhibit dynamic changes.
[0073] For example, the target dynamic change degree determination module includes a floating point displacement feature extraction module, a cyst septum deformation feature extraction module, a lesion deformation feature extraction module, a third feature fusion module connected to the floating point displacement feature extraction module, the cyst septum deformation feature extraction module, and the lesion deformation feature extraction module, and a first dynamic change degree determination module connected to the third feature fusion module.
[0074] Figure 3 A schematic diagram of a target dynamic change degree determination module according to an embodiment of the present invention is shown. Figure 3 As shown, the outputs of the floating point displacement feature extraction module, the cyst septum deformation feature extraction module, and the lesion deformation feature extraction module are all connected to the input of the third feature fusion module. The output of the third feature fusion module is connected to the input of the first dynamic change degree determination module.
[0075] In step S123 above, the degree of dynamic change of each target ultrasound image is determined based on the first fusion feature, including steps S1231 to S1235.
[0076] In step S1231, the first fused feature is processed by the floating point displacement feature extraction module to obtain the floating point displacement feature.
[0077] Floating dot displacement features are used to represent the image characteristics of moving floating dot objects in a target ultrasound image sequence. For example, these floating dot objects can be free, fine structures existing within a fluid-filled anechoic area, not fixedly adhered to the cyst wall or surrounding tissue, and capable of displacement under fluid influence. Specifically, floating dot objects may include: desquamated epithelium and protein deposits within simple ovarian cysts, coagulation microparticles within hemorrhagic cysts and corpus luteum hematomas, free lipid droplets and hair debris within ovarian teratomas, and inflammatory cell debris in hydrosalpinx, etc.
[0078] This invention provides an example for reference, such as Figure 3 As shown, the floating point displacement feature extraction module can use a 3×3 convolution kernel to perform feature extraction processing on the input first fused feature, obtain the feature extraction result, and use this feature extraction result as the floating point displacement feature. It is understandable that the size of the convolution kernel in the floating point displacement feature extraction module can be determined according to the actual situation.
[0079] In step S1232, the cyst septum deformation feature extraction module is used to extract features from the first fused features to obtain the cyst septum deformation features.
[0080] Cyst septum deformation features are used to represent the image characteristics of deformed cyst septa in a target ultrasound image sequence. For example, cyst septum deformation can be the dynamic morphological changes of fibrous septa (i.e., multilocular cystic septa) that separate different cystic cavities within a multilocular cystic lesion of the ovary, as presented in continuous time-series ultrasound images. Specifically, cyst septum deformation may include: changes in the curvature of the septum, overall swinging and displacement, dynamic increases and decreases in thickness, local irregular twisting, and tightening or loosening deformation caused by changes in tension. This type of deformation stems from the fact that the septum is not rigidly adhered and can dynamically change with cystic fluid pressure and fluid movement.
[0081] The receptive field of the cyst septum deformation feature extraction module is larger than that of the floating point displacement feature extraction module. For example, Figure 3As shown, the cyst septum deformation feature extraction module can perform feature extraction processing on the input first fused feature using a 3×3 dilated convolution kernel with a dilation rate of 2, obtaining the feature extraction result, which is then used as the cyst septum deformation feature. It is understood that the size and dilation rate of the dilated convolution kernel in the cyst septum deformation feature extraction module can be determined according to the actual situation, ensuring that the receptive field of the cyst septum deformation feature extraction module is larger than that of the floating point displacement feature extraction module.
[0082] In step S1233, the lesion deformation feature extraction module extracts features from the first fusion feature to obtain the lesion deformation feature.
[0083] Lesion deformation features are used to represent the image characteristics of deformed lesions in a target ultrasound image sequence. For example, lesion deformation can be the non-fixed, dynamically changing morphological features of the lesion body in a continuous time-series ultrasound image. Specific examples include overall lesion displacement, dynamic changes in contour shape, and changes in boundary smoothness and clarity.
[0084] The receptive field of the lesion deformation feature extraction module is larger than that of the cyst septum deformation feature extraction module. For example, Figure 3 As shown, the lesion deformation feature extraction module can perform feature extraction processing on the input first fusion feature using a 3×3 dilated convolution kernel with a dilation rate of 5, and obtain the feature extraction result, which is then used as the lesion deformation feature. It is understood that the size and dilation rate of the dilated convolution kernel in the lesion deformation feature extraction module can be determined according to the actual situation, ensuring that the receptive field of the lesion deformation feature extraction module is larger than that of the cyst septum deformation feature extraction module.
[0085] In step S1234, the displacement features of floating point objects, deformation features of cyst septa, and deformation features of lesions are fused by the third feature fusion module to obtain the third fused feature.
[0086] This invention provides an example for reference, such as Figure 3 As shown, the third feature fusion module can receive the floating point displacement features output by the floating point displacement feature extraction module, the cyst septum deformation features output by the cyst septum deformation feature extraction module, and the lesion deformation features output by the lesion deformation feature extraction module. Then, the received floating point displacement features, cyst septum deformation features, and lesion deformation features are subjected to feature fusion processing (e.g., feature stitching, weighted feature addition, etc.) to obtain the fused image features. This fused image feature is used as the third fused feature and output to the first dynamic change degree determination module.
[0087] In step S1235, the first dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the third fusion feature.
[0088] This invention provides an example for reference, such as Figure 3 As shown, the first dynamic change degree determination module can receive the third fused feature output by the third feature fusion module. Then, dimensionality reduction processing is performed on the third fused feature to obtain a dimensionality-reduced feature vector. Finally, the dimensionality-reduced feature vector is mapped to a numerical sequence corresponding to the number of frames in the target ultrasound image sequence through two fully connected layers. The values in this numerical sequence correspond one-to-one with each frame of the target ultrasound image sequence according to the image acquisition time sequence, and each value represents the target dynamic change degree of its corresponding target ultrasound image.
[0089] According to the above-described scheme of the present invention, the first fused feature is extracted using a floating point displacement feature extraction module to obtain the floating point displacement feature. Then, the first fused feature is extracted using a cyst septum deformation feature extraction module to obtain the cyst septum deformation feature. Next, the first fused feature is extracted using a lesion deformation feature extraction module to obtain the lesion deformation feature. Then, the floating point displacement feature, cyst septum deformation feature, and lesion deformation feature are fused using a third feature fusion module to obtain the third fused feature. Finally, the target dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the third fused feature. The floating point displacement feature extraction module in the above scheme can effectively extract the image features of moving floating points in the target ultrasound image sequence; the cyst septum deformation feature extraction module can effectively extract the image features of deformed cyst septum in the target ultrasound image sequence; and the lesion deformation feature extraction module can be used to extract the image features of deformed lesions in the target ultrasound image sequence. Therefore, the above scheme can effectively extract image features at multiple scales to improve the multi-scale representation ability of the obtained third fusion feature, improve the accuracy of the obtained target dynamic change degree, and help improve the accuracy of the final obtained target region.
[0090] For example, in step S125 above, determining and displaying at least one target region in the last target ultrasound image in the target ultrasound image sequence based on the second fusion feature includes: performing compression fusion processing on the second fusion feature in the channel dimension, and determining and displaying at least one target region in the last target ultrasound image in the target ultrasound image sequence based on the second fusion feature after compression fusion processing.
[0091] This invention provides an example for reference. The second fusion feature can be compressed and fused along the channel dimension using a 1×1 convolution kernel to obtain the compressed and fused second fusion feature. Based on the compressed and fused feature, at least one target region in the last frame of the target ultrasound image sequence is detected. Specifically, based on the compressed and fused second fusion feature, at least one target region in the last target ultrasound image of the target ultrasound image sequence is determined and displayed (the specific process can be found in step S125 above, which describes determining and displaying at least one target region in the last target ultrasound image of the target ultrasound image sequence based on the second fusion feature; this will not be elaborated upon here).
[0092] According to the above-described scheme of the present invention, the second fusion feature can be compressed and fused in the channel dimension, and based on the compressed and fused second fusion feature, at least one target region in the last target ultrasound image in the target ultrasound image sequence can be determined and displayed. The above scheme can compress and fused the second fusion feature in the channel dimension, effectively reducing the amount of data in the second fusion feature, improving the efficiency of target region determination, and meeting the requirements of real-time ultrasound detection scenarios.
[0093] For example, the ultrasound image processing method further includes: for each training ultrasound image sequence, determining the predicted dynamic change degree sequence corresponding to the training ultrasound image sequence through the target dynamic change degree determination module to be trained, and determining the dynamic change degree loss based on the predicted dynamic change degree sequence and the training dynamic change degree sequence corresponding to the training ultrasound image sequence.
[0094] In any embodiment of the present invention, any module trained may belong to the trained object detection model, and any module to be trained may belong to the target detection model to be trained. For any module, the module parameters of the module to be trained may be the same as or different from the module parameters of the trained module. Specifically, if the module has learnable module parameters, then the module parameters of the module to be trained are different from the module parameters of the trained module. If the module has fixed module parameters (e.g., parameters used only for feature concatenation), then the module parameters of the module to be trained are the same as the module parameters of the trained module.
[0095] The predicted dynamic change sequence corresponding to the training ultrasound image sequence is used to represent the predicted degree of dynamic change of each training ultrasound image in the sequence. For example, the predicted dynamic change sequence corresponding to the training ultrasound image sequence can be output by an object detection model. This predicted dynamic change sequence includes multiple predicted dynamic change levels, and different predicted dynamic change levels can correspond to different training ultrasound images. The predicted dynamic change level can be used to represent the prediction result of the object detection model on the degree of dynamic change of the corresponding training ultrasound image.
[0096] This invention provides an example for reference. The dynamic change degree loss can be determined based on the difference between the predicted dynamic change degree sequence corresponding to each training ultrasound image sequence and the training dynamic change degree sequence corresponding to that training ultrasound image sequence. Specifically, for example, a multi-class cross-entropy loss function can be used as the loss function for determining the dynamic change degree loss. This loss function measures the difference between the predicted confidence distribution of all values of the dynamic change degree output by the target dynamic change degree determination module, normalized by a normalized exponential function (such as softmax), and the one-hot encoding distribution of the pre-labeled training dynamic change degree sequence of the training ultrasound image sequence. The calculated loss value is used as the aforementioned dynamic change degree loss. It is understood that the aforementioned all values of the dynamic change degree represent the range of values for both the predicted and training dynamic change degrees.
[0097] The dynamic change level loss is part of the overall loss used to adjust the model parameters of the target detection model to be trained. For example, the classification loss can be determined based on the category label of the training region corresponding to the last training ultrasound image in the training ultrasound image sequence (the category label is a manually pre-labeled label in the training ultrasound image sequence, used to clearly indicate the true attribute category of the training region in the last training ultrasound image) and the predicted category result output by the region detection module (the predicted category result is the probability distribution or final classification of the region belonging to each category after the target detection model performs forward propagation on the training region or detected candidate region in the last training ultrasound image of the input training ultrasound image sequence). The bounding box regression loss is then determined based on the coordinate label of the training region corresponding to the last training ultrasound image and the predicted region coordinates output by the target detection model. The classification loss and the bounding box regression loss are then weighted and summed to obtain the main loss of the target detection model (i.e., the loss value corresponding to step S120 above). Finally, the main loss of the target detection model and the dynamic change level loss are weighted and summed according to a preset weighting coefficient to obtain the overall loss used to adjust the model parameters of the target detection model.
[0098] This invention provides a training process for an object detection model: First, a first-stage training dataset is acquired. This first-stage training dataset may include multiple training ultrasound images, each labeled with: the training dynamic change degree corresponding to the training ultrasound image, and at least one training region in the training ultrasound image. The object detection model is trained using this first-stage training dataset until the number of iterations reaches a first preset number. Then, the object detection model is trained again using the first training dataset from step S120 until the number of iterations reaches a second preset number. Finally, the object detection model is trained again using the first training dataset from step S120 until the number of iterations reaches a third preset number or the overall loss of the object detection model tends to stabilize. It is understood that when the number of iterations is less than the first preset number, multiple first feature extraction modules, first feature fusion modules, target dynamic change degree determination modules, and second feature fusion modules in the object detection model can be frozen, and the module parameters of the region detection module are adjusted only according to the overall loss of the object detection model with a first learning rate. If the number of iterations is greater than or equal to a first preset number and less than or equal to a second preset number, the region detection module in the target detection model can be frozen, and the module parameters of multiple first feature extraction modules, first feature fusion modules, target dynamic change degree determination modules, and second feature fusion modules in the target detection model can be adjusted only according to the overall loss of the target detection model using a second learning rate. If the number of iterations is greater than the second preset number, the module parameters of multiple first feature extraction modules, first feature fusion modules, target dynamic change degree determination modules, second feature fusion modules, and region detection modules in the target detection model can be adjusted according to the overall loss of the target detection model using a third learning rate. It is understood that the first preset number is less than the second preset number, the second preset number is less than the third preset number, the first learning rate is greater than the second learning rate, and the second learning rate is greater than the third learning rate.
[0099] According to the above-described scheme of the present invention, for each training ultrasound image sequence, a target dynamic change degree determination module to be trained can determine the predicted dynamic change degree sequence corresponding to the training ultrasound image sequence, and based on the predicted dynamic change degree sequence and the training dynamic change degree sequence corresponding to the training ultrasound image sequence, a dynamic change degree loss can be determined. Based on the dynamic change degree loss, the above scheme can adjust the parameters of the target detection model to be trained, thereby enabling the trained target detection model to combine the dynamic information of the target ultrasound image changing over time and the static visual features of the target ultrasound image to determine the target region with higher accuracy, especially in scenarios where some lesion tissue exhibits dynamic changes.
[0100] For example, the trained target detection model is also trained based on the ultrasound gain level corresponding to each training ultrasound image sequence. For each training ultrasound image sequence, the ultrasound gain level corresponding to the training ultrasound image sequence is used to represent the full-image gain level of the training ultrasound image sequence annotation.
[0101] For example, the ultrasound gain level corresponding to the training ultrasound image sequence can be used to represent the full-image gain level of the training ultrasound image sequence. The full-image gain level of the training ultrasound image sequence can be determined based on the full-image gain of each training ultrasound image in the training ultrasound image sequence, and the full-image gain of each training ultrasound image in the same training ultrasound image sequence is the same. Taking the gain level as including three levels, namely low level, medium level and high level, as an example, the full-image gain value range represented by the full-image gain level corresponding to different gain levels is different. If the full-image gain represented by the full-image gain level corresponding to the low level is greater than or equal to 0 dB and less than or equal to a dB, the full-image gain represented by the full-image gain level corresponding to the medium level is greater than a dB and less than or equal to b dB, and the full-image gain represented by the full-image gain level corresponding to the high level is greater than b dB, the ultrasound gain level can be labeled as low level, medium level or high level according to the interval in which the actual full-image gain of the training ultrasound image sequence falls. It is understandable that the specific values of a decibel and b decibel mentioned above can be determined by the user based on the actual situation.
[0102] This invention provides a training method for an object detection model for reference. A second training dataset is obtained. The object detection model is iteratively trained using the second training dataset, and the model parameters are adjusted using a loss function until training is complete. The second training dataset includes multiple training ultrasound image sequences. The labels for each training ultrasound image sequence may include: the training dynamic change degree sequence corresponding to the training ultrasound image sequence, at least one training region in the last training ultrasound image in the training ultrasound image sequence, and the ultrasound gain level corresponding to the training ultrasound image sequence. The conditions for training completion may include the loss value calculated by the loss function stabilizing, the number of iterations reaching a preset number, etc.
[0103] This invention provides an example for reference. Taking the use of a trained target detection model to determine the lesion region located in the ovary in the last target ultrasound image of a target ultrasound image sequence as an example, an ultrasound image sequence of the ovary region of the target object can be acquired through the probe of an ultrasound device, serving as the aforementioned target ultrasound image sequence. Then, this target ultrasound image sequence is input into the trained target detection model. The trained target detection model can output and display, on a display device, the last frame of the target ultrasound image sequence, and at least one ovarian lesion region (i.e., the target region) within that frame.
[0104] According to the above-described scheme of the present invention, a trained target detection model can be obtained by training based on the ultrasound gain level corresponding to each training ultrasound image sequence. The trained target detection model not only considers information such as tissue structure morphology in the target ultrasound image, but also adaptively corrects for differences in ultrasound image performance caused by different ultrasound gain levels. Therefore, the above scheme helps reduce the impact of image quality fluctuations in the target ultrasound image on the output of the trained target detection model, and can combine the dynamic information of the target ultrasound image changing over time with the static visual features of the target ultrasound image to determine target regions with higher accuracy and robustness, especially in scenarios where some diseased tissues exhibit dynamic changes.
[0105] For example, the trained region detection module includes multiple trained second feature extraction modules, multiple trained third feature extraction modules, a trained ultrasound gain feature fusion module connected to the multiple trained third feature extraction modules, multiple trained fourth feature fusion modules, and multiple trained first detection head modules.
[0106] Figure 4 A schematic diagram of a region detection module according to an embodiment of the present invention is shown. Figure 4 Taking the following example: Second feature extraction modules A1 to A5 are multiple sequentially connected trained second feature extraction modules; third feature extraction modules B1 to B3 are multiple trained third feature extraction modules; fourth feature fusion modules C1 to C3 are multiple trained fourth feature fusion modules; and first detection head modules D1 to D3 are multiple trained first detection head modules. The output of second feature extraction module A1 is connected to the input of second feature extraction module A2; the output of second feature extraction module A2 is connected to the input of second feature extraction module A3; the output of second feature extraction module A3 is connected to the input of second feature extraction module A4; and the output of second feature extraction module A4 is connected to the input of second feature extraction module A5.
[0107] For each of the multiple trained third feature extraction modules, the trained third feature extraction module is connected to a trained second feature extraction module, and this trained second feature extraction module is different from the trained second feature extraction modules connected to the other trained third feature extraction modules. For example, such as Figure 4 As shown, the third feature extraction module B1 is connected to the second feature extraction module A3, the third feature extraction module B2 is connected to the second feature extraction module A4, and the third feature extraction module B3 is connected to the second feature extraction module A5. Different trained third feature extraction modules are connected to different trained second feature extraction modules.
[0108] For each of the multiple trained fourth feature fusion modules, the trained fourth feature fusion module is connected to both a trained ultrasound gain feature fusion module and a trained second feature extraction module. Furthermore, the trained second feature extraction module connected to the trained fourth feature fusion module is different from the trained second feature extraction modules connected to the other trained fourth feature fusion modules. For example, ... Figure 4 As shown, the fourth feature fusion module C1 is connected to the ultrasound gain feature fusion module and the second feature extraction module A3. The fourth feature fusion module C2 is connected to the ultrasound gain feature fusion module and the second feature extraction module A4. The fourth feature fusion module C3 is connected to the ultrasound gain feature fusion module and the second feature extraction module A5. Different trained fourth feature fusion modules are connected to different trained second feature extraction modules.
[0109] For each of the multiple trained first detection head modules, the trained first detection head module is connected to a trained fourth feature fusion module, and the trained fourth feature fusion module connected to the trained first detection head module is different from the trained fourth feature fusion modules connected to the other trained first detection head modules. For example, as... Figure 4 As shown, the first detection head module D1 is connected to the fourth feature fusion module C1. The first detection head module D2 is connected to the fourth feature fusion module C2. The first detection head module D3 is connected to the fourth feature fusion module C3. It can be understood that each trained first detection head module can be used to output the detection results of all candidate regions in the last target ultrasound image of the target ultrasound image sequence at its corresponding scale (this scale can be determined according to the receptive field scale corresponding to the image features input to the first detection head module). The detection results include: the position coordinates corresponding to each candidate region, and the confidence score corresponding to each candidate region. This confidence score can be used to characterize the probability that the candidate region is the target region. Figure 4As shown, when the receptive field scales corresponding to the image features input from different first detection head modules are different, the sizes of the candidate regions output by different trained first detection head modules will also differ. It can also be understood that when multiple first detection head modules determine multiple candidate regions and their corresponding confidence levels, the object detection model can output candidate regions with confidence levels greater than a preset confidence threshold as target regions. The aforementioned preset confidence threshold can be set according to actual circumstances.
[0110] For example, the input of each trained second feature extraction module can be connected to the output of the second feature fusion module. Multi-scale, layer-by-layer feature extraction of the input region detection module's features can be performed through the aforementioned second feature extraction modules A1 to A5. The receptive field scales corresponding to the different second feature extraction modules are different, arranged in ascending order of receptive field scale as follows: second feature extraction module A1, second feature extraction module A2, second feature extraction module A3, second feature extraction module A4, and second feature extraction module A5. Specifically, second feature extraction module A1 receives the second fused feature output from the second feature fusion module, performs preliminary feature extraction based on its own receptive field, generates and outputs fourth-scale image features to second feature extraction module A2. After receiving the fourth-scale image features, second feature extraction module A2 performs further feature extraction on the fourth-scale image features to obtain fifth-scale image features, and transmits these fifth-scale image features to second feature extraction module A3. Based on this principle, the sixth-scale image features output by the second feature extraction module A3, the seventh-scale image features output by the second feature extraction module A4, and the eighth-scale image features output by the second feature extraction module A5 are obtained sequentially.
[0111] like Figure 4 As shown, for each second feature extraction module whose output is connected to the input of the third feature extraction module, the output of the second feature extraction module can also be connected to a fourth feature fusion module, so that the image features output by the second feature extraction module are input to the fourth feature fusion module connected to it. Specifically, for example... Figure 4 If the output of the second feature extraction module A3 is connected to the input of the third feature extraction module B1, then the output of the second feature extraction module A3 is also connected to the input of the fourth feature fusion module C1. For example... Figure 4 As shown, the input terminals of different third feature extraction modules can be connected to the output terminals of different second feature extraction modules, and the output terminals of different third feature extraction modules are all connected to... Figure 4 The input terminal of the ultrasonic gain feature fusion module is connected. For example... Figure 4As shown, the third feature extraction module B1 can extract features from the image features output by the second feature extraction module A3 (i.e., the sixth-scale image features mentioned above) and output the feature extraction results to the ultrasound gain feature fusion module. The third feature extraction module B2 can extract features from the image features output by the second feature extraction module A4 (i.e., the seventh-scale image features mentioned above) and output the feature extraction results to the ultrasound gain feature fusion module. The third feature extraction module B3 can extract features from the image features output by the second feature extraction module A5 (i.e., the eighth-scale image features mentioned above) and output the feature extraction results to the ultrasound gain feature fusion module. It can be understood that different third feature extraction modules can adapt to receptive fields of different scales and can extract multi-scale features at different levels. The later the third feature extraction module is, the larger its corresponding receptive field scale, and the more the extracted features are biased towards global semantic information. Since there are differences in the receptive field scale and image feature size of the image features output by each third feature extraction module, the ultrasound gain feature fusion module can first upsample and align the image features output by each third feature extraction module to ensure that the size of each image feature input to the ultrasound gain feature fusion module is consistent. The aligned multi-scale image features are then fused to obtain the fused image features. It is understood that the above feature fusion methods may include feature concatenation, weighted feature addition, etc.
[0112] For example, Figure 4The input of the fourth feature fusion module C1 is connected to the output of the ultrasound gain feature fusion module and the output of the second feature extraction module A3, and its output is connected to the input of the first detection head module D1. The input of the fourth feature fusion module C2 is connected to the output of the ultrasound gain feature fusion module and the output of the second feature extraction module A4, and its output is connected to the input of the first detection head module D2. The input of the fourth feature fusion module C3 is connected to the output of the ultrasound gain feature fusion module and the output of the second feature extraction module A5, and its output is connected to the input of the first detection head module D3. The fourth feature fusion module C1 can receive the fused image features output by the ultrasound gain feature fusion module and the sixth-scale image features output by the second feature extraction module A3. Furthermore, the fused image features are upsampled so that the size of the upsampled fused image features is the same as the size of the sixth-scale image features. Then, the upsampled and fused image features are fused with the sixth-scale image features (e.g., feature stitching, weighted feature addition, etc.). The upsampled and fused image features are fused with the sixth-scale image features into a single image feature, which is then output to the first detection head module D1. Based on the same principle, the fourth feature fusion module C2 can generate and output image features to the first detection head module D2 based on the input fused image features and the seventh-scale image features. The fourth feature fusion module C3 can generate and output image features to the first detection head module D3 based on the input fused image features and the eighth-scale image features. The specific structure of the first detection head module is not limited in this embodiment of the invention; it can be used to determine image regions.
[0113] It is understood that the above content is not intended to limit the number of the second feature extraction module, the third feature extraction module, the fourth feature fusion module, and the first detection head module. The number of these modules can be adjusted according to the actual situation.
[0114] According to the above-described scheme of the present invention, the target detection model constructs feature representations at different scales by employing multi-level cascaded second feature extraction modules, and combines multiple sets of third feature extraction modules with different receptive field scales to achieve multi-dimensional feature enhancement. The enhanced features are then fused with image features output from the second feature extraction modules at different levels and input into multiple first detection head modules corresponding to different receptive field scales. This effectively improves the target detection model's ability to extract relevant features of the target region and its robustness. Simultaneously, the division of labor detection mechanism using multiple first detection head modules enables full coverage detection of regions of interest of different sizes in the target ultrasound image. Compared to a single detection head module, this reduces the probability of missed detections in the target region, improving the accuracy and reliability of target region detection, and consequently enhancing the accuracy and robustness of the ultrasound image processing method.
[0115] For example, before the training is completed, the trained region detection module also includes a second detection head module to be trained, which is connected to the ultrasound gain feature fusion module to be trained.
[0116] Figure 5 A schematic diagram of a region detection module according to an embodiment of the present invention is shown. Figure 5 As shown, with Figure 5 The ultrasonic gain feature fusion module in the example is the ultrasonic gain feature fusion module to be trained, and the second detection head module is the second detection head module to be trained. Figure 5 The input end of the second detection head module is connected to the output end of the ultrasonic gain feature fusion module.
[0117] The ultrasound image processing method also includes: using the second detection head module to be trained, based on the gain training features output by the ultrasound gain feature fusion module to be trained, to determine the gain prediction level corresponding to the training ultrasound image sequence.
[0118] The gain training features are obtained based on the second fusion features corresponding to the training ultrasound image sequence, sequentially through the second feature extraction module to be trained, the third feature extraction module to be trained, and the ultrasound gain feature fusion module to be trained. For example, the extraction process of the gain training features (i.e., the fused image features output by the ultrasound gain feature fusion module) is as follows: the second fusion features corresponding to the training ultrasound image sequence are input to the cascaded second feature extraction module to be trained, obtaining multi-scale first image features. Multiple sets of third feature extraction modules to be trained extract the corresponding level of first image features to obtain second image features. Then, the ultrasound gain feature fusion module to be trained fuses multiple sets of second image features to finally obtain the gain training features.
[0119] The gain prediction level indicates the predicted full-image gain of the training ultrasound image sequence. For example, continuing with the gain level example above, a low gain prediction level means the predicted full-image gain of the training ultrasound image sequence falls within the range of full-image gain values corresponding to the low level. A medium gain prediction level means the predicted full-image gain falls within the range of full-image gain values corresponding to the medium level. A high gain prediction level means the predicted full-image gain falls within the range of full-image gain values corresponding to the high level.
[0120] This invention provides an example for reference. See also Figure 5 As shown, with Figure 5 Taking the ultrasound gain feature fusion module (as the ultrasound gain feature fusion module to be trained), the second feature extraction modules A1 to A5 (all second feature extraction modules to be trained), and the third feature extraction modules B1 to B3 (all third feature extraction modules to be trained) as an example, the gain training features output by the ultrasound gain feature fusion module can be determined through the process of determining the fused image features output by the ultrasound gain feature fusion module described above. The specific process will not be elaborated here. The ultrasound gain feature fusion module outputs the gain training features to the second detection head module. The second detection head module can perform global average pooling on the above feature fusion processing results to obtain a global feature vector. Then, the second detection head module performs classification processing on the global feature vector. The second detection head module can use a 1×1 convolutional layer or a fully connected layer to perform classification processing on the global feature vector to obtain the confidence level of each gain level corresponding to the feature fusion processing result. The target detection model to be trained can use the gain level with the highest confidence as the gain prediction level corresponding to the above-mentioned training ultrasound image sequence.
[0121] Understandably, the second detection head module serves as an auxiliary supervision branch during the training phase of the target detection model. It only participates in the training process and not in the target detection flow during the inference phase. During training, the ultrasonic gain feature fusion module can be guided to learn robust feature representations for different gain levels through auxiliary supervision of gain level loss. This allows the trained target detection model to adaptively suppress feature interference caused by gain fluctuations during the inference phase without the need for additional gain prediction steps, thereby improving detection robustness.
[0122] The ultrasound image processing method also includes: determining the gain level loss based on the gain prediction level corresponding to the training ultrasound image sequence and the ultrasound gain level corresponding to the training ultrasound image sequence.
[0123] This invention provides an example for reference. Gain level loss can be determined based on the difference between the predicted gain level corresponding to the training ultrasound image sequence and the corresponding ultrasound gain level. Specifically, as shown above regarding gain levels (i.e., which may include low, medium, and high levels), a multi-class cross-entropy loss function can be used as the loss function to determine the gain level loss. This loss function measures the difference between the distribution of predicted confidence scores for each gain level output by the second detection head module, normalized by a normalized exponential function (such as softmax), and the one-hot encoded distribution of the pre-labeled true ultrasound gain levels (i.e., ultrasound gain levels) of the training ultrasound image sequence. The calculated loss value is used as the aforementioned gain level loss.
[0124] Gain level loss is a portion of the overall loss used to adjust the model parameters of the target detection model before it is fully trained. For example, the main loss, dynamic change degree loss, and gain level loss mentioned above can be weighted and summed according to preset weighting coefficients to obtain the overall loss used to adjust the model parameters of the target detection model.
[0125] According to the above-described scheme of the present invention, the gain prediction level corresponding to the training ultrasound image sequence can be determined based on the gain training features by using the second detection head module to be trained. Furthermore, the gain level loss can be determined based on the gain prediction level and the ultrasound gain level corresponding to the training ultrasound image sequence.
[0126] For example, the trained object detection model is also trained based on the region category corresponding to each training region.
[0127] Region categories are used to represent the disease category or tissue name of the tissue of interest in the training region. For example, the disease category and tissue name of the tissue of interest can be determined according to the training objective of the target detection model. Taking the training objective of the target detection model as identifying lesions in the ovary as an example, the above-mentioned disease categories may include: follicular cyst, luteal cyst, corpus luteum hematoma, etc. The above-mentioned tissue names may include: follicle, corpus luteum, cyst wall, cyst septum, fallopian tube, etc.
[0128] This invention provides a training method for an object detection model for reference. A third training dataset is obtained. The object detection model is iteratively trained using the third training dataset, and the model parameters are adjusted using a loss function until training is complete. The third training dataset includes multiple training ultrasound image sequences. The labels for each training ultrasound image sequence may include: the training dynamic change sequence corresponding to the training ultrasound image sequence, at least one training region in the last training ultrasound image in the training ultrasound image sequence, and the disease name and / or tissue name of the tissue of interest corresponding to each training region. The conditions for training completion may include the loss value calculated by the loss function stabilizing, the number of iterations reaching a preset number, etc.
[0129] According to the above-described scheme of the present invention, the target detection model trained based on the region category corresponding to each training region can determine the region category corresponding to at least one target region in the last target ultrasound image, so as to better assist the user in obtaining relevant information about the target object.
[0130] According to another aspect of the present invention, an ultrasonic device is also provided. Figure 6 A schematic block diagram of an ultrasonic device according to an embodiment of the present invention is shown, such as Figure 6 As shown, the ultrasound device 200 includes a memory 210 and a processor 220, wherein the memory 210 is used to store computer programs; and the processor 220 is used to execute the computer programs to implement the above-described ultrasound image processing method.
[0131] Furthermore, according to another aspect of the present invention, a storage medium is provided, on which program instructions are stored. When the program instructions are executed by a computer or processor, the computer or processor performs corresponding steps of the ultrasound image processing method described in the embodiments of the present invention, and is used to implement corresponding modules in the ultrasound device described in the embodiments of the present invention. The storage medium may, for example, include a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media. According to yet another aspect of the present invention, a computer program product is also provided, including computer program instructions. When the computer program instructions are executed by a computer or processor, the computer or processor performs the corresponding steps mentioned in the ultrasound device described above.
[0132] Those skilled in the art can understand the specific implementation scheme of the storage medium by reading the relevant descriptions of the corresponding steps mentioned above regarding the ultrasonic equipment; for the sake of brevity, they will not be elaborated further here.
[0133] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of the invention. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of the invention. All such changes and modifications are intended to be included within the scope of the invention as claimed in the appended claims.
[0134] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0135] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0136] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0137] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of the invention. However, this approach should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, its inventive point lies in solving the corresponding technical problem with fewer features than all of those in a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0138] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or elements of any method or apparatus so disclosed may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0139] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.
[0140] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the ultrasonic device according to embodiments of the present invention. The present invention can also be implemented as a system program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0141] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0142] The above description is merely a specific embodiment of the present invention or an explanation of that embodiment. The scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for processing ultrasound images, characterized in that, The method includes: Acquire a target ultrasound image sequence of a target object, wherein the target ultrasound image sequence includes target ultrasound images of the target object acquired sequentially; The target ultrasound image sequence is input into a trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence. The trained target detection model is trained based on multiple training ultrasound image sequences, a training dynamic change degree sequence corresponding to each training ultrasound image sequence, and at least one training region in the last training ultrasound image of each training ultrasound image sequence. For each training ultrasound image sequence, the training dynamic change degree sequence corresponding to the training ultrasound image sequence is used to represent the degree of dynamic change of the labeled images corresponding to each of the multiple training ultrasound images in the training ultrasound image sequence. The target region is used to represent the image region in the target ultrasound image where the tissue of interest exists, and the training region is used to represent the image region in the training ultrasound image where the tissue of interest exists.
2. The method as described in claim 1, characterized in that, The trained target detection model includes: multiple first feature extraction modules, a first feature fusion module connected to all of the multiple first feature extraction modules, a target dynamic change degree determination module connected to the first feature fusion module, a second feature fusion module connected to both the first feature fusion module and the target dynamic change degree determination module, and a region detection module connected to the second feature fusion module. The step of inputting the target ultrasound image sequence into the trained target detection model to obtain and display at least one target region in the last target ultrasound image of the target ultrasound image sequence includes: For each of the plurality of first feature extraction modules, the target ultrasound image input to the first feature extraction module is processed by the first feature extraction module to obtain the target ultrasound image features corresponding to the target ultrasound image; The first feature fusion module performs feature splicing processing on the target ultrasound image features corresponding to each target ultrasound image to obtain the first fusion feature corresponding to the target ultrasound image sequence. The target dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the first fusion feature. For each target ultrasound image, the target dynamic change degree corresponding to the target ultrasound image is used to represent the dynamic change degree of the image corresponding to the target ultrasound image. The second feature fusion module determines the second fusion feature corresponding to the target ultrasound image sequence based on the target dynamic change degree corresponding to each target ultrasound image and the first fusion feature. Based on the second fusion feature, the region detection module determines and displays at least one target region in the last target ultrasound image of the target ultrasound image sequence.
3. The method as described in claim 2, characterized in that, The target dynamic change degree determination module includes a floating point displacement feature extraction module, a cyst septum deformation feature extraction module, a lesion deformation feature extraction module, a third feature fusion module connected to the floating point displacement feature extraction module, the cyst septum deformation feature extraction module, and the lesion deformation feature extraction module, and a first dynamic change degree determination module connected to the third feature fusion module. The step of determining the target dynamic change degree corresponding to each target ultrasound image based on the first fused feature includes: The floating point displacement feature extraction module extracts features from the first fused features to obtain floating point displacement features, wherein the floating point displacement features are used to represent the image features of moving floating points in the target ultrasound image sequence. The cyst septum deformation feature extraction module extracts features from the first fused features to obtain cyst septum deformation features. The cyst septum deformation features are used to represent the image features of deformed cyst septum in the target ultrasound image sequence. The receptive field of the cyst septum deformation feature extraction module is larger than that of the floating point displacement feature extraction module. The lesion deformation feature extraction module extracts features from the first fused features to obtain lesion deformation features. The lesion deformation features are used to represent the image features of deformed lesions in the target ultrasound image sequence. The receptive field of the lesion deformation feature extraction module is larger than that of the cyst septum deformation feature extraction module. The third feature fusion module performs feature fusion processing on the displacement features of the floating point-like objects, the deformation features of the cyst septum, and the deformation features of the lesion to obtain the third fused feature. The first dynamic change degree determination module determines the target dynamic change degree corresponding to each target ultrasound image based on the third fusion feature.
4. The method as described in claim 2, characterized in that, The step of determining and displaying at least one target region in the last target ultrasound image in the target ultrasound image sequence based on the second fusion feature includes: The second fusion feature is compressed and fused in the channel dimension, and based on the compressed and fused second fusion feature, at least one target region in the last target ultrasound image in the target ultrasound image sequence is determined and displayed.
5. The method as described in claim 2, characterized in that, The method further includes: For each training ultrasound image sequence, the target dynamic change degree determination module determines the predicted dynamic change degree sequence corresponding to the training ultrasound image sequence. Based on the predicted dynamic change degree sequence and the training dynamic change degree sequence corresponding to the training ultrasound image sequence, the dynamic change degree loss is determined. The predicted dynamic change degree sequence corresponding to the training ultrasound image sequence is used to represent the predicted image dynamic change degree corresponding to each training ultrasound image in the training ultrasound image sequence. The dynamic change degree loss is part of the overall loss used to adjust the model parameters of the target detection model to be trained.
6. The method as described in claim 2, characterized in that, The trained target detection model is also trained based on the ultrasound gain level corresponding to each training ultrasound image sequence. For each training ultrasound image sequence, the ultrasound gain level corresponding to the training ultrasound image sequence is used to represent the full-image gain level of the training ultrasound image sequence.
7. The method as described in claim 6, characterized in that, The region detection module includes multiple sequentially connected trained second feature extraction modules, multiple trained third feature extraction modules, a trained ultrasound gain feature fusion module connected to all of the trained third feature extraction modules, multiple trained fourth feature fusion modules, and multiple trained first detection head modules. For each of the trained third feature extraction modules, this trained third feature extraction module is connected to a trained second feature extraction module, and this trained second feature extraction module is different from the other trained third feature extraction modules. Similarly, for each of the trained fourth feature fusion modules... The fourth feature fusion module, after training, is connected to both the trained ultrasound gain feature fusion module and a trained second feature extraction module. The trained second feature extraction module connected to the fourth feature fusion module is different from the trained second feature extraction modules connected to the other trained fourth feature fusion modules. For each of the plurality of trained first detection head modules, the trained first detection head module is connected to a trained fourth feature fusion module, and the trained fourth feature fusion module connected to the first detection head module is different from the trained fourth feature fusion modules connected to the other trained first detection head modules.
8. The method as described in claim 7, characterized in that, Before training is complete, the trained region detection module also includes a second detection head module to be trained, which is connected to the ultrasound gain feature fusion module to be trained. The method further includes: The gain prediction level corresponding to the training ultrasound image sequence is determined by the second detection head module to be trained, based on the gain training features output by the ultrasound gain feature fusion module to be trained. The gain training features are obtained by sequentially passing through the second feature extraction module, the third feature extraction module, and the ultrasound gain feature fusion module to be trained, based on the second fusion features corresponding to the training ultrasound image sequence. The gain prediction level is used to represent the full-image gain degree predicted by the training ultrasound image sequence. The method further includes: Based on the gain prediction level corresponding to the training ultrasound image sequence and the ultrasound gain level corresponding to the training ultrasound image sequence, a gain level loss is determined, wherein the gain level loss is part of the overall loss used to adjust the model parameters of the target detection model to be trained.
9. The method as described in claim 1, characterized in that, The trained target detection model is also trained based on the region category corresponding to each training region. The region category is used to represent the disease category or organization name of the organization of interest in the training region.
10. An ultrasonic device, characterized in that, The device includes a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program to implement the ultrasound image processing method as described in any one of claims 1-9.
11. A storage medium storing computer program instructions, characterized in that, The computer program instructions, when executed, are used to perform the ultrasound image processing method as described in any one of claims 1-9.
12. A computer program product comprising computer program instructions, characterized in that, The computer program instructions, when executed by a processor, are used to perform the ultrasound image processing method as described in any one of claims 1-9.