Ultrasonic image processing method and device, computer equipment and storage medium

By performing multi-category semantic segmentation and measuring the diameter table length of the ultrasound image collection, the problem of inaccurate ultrasound image processing is solved, and higher consistency and accuracy of inspection results are achieved.

CN120339371AActive Publication Date: 2025-07-18REPRODUCTIVE & GENETIC HOSPITAL OF CITIC XIANGYA CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510312291.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18
Estimated Expiration
2045-03-17

Smart Images

  • Figure CN120339371A_ABST
    Figure CN120339371A_ABST
Patent Text Reader

Abstract

The invention relates to an ultrasonic image processing method and device, computer equipment and a storage medium. The method comprises the steps that an ultrasonic image set is obtained, each ultrasonic image in the ultrasonic image set comprises multiple classes of recognition targets, multi-class semantic segmentation is conducted on the ultrasonic images in the ultrasonic image set on different preset segmentation scales, and a segmentation result set of the recognition targets of all classes is obtained; and for the segmentation result set of each category of identification targets, measuring the diameter table length of the identification target of each segmentation result in the segmentation result set, and determining the measured maximum diameter table length as the target diameter table length of the category of identification targets. By adopting the method, the ultrasonic image processing accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to an ultrasonic image processing method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] With the development of ultrasonic technology, ultrasonic technology has gradually been introduced into clinical applications, and the detection results of ultrasonic images can usually be used as important auxiliary information for doctors' diagnosis and treatment. For example, ultrasonic technology is used for prenatal examination of pregnant women, and the growth and development of the fetus can be evaluated using ultrasonic images.

[0003] However, in traditional solutions, the processing and analysis of ultrasonic images may not be accurate enough, and may also result in a low consistency of inspection results. Summary of the Invention

[0004] Based on this, it is necessary to provide an ultrasonic image processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of ultrasonic image processing for the above technical problems.

[0005] In a first aspect, this application provides an ultrasonic image processing method. The method includes:

[0006] Obtain a set of ultrasonic images, where each ultrasonic image in the set of ultrasonic images contains multiple categories of recognition targets;

[0007] Perform multi-category semantic segmentation on each ultrasonic image in the set of ultrasonic images at different preset segmentation scales to obtain a set of segmentation results of the recognition targets of each category;

[0008] For each set of segmentation results of the recognition targets of each category, measure the radial table length of the recognition target of each segmentation result in the set of segmentation results, and determine the maximum measured radial table length as the target radial table length of the recognition target of the category.

[0009] In a second aspect, this application also provides an ultrasonic image processing apparatus. The apparatus includes:

[0010] A data acquisition module, configured to obtain a set of ultrasonic images, where each ultrasonic image in the set of ultrasonic images contains multiple categories of recognition targets;

[0011] A semantic segmentation module, configured to perform multi-category semantic segmentation on each ultrasonic image in the set of ultrasonic images at different preset segmentation scales to obtain a set of segmentation results of the recognition targets of each category;

[0012] A data measurement module, which is used to measure the diameter length of the recognition target of each segmentation result in the segmentation result set of the recognition targets of each category, and determine the target diameter length of the recognition target of the category as the maximum measured diameter length.

[0013] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above-mentioned embodiment of the ultrasonic image processing method are implemented.

[0014] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned embodiment of the ultrasonic image processing method are implemented.

[0015] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned embodiment of the ultrasonic image processing method are implemented.

[0016] For the above ultrasonic image processing method, device, computer device, storage medium and computer program product, the present application obtains a set of ultrasonic images, where each ultrasonic image includes recognition targets of multiple categories, and then performs multi-category semantic segmentation on each ultrasonic image at different preset segmentation scales. A large segmentation scale can capture the global information of the ultrasonic image, and a small segmentation scale can capture the detailed information of the ultrasonic image, so that the recognition targets of each category can be more comprehensively and accurately recognized, thereby improving the accuracy of ultrasonic image processing and analysis. Further, for the segmentation result set of the recognition targets of each category, by measuring the diameter length of the recognition target in each segmentation result and determining the maximum diameter length as the target diameter length of the recognition target of the category, this quantitative measurement method is calculated based on objective data, reducing the error that may be brought by the doctor's subjective judgment and further improving the accuracy of the ultrasonic image analysis result. Description of the Drawings

[0017] Figure 1 It is an application environment diagram of the ultrasonic image processing method in an embodiment;

[0018] Figure 2 It is a flowchart of the ultrasonic image processing method in an embodiment;

[0019] Figure 3 It is a flowchart of the ultrasonic image processing method in another embodiment;

[0020] Figure 4Schematic flowchart of the semantic segmentation step in an embodiment;

[0021] Figure 5 Schematic flowchart of the class presence probability prediction step in an embodiment;

[0022] Figure 6 Schematic flowchart of the ultrasonic image processing method in a detailed embodiment;

[0023] Figure 7 Schematic diagram of the multi-class semantic segmentation in an embodiment;

[0024] Figure 8 Block diagram of the structure of the ultrasonic image processing device in an embodiment;

[0025] Figure 9 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0026] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0027] The ultrasonic image processing method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers.

[0028] Specifically, an operator can collect an ultrasonic image set through an ultrasonic acquisition device. Each ultrasonic image in the ultrasonic image set contains multi-class recognition targets. The operator can upload the ultrasonic image set to the server 104 through the terminal 102. The server 104 performs multi-class semantic segmentation on each ultrasonic image in the ultrasonic image set at different preset segmentation scales to obtain a segmentation result set of the recognition targets of each class. Further, for each segmentation result set of the recognition targets of each class, the server 104 measures the radial table length of the recognition target of each segmentation result in the segmentation result set, and determines the maximum measured radial table length as the target radial table length of the recognition target of the class.

[0029] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0030] In one embodiment, as Figure 2 shown, an ultrasonic image processing method is provided. Taking the method applied to the Figure 1 server 104 in it as an example, it includes the following steps:

[0031] S100, obtain an ultrasonic image set.

[0032] Among them, each ultrasonic image in the ultrasonic image set contains multiple categories of recognition targets. The ultrasonic image set includes multiple ultrasonic images. The ultrasonic image is a picture obtained by using ultrasonic waves to image the internal organs or tissues of the human body. In this embodiment, taking the application of ultrasonic technology to early pregnancy embryo examination as an example, each ultrasonic image may include multiple categories of recognition targets, such as gestational sac, yolk sac, and germ.

[0033] Specifically, it can be that a medical staff manually scans the fetus in the pregnant woman's abdomen through an ultrasonic imaging device, so as to obtain multiple frames of ultrasonic images, forming an ultrasonic image set. For example, in a prenatal examination, the ultrasonic imaging device collected 10 frames of ultrasonic images, and each ultrasonic image contains multiple categories of recognition targets such as gestational sac, yolk sac, and germ.

[0034] S200, perform multi-category semantic segmentation on each ultrasonic image in the ultrasonic image set at different preset segmentation scales, and obtain a segmentation result set of each category of recognition target.

[0035] Among them, the goal of semantic segmentation is to divide each pixel point in the ultrasonic image into a specific category, and multi-category semantic segmentation means classifying the pixel points in the ultrasonic image into different categories. For example, classifying the pixel points corresponding to the gestational sac, yolk sac, germ, etc. in the ultrasonic image into the corresponding categories. When performing multi-category semantic segmentation, in order to analyze the ultrasonic image from different detail levels, different segmentation scales can be preset. The larger the segmentation scale, the more it means paying attention to the macroscopic and overall perspective during the segmentation process. The smaller the segmentation scale, the more it means paying attention to the detailed part of the ultrasonic image and being able to discover tiny features. For example, in this embodiment, 4 different preset segmentation scales can be set to achieve semantic segmentation from micro to macro.

[0036] The segmentation result set refers to the summary of all segmentation results obtained for each category after performing multi-class semantic segmentation on each ultrasound image in the ultrasound image set. For example, for the category of gestational sac, semantic segmentation is performed on the ultrasound images in the ultrasound image set, and the segmentation result set of the gestational sac finally obtained contains each ultrasound image marked with the gestational sac.

[0037] Exemplarily, for each ultrasound image in the ultrasound image set, multi-class semantic segmentation operations are performed at different preset segmentation scales. For example, 4 different segmentation scales from small to large can be preset in advance. At each segmentation scale, the server divides the pixel points in the ultrasound image into corresponding categories. For example, a part is divided into the gestational sac, a part is divided into the yolk sac, a part is divided into the embryo bud, a part is divided into the background, etc. It should be noted that in multi-class semantic segmentation, the recognition targets of various categories can be marked simultaneously in the segmentation result corresponding to each ultrasound image, or the recognition targets of different categories can be marked separately, and then the segmentation result sets of different categories can be obtained. The segmentation result sets of different categories can also be output through different independent output channels. In the segmentation result set of each category, it contains the segmentation results of multiple ultrasound images for this category at different segmentation scales. Since the ultrasound image may be scaled during semantic segmentation, different segmentation scales will affect the scaling degree. Therefore, the segmentation result can be further downsampled to restore it to the size of the input ultrasound image, and by averaging the segmentation results of different segmentation scales, for example, averaging the corresponding pixel values of the ultrasound images at different segmentation scales, the segmentation result set of the recognition target of each category is obtained (taking 4 preset segmentation scales as an example, one ultrasound image should correspond to 4 segmentation results. After the averaging operation, one ultrasound image corresponds to one segmentation result, and the segmentation result set includes the segmentation results corresponding to each of the multiple ultrasound images).

[0038] S300, for the segmentation result set of the recognition target of each category, measure the diameter length of the recognition target of each segmentation result in the segmentation result set, and determine the target diameter length of the recognition target of the category as the maximum diameter length measured.

[0039] Among them, the diameter length refers to the longest radial distance of the recognition target in the segmentation result. For example, if the gestational sac is included in the segmentation result, the length measured along the longest direction of the gestational sac is the diameter length of the gestational sac (recognition target).

[0040] Exemplarily, the set of segmentation results of the recognition targets of each category should include multiple segmentation results. Taking the gestational sac as the recognition target as an example, for each segmentation result (each processed ultrasound image), determine the length of the gestational sac in the direction of the longest axis in this segmentation result, and this length is the diameter table length of the gestational sac. Each segmentation result corresponds to a diameter table length, so finally a set of diameter table lengths can be obtained. Further, find the largest diameter table length in the set of diameter table lengths and determine it as the target diameter table length of the recognition target of this category. For example, compare the diameter table lengths of the gestational sac measured based on each ultrasound image and select the largest one as the target diameter table length of the gestational sac. This is because the largest diameter table length can, to a certain extent, most representatively reflect the maximum size of the recognition target of this category in the natural state, which is of great significance for subsequent clinical analysis and clinical evaluation.

[0041] The above ultrasonic image processing method is different from traditional manual analysis. In this application, by obtaining a set of ultrasonic images, each of which includes recognition targets of multiple categories, and then performing multi-category semantic segmentation on each ultrasonic image at different preset segmentation scales. The large segmentation scale can capture the global information of the ultrasonic image, and the small segmentation scale can capture the detailed information of the ultrasonic image, so that the recognition targets of each category can be recognized more comprehensively and accurately, thereby improving the accuracy of ultrasonic image processing and analysis. Further, for the set of segmentation results of the recognition targets of each category, by measuring the diameter table length of the recognition target in each segmentation result and determining the largest diameter table length as the target diameter table length of the recognition target of this category, this quantitative measurement method is calculated based on objective data, reducing the error that may be brought by the doctor's subjective judgment and further improving the accuracy of ultrasonic image analysis.

[0042] In one embodiment, as Figure 3 shown, S200 includes:

[0043] S210, taking the set of ultrasonic images as the input, calling the trained multi-category semantic segmentation model, and obtaining the set of segmentation results of the recognition targets of each category.

[0044] Among them, the trained multi-class semantic segmentation model includes multiple independent output channels, and each independent output channel is used to output a set of segmentation results of the recognition targets of a category. The trained multi-class semantic segmentation model is trained based on a set of historical ultrasound images with category labels. The multi-class semantic segmentation model is a model constructed based on deep learning technology, which can analyze the input ultrasound images and divide each pixel in the ultrasound images into different categories, so as to achieve semantic segmentation of different targets in the ultrasound images. For example, the multi-class semantic segmentation model in this embodiment can distinguish different categories of recognition targets such as gestational sac, yolk sac, and germ. Further, in order to facilitate subsequent analysis of the segmentation results of different categories of recognition targets respectively, the multi-class semantic segmentation model in this embodiment is designed to have multiple independent output channels, and each independent output channel is specifically responsible for outputting the segmentation results of the recognition targets of a category. For example, the segmentation results of the gestational sac, yolk sac, and germ are output through different independent output channels.

[0045] Exemplarily, the multi-class semantic segmentation model in this application can be constructed and trained in the following manner. First, a multi-class semantic segmentation model to be trained can be constructed based on a deep learning network. The multi-class semantic segmentation model to be trained includes a Backbone part, a Neck part, and a Prediction part. Among them, the Backbone part is a combination of multiple convolutional modules and residual modules, which is used to extract high-dimensional features (forming a feature map) in the ultrasound images. Different convolutional modules have different parameters such as convolutional kernel size and stride, and can extract features of different scales and levels. The residual module can enable the deep learning network to directly learn the residuals between the input and the output, reducing the training difficulty of the deep learning network. The Neck part is used to perform multiple upsampling and downsampling operations on the high-dimensional features in the ultrasound images. Upsampling refers to increasing the resolution of the feature map so that the feature map can contain more detailed information, while downsampling is to reduce the resolution of the feature map, reduce the amount of data, and at the same time extract more abstract features. Through repeated upsampling and downsampling, the Neck part can fuse features of multiple granularities. The Prediction part can output the category probability of each pixel point based on the Softmax function. For each pixel point in the ultrasound image, after being processed by the previous Backbone part and Neck part, a feature representation of the pixel point will be obtained. These feature representations are input into the Softmax function, and the Softmax function will calculate the probability that the pixel point belongs to each category. For example, in ultrasound image segmentation, after each pixel point is processed by the Softmax function, probability values belonging to different categories such as gestational sac, yolk sac, germ, and background will be obtained.

[0046] After constructing the multi-class semantic segmentation model to be trained, appropriate training data should be selected to train it. For example, the training data can be 1000 historical ultrasound images containing normal early pregnancy samples, and 500 historical ultrasound images for each type of early pregnancy abnormal sample (500 ultrasound images containing abnormal gestational sacs, 500 ultrasound images containing abnormal yolk sacs, 500 ultrasound images containing abnormal embryos, etc.). Each historical ultrasound image carries a class label, that is, in each historical ultrasound image, the gestational sac, yolk sac, and embryo in it are labeled. Using the above training data to train the multi-class semantic segmentation model, since the training data contains a large number of ultrasound images and corresponding pixel-level class annotation information, during the training process, the multi-class semantic segmentation model can gradually learn the features and patterns of the recognition targets of different classes in the training data according to the input ultrasound images and the corresponding annotation information, thereby improving the accuracy and performance of multi-class semantic segmentation of ultrasound images, and finally obtaining the trained multi-class semantic segmentation model.

[0047] Further, taking the obtained ultrasound image data set as the input data and inputting it into the trained multi-class semantic segmentation model, the trained multi-class semantic segmentation model will classify each pixel in the ultrasound image based on the knowledge learned during the training process to determine which class of recognition target it belongs to. Since the trained multi-class semantic segmentation model has multiple independent output channels, and each independent output channel corresponds to a class of recognition target, when the model processes the ultrasound image, each independent output channel will separately output the segmentation result of the recognition target of the class responsible for that channel.

[0048] In this embodiment, the trained multi-class semantic segmentation model is trained based on the historical ultrasound image set with class labels. Therefore, the trained multi-class semantic segmentation model can capture the feature information of different classes of recognition targets, and each independent output channel of the model is responsible for outputting the segmentation result of a class of recognition target, which can enable the model to focus on the feature learning and segmentation of each class, thereby improving the accuracy of the segmentation result of each class.

[0049] In one embodiment, as Figure 4 shown, when the trained multi-class semantic segmentation model is called, the following steps are executed:

[0050] S220, perform multi-class semantic segmentation on each ultrasound image in the ultrasound image set at different preset segmentation scales to obtain a set of segmentation results corresponding to different preset segmentation scales.

[0051] S230 , performing upsampling operations on different segmentation result sets to obtain segmentation result sets corresponding to different preset segmentation scales at the original size.

[0052] S240, averaging the segmentation result sets corresponding to different preset segmentation scales at the original size to obtain segmentation result sets of each category of recognition targets.

[0053] The original size is the size of the ultrasound image in the ultrasound image set. When the trained multi-category semantic segmentation model is called, the model will perform multi-category semantic segmentation on each ultrasound image in the ultrasound image set at different preset segmentation scales. Different preset segmentation scales mean that the model will analyze the ultrasound image from multiple resolutions or receptive fields.

[0054] Specifically, four different segmentation scales can be set. A larger segmentation scale can quickly capture the overall structure and general contour information in the ultrasound image, while a smaller segmentation scale can mine the detailed features of the ultrasound image, and can more accurately divide the edges and subtle structures of the identified target. For example, in early pregnancy ultrasound images, the scale of the gestational sac is much larger than that of the yolk sac and the embryo. Single-scale semantic segmentation tends to focus on predicting the segmentation results of the gestational sac and ignore the yolk sac and the embryo. Therefore, by allowing the multi-category semantic segmentation model to perform semantic segmentation on ultrasound images at different segmentation scales, the accuracy of the segmentation results can be improved.

[0055] Furthermore, after obtaining the set of segmentation results corresponding to different preset segmentation scales, since the size of the segmentation result is usually different from the size of the original ultrasound image when semantic segmentation is performed at different scales, it is necessary to perform an upsampling operation on it to restore the size of the segmentation result to the size of the input ultrasound image, so that the segmentation results at different scales can be subsequently comprehensively analyzed on the basis of the same size.

[0056] After obtaining the set of segmentation results corresponding to different preset segmentation scales under the original size, the average can be calculated to synthesize the segmentation information at different scales. For example, in the set of segmentation results corresponding to different preset segmentation scales, it includes the probabilities of each pixel point on each ultrasound image belonging to each category, such as the probabilities of each pixel point belonging to the gestational sac, yolk sac, and embryo respectively. Taking the average means that for each ultrasound image, the probabilities of each pixel point belonging to each category obtained by analysis at different segmentation scales are averaged correspondingly as the final segmentation result. Specifically, assuming that four segmentation scales are preset, for a certain pixel point on a certain ultrasound image, its probabilities of belonging to the gestational sac, yolk sac, embryo, and background at these four segmentation scales are (0.1, 0.5, 0.4, 0), (0.2, 0.6, 0.2, 0), (0.1, 0.6, 0.3, 0), and (0.1, 0.4, 0.5, 0) respectively. Then, after the averaging operation, the probabilities of this pixel point belonging to the gestational sac, yolk sac, embryo, and background are (0.125, 0.525, 0.35, 0).

[0057] In this embodiment, by performing semantic segmentation at different preset segmentation scales and synthesizing the information of different segmentation scales through the averaging operation, the limitations brought by a single segmentation scale can be effectively reduced, misjudgment can be reduced, and the semantic segmentation accuracy of various recognition targets can be improved.

[0058] In one embodiment, as Figure 5 shown, the trained multi-class semantic segmentation model includes a multi-layer perceptron. When the trained multi-class semantic segmentation model is called, the following steps are also executed:

[0059] S250, for each ultrasound image, perform feature extraction and pooling operations on the ultrasound image to obtain feature vectors of recognition targets of different categories at different preset segmentation scales, and splice the feature vectors of recognition targets of the same category at different preset segmentation scales to obtain the vector splicing results of recognition targets of each category.

[0060] S260, taking the vector splicing results of recognition targets of each category as input, call the multi-layer perceptron to obtain the category existence probabilities of recognition targets of different categories in each ultrasound image.

[0061] S270, if the category probability of the recognition target is less than the preset category probability threshold, then filter out the segmentation result corresponding to the recognition target.

[0062] Among them, a multi-layer perceptron is a network structure composed of multiple layers of neurons, which can be used to predict the probability of the presence of different categories of recognition targets in an ultrasound image based on the input feature vector. Feature extraction refers to extracting information from an ultrasound image that can represent the image features. The pooling operation is a downsampling operation that can aggregate local regions of the feature map, reduce the data volume and computational complexity, and at the same time retain the main feature information.

[0063] Exemplarily, for each ultrasound image in the ultrasound image set, first perform feature extraction and pooling operations. For example, the convolutional module in the multi-class semantic segmentation model encodes (feature extraction) the input ultrasound image, converts the image information into a feature vector, and then decodes the feature vector at different preset segmentation scales to restore the encoded feature vector to a feature representation related to the image semantics, obtaining the decoded feature vector. In addition to being sent to the convolutional head to output the segmentation result, the decoded feature vector is also input into a global average pooling layer, which will perform a global average pooling operation on the decoded feature vector. Assuming that the size of the input feature map is H*W*C (height * width * number of channels), after passing through the global average pooling layer, each channel will obtain an average value, and finally output a 1*C feature vector. Assuming that the number of preset segmentation scales is 4, after each ultrasound image undergoes the above processing, four groups of feature vectors can be obtained. Each group of feature vectors includes feature vectors of different categories of recognition targets. Vector splicing is performed on the feature vectors of the same category of recognition target at different preset segmentation scales, and a vector splicing result of 1*(C1 + C2 + C3 + C4) can be obtained.

[0064] Further, input the obtained vector concatenation result into a multi-layer perceptron. The multi-layer perceptron can perform non-linear transformation and learning on the input vector concatenation result to predict the existence probabilities of the gestational sac, yolk sac, and embryo in the ultrasound image. For example, input the concatenated 1*(C1 + C2 + C3 + C4) vector into the multi-layer perceptron, and the multi-layer perceptron outputs a 1*3 feature vector. The three values in this feature vector respectively represent the probabilities of the existence of the gestational sac, yolk sac, and embryo in this frame of ultrasound image. If the existence probability of an identified target of a certain category in the ultrasound image is less than a preset category probability threshold, for example, the existence probability of the gestational sac is less than 0.5, it can be considered that there is actually no gestational sac in the ultrasound image. Regardless of the corresponding segmentation result of the gestational sac, this segmentation result should be filtered, which means suppressing this segmentation result, and the segmentation result of the identified target of this category should not be used as the final valid output. This is because when there is no gestational sac in the ultrasound image, the segmentation output head may still predict the segmentation result of the gestational sac, which will lead to serious mismeasurements. Therefore, suppressing the segmentation result of the gestational sac can reduce possible misidentifications. It should be noted that since there are multiple frames of input ultrasound images, it is also possible that the existence probability of the gestational sac in some ultrasound images is less than 0.5, while the existence probability of the gestational sac in other ultrasound images is greater than or equal to 0.5. At this time, the way of the minority obeying the majority principle can be used to determine whether to suppress the segmentation result of the gestational sac.

[0065] In this embodiment, the global average pooling layer is used to process the feature vector, reducing the redundant information of the features. And by concatenating the feature vectors at multiple different segmentation scales, the recognition ability for gestational sacs, yolk sacs, and embryos of different sizes and shapes can be enhanced. Subsequently, according to the existence probability of the identified target output by the multi-layer perceptron, performing a suppression operation on the segmentation result can effectively reduce subsequent misjudgment situations and improve the efficiency and accuracy of ultrasound image processing.

[0066] In one embodiment, S300 includes: for each segmentation result, extract the contour pixel points of the identified target in the segmentation result, determine the straight-line distances of each pixel point in the contour pixel points, and determine the longest straight-line distance as the diameter length of the identified target.

[0067] Continuing with the above embodiment, after obtaining the segmentation result, first, it is necessary to extract the contour pixel points of the identified target from the segmentation result, for example, through an edge detection algorithm, etc., so as to accurately find the boundary of the identified target. Then, calculate the straight-line distances between any two pixel points in these contour pixel points, and determine the longest straight-line distance among them as the diameter length of the identified target. The diameter length reflects the maximum size of the identified target in a certain direction.

[0068] In this embodiment, through pixel-level distance calculation, compared with estimating the size of the recognition target solely by visual observation, the diameter length calculated at the pixel level is more objective and accurate, providing a reliable basis for subsequent analysis.

[0069] To make a clearer description of the ultrasonic image processing method provided in this application, the following will be combined with the attached Figure 6 and One A detailed embodiment is explained, and the detailed embodiment includes the following steps:

[0070] S601. Obtain a set of ultrasonic images, and perform multi-class semantic segmentation on each ultrasonic image in the set of ultrasonic images at different preset segmentation scales to obtain a set of segmentation results corresponding to different preset segmentation scales.

[0071] S602. Perform an upsampling operation on different sets of segmentation results to obtain a set of segmentation results corresponding to different preset segmentation scales at the original size, where the original size is the size of the ultrasonic images in the set of ultrasonic images.

[0072] S603. Take the average of the sets of segmentation results corresponding to different preset segmentation scales at the original size to obtain a set of segmentation results of the recognition targets of each category.

[0073] S604. For each ultrasonic image, perform feature extraction and pooling operations on the ultrasonic image to obtain feature vectors of the recognition targets of different categories at different preset segmentation scales, and splice the feature vectors of the recognition targets of the same category at different preset segmentation scales to obtain the vector splicing results of the recognition targets of each category.

[0074] S605. Use the vector splicing results of the recognition targets of each category as input, call a multi-layer perceptron to obtain the category existence probabilities of the recognition targets of different categories in each ultrasonic image.

[0075] S606. If the category probability of the recognition target is less than the preset category probability threshold, filter out the segmentation result of the recognition target.

[0076] S607. For the set of segmentation results of the recognition targets of each category, for each segmentation result, extract the contour pixel points of the recognition target in the segmentation result, and determine the straight-line distances of each pixel point in the contour pixel points.

[0077] S608. Determine the longest straight-line distance as the diameter length of the recognition target, and determine the maximum measured diameter length as the target diameter length of the recognition target of the category.

[0078] Among them, in the detailed embodiment, a trained multi-class semantic segmentation model is called, and the model architecture is as Figure 7As shown, the image input on the left is an ultrasound image, which is the input data of the entire model. The ultrasound image needs to undergo feature extraction in multiple stages. The segmentation scales in each feature extraction stage are different, and the feature extraction in each stage can be achieved through multiple convolutional modules, such as deformable convolutional modules, etc., to capture features more accurately. The features extracted in each feature extraction stage can respectively pass through a convolutional head and an average pooling layer. The convolutional head outputs the segmentation results corresponding to different preset segmentation scales, and the average pooling layer outputs the feature vectors after pooling. The feature vectors corresponding to different preset segmentation scales are concatenated and then input into a multi-layer perceptron to predict the probability of the existence of recognition targets of different categories in each ultrasound image.

[0079] It should be noted that when the multi-class semantic segmentation model is used for early pregnancy ultrasound image processing, there may be artifacts in the originally collected ultrasound images. For example, there are artifacts between the germ and the gestational sac, which may lead to inaccurate measurement results. Therefore, before inputting the ultrasound image set into the multi-class semantic segmentation model, it is necessary to perform data cleaning on the ultrasound image set to screen out the ultrasound images with poor quality. And, considering that there may be a situation where the yolk sac and the germ are adhered together in the early pregnancy ultrasound image, at this time, the multi-class semantic segmentation model may segment the yolk sac together as the germ, resulting in an overly large measured length of the germ diameter table. Therefore, when training the multi-class semantic segmentation model, ultrasound images containing the above situation can be added to the training data, or when applying the trained multi-class semantic segmentation model, the ultrasound images with sudden changes in the measured length of the germ diameter table can be discarded or suppressed.

[0080] In addition, the early pregnancy ultrasound image may also include a germ with lower limbs, which may cause the positioning to the lower limbs when measuring the length of the germ diameter table, resulting in a larger measurement result. Therefore, ultrasound images containing a germ with lower limbs can be added to the training data of the multi-class semantic segmentation model, and the lower limbs of the germ can be removed when annotating the ultrasound image. Or when applying the trained multi-class semantic segmentation model, the lower limbs of the germ in the segmentation result can be manually removed.

[0081] When applying the ultrasound image processing method provided in this application to the segmentation and measurement of early pregnancy ultrasound images, after multiple experimental verifications, the precision coefficients of the multi-class semantic segmentation model for segmenting the gestational sac, yolk sac, and germ reach 0.972, 0.908, and 0.826 respectively, and the average segmentation precision coefficient reaches 0.902. And compared with the existing models, the multi-class semantic segmentation model of this application has lighter parameters, and the number of parameters is 9.03* , and the speed can reach 12 milliseconds per frame, which can effectively improve the segmentation precision and segmentation efficiency of the three anatomical structures of the gestational sac, yolk sac, and germ.

[0082] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0083] Based on the same inventive concept, an embodiment of the present application also provides an ultrasonic image processing device for implementing the ultrasonic image processing method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the ultrasonic image processing device provided below can refer to the limitations on the ultrasonic image processing method in the foregoing, and will not be repeated here.

[0084] In one embodiment, as Figure 8 shown, an ultrasonic image processing device 800 is provided, including: a data acquisition module 810, a semantic segmentation module 820, and a data measurement module 830, where:

[0085] The data acquisition module 810 is configured to acquire a set of ultrasonic images, and each ultrasonic image in the set of ultrasonic images includes multiple categories of recognition targets.

[0086] The semantic segmentation module 820 is configured to perform multi-category semantic segmentation on each ultrasonic image in the set of ultrasonic images at different preset segmentation scales to obtain a set of segmentation results of the recognition targets of each category.

[0087] The data measurement module 830 is configured to measure the diameter table length of the recognition target of each segmentation result in the set of segmentation results for each category of recognition target, and determine the maximum measured diameter table length as the target diameter table length of the recognition target of the category.

[0088] In one embodiment, the semantic segmentation module 820 is further configured to use the set of ultrasonic images as input, call a trained multi-category semantic segmentation model, and obtain a set of segmentation results of the recognition targets of each category. The trained multi-category semantic segmentation model includes multiple independent output channels, and each independent output channel is used to output a set of segmentation results of the recognition targets of one category. The trained multi-category semantic segmentation model is trained based on a set of historical ultrasonic images with category labels.

[0089] In one embodiment, the semantic segmentation module 820 is further configured to perform multi-class semantic segmentation on each ultrasound image in the ultrasound image set at different preset segmentation scales, obtain a set of segmentation results corresponding to different preset segmentation scales, perform an upsampling operation on different sets of segmentation results to obtain a set of segmentation results corresponding to different preset segmentation scales at the original size, where the original size is the size of the ultrasound images in the ultrasound image set, and average the set of segmentation results corresponding to different preset segmentation scales at the original size to obtain a set of segmentation results of the recognition targets of each category.

[0090] In one embodiment, the trained multi-class semantic segmentation model includes a multi-layer perceptron. The semantic segmentation module 820 is further configured to, for each ultrasound image, perform feature extraction and pooling operations on the ultrasound image to obtain feature vectors of the recognition targets of different categories at different preset segmentation scales, splice the feature vectors of the recognition targets of the same category at different preset segmentation scales to obtain a vector splicing result of the recognition targets of each category, use the vector splicing result of the recognition targets of each category as an input to call the multi-layer perceptron to obtain the category existence probabilities of the recognition targets of different categories in each ultrasound image, and if the category probability of the recognition target is less than a preset category probability threshold, filter out the segmentation result of the recognition target.

[0091] In one embodiment, the data measurement module 830 is further configured to, for each segmentation result, extract the contour pixel points of the recognition target in the segmentation result, determine the straight-line distances of the pixel points in the contour pixel points, and determine the diameter table length of the recognition target as the longest straight-line distance.

[0092] Each module in the above ultrasound image processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form so that the processor can call and execute the operations corresponding to the above modules.

[0093] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as an ultrasound image set. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an ultrasound image processing method.

[0094] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0095] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above-mentioned ultrasound image processing embodiment.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above-mentioned ultrasound image processing embodiment.

[0097] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned ultrasound image processing embodiment.

[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0099] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0100] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0101] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An ultrasonic image processing method, characterized in that, The method includes: Obtaining an ultrasonic image set, where each ultrasonic image in the ultrasonic image set contains multiple categories of recognition targets; Performing multi-category semantic segmentation on each ultrasonic image in the ultrasonic image set at different preset segmentation scales to obtain a set of segmentation results of the recognition targets of each category; For the set of segmentation results of the recognition targets of each category, measuring the diameter length of the recognition target of each segmentation result in the set of segmentation results, and determining the target diameter length of the recognition target of the category as the maximum measured diameter length.

2. The method according to claim 1, wherein The performing multi-category semantic segmentation on each ultrasonic image in the ultrasonic image set at different preset segmentation scales to obtain a set of segmentation results of the recognition targets of each category includes: Taking the ultrasonic image set as the input, calling a trained multi-category semantic segmentation model to obtain a set of segmentation results of the recognition targets of each category; Among them, the trained multi-category semantic segmentation model includes multiple independent output channels, and each independent output channel is used to output a set of segmentation results of the recognition targets of one category. The trained multi-category semantic segmentation model is trained based on a historical ultrasonic image set carrying category labels.

3. The method according to claim 2, characterized in that When the trained multi-category semantic segmentation model is called, the following steps are performed: Performing multi-category semantic segmentation on each ultrasonic image in the ultrasonic image set at different preset segmentation scales to obtain a set of segmentation results corresponding to different preset segmentation scales; Performing an upsampling operation on the different sets of segmentation results to obtain a set of segmentation results corresponding to different preset segmentation scales in the original size, where the original size is the size of the ultrasonic images in the ultrasonic image set; Averaging the sets of segmentation results corresponding to different preset segmentation scales in the original size to obtain a set of segmentation results of the recognition targets of each category.

4. The method according to claim 3, wherein The trained multi-category semantic segmentation model includes a multi-layer perceptron. When the trained multi-category semantic segmentation model is called, the following steps are also performed: For each ultrasonic image, performing feature extraction and pooling operations on the ultrasonic image to obtain feature vectors of the recognition targets of different categories at different preset segmentation scales, and splicing the feature vectors of the recognition targets of the same category at different preset segmentation scales to obtain a vector splicing result of the recognition targets of each category; Taking the vector splicing result of the recognition targets of each category as the input, calling the multi-layer perceptron to obtain the existence probabilities of the recognition targets of different categories in each ultrasonic image; If the category probability of the recognition target is less than a preset category probability threshold, then filtering out the segmentation result of the recognition target.

5. The method according to any one of claims 1 to 4, characterized in that, The measuring the diameter length of the recognition target of each segmentation result in the set of segmentation results includes: For each segmentation result, extracting the contour pixel points of the recognition target in the segmentation result; Determining the straight-line distances of the pixel points in the contour pixel points, and determining the longest straight-line distance as the diameter length of the recognition target.

6. An ultrasonic image processing device, characterized in that, The device includes: A data acquisition module for obtaining an ultrasonic image set, where each ultrasonic image in the ultrasonic image set contains multiple categories of recognition targets; A semantic segmentation module, configured to perform multi-class semantic segmentation on each ultrasound image in the ultrasound image set at different preset segmentation scales, so as to obtain a segmentation result set of recognition targets of each category; A data measurement module, configured to measure the diameter length of the recognition target of each segmentation result in the segmentation result set for the segmentation result set of the recognition target of each category, and determine the maximum measured diameter length as the target diameter length of the recognition target of the category.

7. The device according to claim 6, characterized in that, The semantic segmentation module is further configured to use the ultrasound image set as an input, call a trained multi-class semantic segmentation model, and obtain a segmentation result set of recognition targets of each category, where the trained multi-class semantic segmentation model includes a plurality of independent output channels, and each independent output channel is configured to output a segmentation result set of recognition targets of one category, and the trained multi-class semantic segmentation model is trained based on a historical ultrasound image set carrying category labels.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Adaptive multi-scale respiration monitoring method based on camera

    CN113256648A

  • Real-time semantic segmentation method based on multi-scale segmentation fusion

    CN113392840A

  • Method, system and device for automatically measuring gestational sac and germ and medium

    CN116392166A

  • Bird image recognition method and system for preventing misjudgment

    CN116740758A

  • Image segmentation method and device, equipment and storage medium

    CN117115900A