A fetal heart disease ultrasound image quality detection method based on multi-task learning

By employing a multi-task learning approach, combined with convolutional neural network feature sharing and attention mechanisms, a fetal cardiac disease ultrasound image quality detection model was constructed. This model addresses the issue of fetal ultrasound image analysis relying on human experience, achieving efficient image quality assessment and detection of key anatomical structures, reducing the burden on doctors and improving detection efficiency.

CN117274667BActive Publication Date: 2025-10-24HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311112939.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-10-24
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

In existing technologies, the interpretation of fetal ultrasound images mainly relies on the clinical experience of ultrasound doctors, which has the problems of strong subjectivity, time-consuming and labor-intensive, and consuming expert resources. This is especially prominent when novice doctors are making assessments. Furthermore, existing neural network models are difficult to achieve the flexibility and convenience of multi-task detection.

Method used

A multi-task learning approach is adopted, which utilizes the feature sharing property of convolutional neural networks. An attention mechanism is combined with the YOLOv5 multi-scale single-stage detection network to construct a shared underlying network. This enables an end-to-end multi-task learning model for fetal ultrasound image quality assessment and detection of key anatomical structures in fetal cardiac ultrasound images. Image detection is completed through a shared underlying network, a key anatomical structure detector, and an ultrasound section classifier.

Benefits of technology

It enables accurate ultrasound planar quality assessment and target detection in a very short time, reducing the workload of clinical ultrasound physicians, improving ultrasound image quality and detection efficiency, and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274667B_ABST
    Figure CN117274667B_ABST
Patent Text Reader

Abstract

The application provides a fetal heart disease ultrasonic image quality detection method based on multi-task learning, which comprises the following steps: constructing a multi-task learning network model; inputting a fetal ultrasonic data set into the multi-task learning network model for detection, and performing feature extraction through a shared bottom layer network; sharing features of a feature map after the feature extraction through a shared network characteristic, and inputting the shared feature map into a key anatomical structure detector and an ultrasonic section classifier respectively to perform corresponding task detection; and analyzing the output image through an ultrasonic image quality analysis system interface to complete auxiliary ultrasonic image detection. The application realizes the fetal heart disease ultrasonic image quality control task of single model completing fetal ultrasonic image quality evaluation and fetal heart ultrasonic image key anatomical structure detection through a multi-task learning model, and can accurately complete ultrasonic plane quality evaluation and ultrasonic image target detection in a very short time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical ultrasound image quality detection, and particularly relates to a fetal heart disease ultrasound image quality detection method based on multi-task learning. BACKGROUND

[0002] The working pressure of ultrasound image doctors is increasing. With the development and progress of science and technology, the level of medical service is improving, various medical services and protection are improving, and health knowledge is popularizing. People's cognition of physical and mental health is also from shallow to deep. Therefore, more and more people begin to pay attention to physical and mental health, especially pregnant women. More and more pregnant women perform prenatal fetal health detection and evaluation, which brings unprecedented working pressure to hospitals and doctors. At present, prenatal examination mainly depends on ultrasound detection, and the analysis and health evaluation of ultrasound images are mainly completed by doctors with rich clinical ultrasound image analysis experience and rich human anatomy professional knowledge. This not only occupies a large number of talent resources, but also reduces the work efficiency.

[0003] Prenatal ultrasound examination is a powerful tool for preventing birth defects and evaluating fetal health. Benefiting from the advantages of non-invasiveness, real-time and low cost, ultrasound examination has always been the preferred imaging mode for prenatal fetal examination. During prenatal examination, the fetus is scanned by an ultrasound device to obtain fetal ultrasound images. Ultrasound doctors can predict and evaluate the development of the fetus through the organ position, organ shape and other images of the image. However, the standard section mainly depends on the clinical experience of ultrasound image doctors at present, which not only has great subjectivity, but also easily obtains poor ultrasound sections. In addition, in clinical practice, due to the rich expert ultrasound doctors, the quality of ultrasound images of novice doctors is evaluated, which not only consumes time and effort, but also occupies clinical expert resources.

[0004] At present, the ultrasound image analysis process mainly depends on the rich clinical experience and rich human anatomy knowledge of ultrasound doctors to manually screen standard section ultrasound images, and then evaluate the health degree of fetal development. Accurate evaluation is based on accurate standard ultrasound images, which is very challenging for novice doctors. The results of manual analysis and evaluation of ultrasound images not only have strong subjective factors, but also are very time-consuming and labor-intensive. How to reduce the working pressure of ultrasound doctors and release the occupation of ultrasound doctor resources, and improve the quality of ultrasound sections has become a real demand for clinical examination.

[0005] Although Artificial Neural Network (ANN) was proposed as early as the 1940s, researchers have also proposed various innovative network algorithms and solved various neural network problems, but due to the limitations of computer computing power and insufficient related data, etc., the neural network algorithm has failed to achieve a breakthrough in many application fields or scenarios, especially in the medical field. In recent years, with the breakthroughs in computing power and technology, artificial intelligence algorithm depth optimization, network big data development and other fields, artificial neural networks have once again been pushed into a development boom. At the same time, with the rapid development of the country, people's life is becoming more and more prosperous, and more and more people begin to pay attention to prenatal fetal examination, so clinical prenatal ultrasound examination has accumulated a large amount of data, thus promoting the development and application of artificial intelligence in prenatal ultrasound auxiliary diagnosis.

[0006] With the support of image processing, machine learning, attention mechanism, deep learning and other technologies, scholars have proposed many neural network intelligent auxiliary diagnosis systems applied in medical ultrasound images. For example, the prior art proposes an end-to-end fetal head ultrasound standard plane detection (USPD) model based on multi-task learning and hybrid knowledge graph, which reduces the missed detection rate and false detection rate, improves the reliability and interpretability of end-to-end automatic detection, and improves the detection accuracy of relatively small and morphologically variable anatomical structures; the prior art proposes a new multi-task faster regional convolutional neural network (MFR-CNN) framework to realize automatic quality assessment, and proposes a transfer learning segmentation training method, and the network model adds a clinical prior knowledge module to improve the accuracy of the detection result. In the shared residual feature layer, a multi-scale channel and spatial hybrid attention mechanism module is added, which enables the network to focus on the details of the target during network training, while improving the feature expression ability of the network model; the prior art proposes to combine the feature pyramid network (FPN), MobileNet and UNet (MobileUNet-FPN) to segment 13 key cardiac structures, and the experimental results show that the proposed model has excellent performance on fetal A4C and femur length images; the prior art proposes an automatic fetal ultrasound standard plane recognition (FUSPR) model based on deep learning in an industrial internet of things (IIoT) environment, wherein the FUSPR model is composed of a convolutional neural network (CNN) component and a recurrent neural network (RNN) component, which respectively learn the spatial and temporal features of the ultrasound video stream. The results of a large number of experiments on more than 1000 ultrasound videos show that the FUSPR model is superior to the competitive baseline in terms of accuracy and performance; the prior art proposes a hybrid classification framework for cardiac cycle detection that performs detection tasks using a hybrid classification framework, and the experimental data show that the average classification result accuracy of the model reaches 94.84%, and the average detection error for some structures and some frames is 1.25 and 0.80 frames, respectively; the prior art proposes an abnormal threshold method based on multiple classifiers, and uses the XGBoost algorithm as a sub-classifier to process massive ECG data to realize heart disease classification problems, and the experimental results show that this method can effectively improve the classification accuracy. As can be seen, many experiments show that artificial neural networks have great potential and feasibility in medical ultrasound images for image quality detection and obtaining key target anatomical structures.

[0007] At present, since most artificial intelligence neural networks still belong to a single task network that can only complete one or a class of tasks, when multiple tasks such as image quality assessment and target detection of a picture need to be completed at the same time, two network models respectively for image quality assessment and target detection need to be used to detect a picture, which not only consumes time and effort, but also requires the computer to calculate the same picture twice, greatly reducing the flexibility and convenience of artificial intelligence neural networks.

[0008] Therefore, and inspired by the above experimental results and single task neural networks, the present application proposes to utilize the feature sharing characteristics of convolutional neural networks to realize bottom layer feature sharing in the YoloV5 multi-scale single stage detection network combined with attention mechanism, and to realize an end-to-end multi-task learning model (GBI-YoloV5) of a single model for fetal heart disease ultrasound image quality control task of fetal ultrasound image quality assessment and fetal heart ultrasound image key anatomical structure detection. SUMMARY

[0009] In order to overcome the shortcomings of the prior art, the present application provides a fetal heart disease ultrasound image quality detection method based on multi-task learning, which utilizes the shared feature principle of multiple single task learning models and convolutional neural networks, proposes and realizes an end-to-end multi-task network model of a single model, utilizes the feature sharing characteristics of convolutional neural networks to realize bottom layer feature sharing in the YoloV5 multi-scale single stage detection network combined with attention mechanism, and realizes a single model for fetal heart disease ultrasound image quality control task of fetal ultrasound image quality assessment and fetal heart ultrasound image key anatomical structure detection through a multi-task learning model, which can accurately complete ultrasound plane quality assessment and ultrasound image target detection in a very short time.

[0010] The application provides a fetal heart disease ultrasound image quality detection method based on multi-task learning, which comprises the following steps: constructing a multi-task learning network model; wherein the multi-task learning network model comprises a shared bottom layer network, a key anatomical structure detector and an ultrasound section classifier; fetal ultrasound data set is input into the multi-task learning network model for detection, and feature extraction is performed through the shared bottom layer network; the feature map after completing the feature extraction is shared through the shared network characteristics, and the shared feature map is input into the key anatomical structure detector and the ultrasound section classifier respectively, and the corresponding task detection is performed; after the key anatomical structure detector outputs the detected and framed target area image, the quality and type of the image are output in the ultrasound section classifier, and the output image is analyzed by combining the two outputs; the output image and the analysis structure result are analyzed through the ultrasound image quality analysis system interface to complete the auxiliary ultrasound image detection.

[0011] Further, the shared bottom layer network comprises: an input layer, the image input into the input layer is the data set image after the original image is preprocessed through mosaic enhancement and image self-adaption; a convolution layer structure, the shared bottom layer network is composed of five convolution layers with a convolution kernel of 3 and one convolution layer with a convolution kernel of 6; a residual network structure, the entire shared bottom layer network is composed of four residual structure blocks with a convolution kernel of 1 and a step of 1, wherein the main part of the residual structure block is composed of N residual blocks without residual structure which are nested continuously.

[0012] An attention mechanism block is introduced before the fifth layer convolution layer of the YoloV5 main network, which is a GAM channel and spatial mixed attention mechanism, used for enhancing feature extraction, and outputting the shared feature map through the fifth convolution layer and the residual layer; wherein the activation function of each layer of the shared bottom layer network is a SiLU activation function.

[0013] Further, the attention mechanism block is a GAM mixed attention mechanism, which integrates a multi-layer perception module. When the input feature map is output after the GAM channel, it is multiplied with the original feature map according to the feature points, and the obtained feature map is consistent in size with the original feature map, and the obtained feature map is used as the input of the GAM mixed attention mechanism. In the GAM mixed attention mechanism, a convolution layer with a convolution kernel of 7 is used to increase and decrease the dimension of the feature map, that is, the original convolution is replaced by dilated convolution to improve the convolution speed.

[0014] Further, after generating the shared bottom network, multi-scale features are combined with SPPF, FPN and PAN, and a key anatomical structure detector is generated using 9 groups of anchor frames provided by YoloV5; wherein, there is a transition layer of spatial pyramid pooling (SPPF) between the shared bottom network and the key anatomical structure detector, which is used to fix the size of the shared feature layer and enhance feature extraction, providing rich feature layers for subsequent FPN.

[0015] Further, the FPN includes convolutional layers, residual layers, up-sampling layers and feature fusion layers, which adjust the channel number and size of the feature layer from the SPPF layer using convolutional layers and up-sampling layers, fuse different scale features from the shared network, and then extract the fused feature layer using residual layers, thereby realizing the construction of FPN; wherein, three different scale feature layers are still output in the FPN, the largest scale feature layer of the FPN is used as the input feature of the PAN and the output of the target detection network, and the medium and small scale feature layers are used as the input feature layers of the PAN; the PAN includes convolutional layers, residual layers and feature fusion layers, which adjust the channel number of the largest scale feature from the SPPF layer using convolutional layers, fuse it with the features from other scales of the PAN, and extract the fused features through residual layers, thereby constructing the PAN; wherein, two different scale feature layers are output in the PAN structure for target detection.

[0016] Further, the multi-task learning network model is trained, including: step S11, initializing the shared bottom network of GBI-YoloV5 using the excellent model parameters of YoloV5 network open source, freezing the backbone network parameters, and training the key anatomical region detection module separately; step S12, initializing the shared bottom network of the network using the network model trained in step S11, continuing to train the key anatomical structure detector separately, and this time the shared bottom network is not frozen; step S13, initializing the shared bottom network using the network model trained in step S12, freezing the shared bottom network and the key anatomical structure detector, and training the ultrasound section classifier separately; step S14, the model trained in step S13 is the model training result of the multi-task learning network model, and the entire network can be initialized using this result to analyze the fetal heart ultrasound image quality detection.

[0017] Further, after the shared feature map is input into the ultrasound section classifier, the ultrasound section classifier uses Conv Block residual network and Identity Block residual network to learn features and adjust the size of the features so as to be equal to other scale features, the adjusted feature map is added to the scale feature map according to the feature points for feature fusion, after the three scale features are fused, the feature map is flattened by using average pooling, and finally a fully connected layer is used to output the classification result.

[0018] Further, when the ultrasound section classifier performs the classification task, a Softmax cross-entropy loss function composed of a softmax and a cross-entropy loss function is used to replace the cross-entropy loss function, and for the softmax, it is represented as formula (3-2):

[0019] (3-2)

[0020] wherein j represents a class index, z is a discrete probability distribution result output by the fully connected layer, is the kth value output by the fully connected layer;

[0021] For the cross-entropy loss function, it is represented as formula (3-3):

[0022] (3-3)

[0023] wherein the output of formula (3-3) is the output in formula (3-2), is a true value of classification, and the output of formula (3-3) is the final loss of classification.

[0024] Further, in the key anatomical structure detector, the loss of the whole target detection includes a classification loss, a target confidence loss and a positioning loss, therefore, the loss function of the whole detector is defined as formula (3-4):

[0025] (3-4)

[0026] wherein N is the number of detection layers, B is the number of targets of an anchor frame to which a certain class label is assigned, SxS is the number of grids divided by the current scale, L box , L obj and L cls represent a target confidence loss, a confidence loss and a classification loss respectively, , and represent the weights of the above three losses respectively;

[0027] For target confidence loss, CIOU rectangular box loss is mainly used, and the calculation principle is expressed as formula (3-5):

[0028] (3-5)

[0029] The confidence loss and classification loss are both solved using the binary cross entropy loss function. The calculation principle of this loss function is expressed as formula (3-6):

[0030] (3-6)

[0031] Among them, n is the number of detection target categories, The label for the target classification, for The probabilities of the labels, b and b gt Represent the anchor box and the real box during detection, W, H, w gt and h gt Represent the width and height of the anchor box and the real box respectively, ρ is the distance between the center points of the anchor box and the real box, d and c are the farthest distances between the boundaries of the two boxes, and α is the weight coefficient.

[0032] A further solution is to build an ultrasound image quality analysis system interface, including: using the Anaconda software source to install the corresponding virtual environment required by PyQt5 for the Python environment, and configuring the corresponding PyQt5 components through the PyCharm software. The components include QT Design and Pyuic components, which are PyQt5 development components; selecting the appropriate window for window component layout by running the QT Design component, saving the design as a ui file after QT Design completes the design, and using the Pyuic component to directly convert the ui into a Python script. The script is stored in the form of a class, so Python can call this script using the class method, and use the Python programming language to complete distributed parallel processing, various control functions, system initialization and other functions, thereby realizing PyQt5 design visualization system design.

[0033] Therefore, compared with the prior art, the application proposes to construct a shared bottom network by using a bottom shared feature principle, and introduce a GAM attention mechanism in the shared bottom network to complete the construction of the shared network for complex learning tasks; an image ultrasound section classifier is realized by combining multi-scale convolution fusion; a key anatomical structure detector is constructed by using FPN and PAN, and the ultrasound section classifier and the key anatomical structure detector are combined with the shared bottom network to realize a multi-task learning network for fetal heart disease ultrasound image quality detection, so as to reduce the work pressure of clinical ultrasound doctors and improve the ultrasound image quality. Moreover, the application uses PyQt5 to program an ultrasound image detection visualization system, and the ultrasound image detection can be realized by controlling a mouse, which greatly improves the use efficiency, applicable scene and reduces the use difficulty.

[0034] In addition, the multi-task learning model proposed by the application is an end-to-end multi-task learning model realized based on a YoloV5 multi-scale single-stage single-task learning network, a ResNet-50 classification network and an attention mechanism, and has good target detection effect and classification effect. The multi-task model can realize automatic positioning and analysis of multiple fetal heart ultrasound image key anatomical structures, and classify the organ types to which the ultrasound images belong. In terms of detection time, the detection time of the entire multi-task model is not inferior to the detection time of the single-task YoloV5, but is superior to the time used for image quality detection and key anatomical region detection by simultaneously using the YoloV5 model for detection and the ResNet-50 model for classification; in terms of detection effect and classification effect, the key anatomical region detection and image quality detection effect of the entire multi-task learning model is obviously superior to the key anatomical region detection effect of the YoloV5 model and the image quality detection effect of the ResNet-50 model.

[0035] The application will be further described in detail below in combination with the drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 is a flowchart of an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning.

[0037] Figure 2 is a network structure diagram of a multi-task learning network model in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning.

[0038] Figure 3 is a structure diagram of a shared bottom network in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning.

[0039] Figure 4is a structural schematic diagram of a GAM mixed attention mechanism in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0040] Figure 5 is a structural schematic diagram of an ultrasound section classifier in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0041] Figure 6 is a structural schematic diagram of a key anatomical structure detector in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0042] Figure 7 is a first structural schematic diagram of training a multi-task learning network model in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0043] Figure 8 is a second structural schematic diagram of training a multi-task learning network model in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0044] Figure 9 is a schematic diagram of a PyQt5 component in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0045] Figure 10 is a schematic diagram of a QT Design component in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0046] Figure 11 is a schematic diagram of a multi-task automatic detection system running result in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application.

[0047] Figure 12 is a structural schematic diagram of an ultrasound section classifier prediction result confusion matrix in an embodiment of a fetal heart disease ultrasound image quality detection method based on multi-task learning of the present application. DETAILED DESCRIPTION

[0048] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0049] Referring toFigure 1 and Figure 2 The application provides a fetal heart disease ultrasound image quality detection method based on multi-task learning, which comprises the following steps:

[0050] Step S1, a multi-task learning network model is constructed; wherein the multi-task learning network model comprises a shared bottom layer network, a key anatomical structure detector and an ultrasound section classifier; Step S2, a fetal ultrasound data set is input into the multi-task learning network model for detection, and feature extraction is performed through the shared bottom layer network; Step S3, the feature map after completing the feature extraction is shared through the shared network characteristics, and the shared feature map is input into the key anatomical structure detector and the ultrasound section classifier, respectively, for detection of corresponding tasks; Step S4, after the key anatomical structure detector outputs the detected and framed target region image, the quality and type of the image are output in the ultrasound section classifier, and the two outputs are combined to analyze the output image; Step S5, the output image and the analysis structure result are analyzed through the ultrasound image quality analysis system interface to complete the auxiliary ultrasound image detection.

[0051] In the embodiment, the shared bottom layer network is the key to convert the single-task learning model into the multi-task learning model, and the shared bottom layer network in the embodiment is improved from the backbone network of the YoloV5 multi-scale and single-stage fusion model, mainly composed of a convolution layer, a pooling layer and a residual structure, and comprising four parts.

[0052] The input layer of the image is the data set image after the original image is subjected to mosaic enhancement and image self-adaptive preprocessing; the convolution layer structure is composed of five convolution layers with a convolution kernel of 3 and one convolution layer with a convolution kernel of 6, and the step is 2; the residual network structure is composed of four residual structure blocks with a convolution kernel of 1 and a step of 1, wherein the main part of the residual structure block is composed of N residual blocks without residual structure; the attention mechanism block is a GAM channel and spatial mixed attention mechanism introduced before the fifth convolution layer of the YoloV5 backbone network, which is used to enhance feature extraction, and outputs the shared feature map through the fifth convolution layer and the residual layer. Figure 3 as shown in the shared bottom layer network structure diagram.

[0053] The attention mechanism block is a GAM mixed attention mechanism, and the attention mechanism structure is as shown in Figure 4The input feature map is subjected to a channel attention, which aims to preserve information and enhance features from the channels of the feature map. The GAM hybrid attention mechanism incorporates a multi-layer perceptron module. After the input feature map is output by the GAM channel, it is multiplied by the original feature map by feature points. The resulting feature map is consistent in size with the original feature map, and the resulting feature map is used as the input to the GAM hybrid attention mechanism. In the GAM hybrid attention mechanism, a convolution layer with a kernel of 7 is used to increase and decrease the dimension of the feature map. In fact, dilated convolution is used instead of the original convolution to improve the speed of convolution. The spatial attention mechanism aims to focus on information in space. Two convolution layers are used for feature fusion. To reduce the negative impact of the max pooling module on the feature map, the pooling layer is removed in the spatial attention of the GAM, further preserving the feature map. The feature map output by the spatial attention mechanism is multiplied by the input spatial attention mechanism information by feature points as the final output. The size of the output feature map is consistent with the size of the input feature map, making the attention mechanism compatible with many deep learning models. The activation function of each layer of the entire bottom shared network is the SiLU activation function, and the overall architecture of the bottom network is similar to the YoloV5 backbone network. Although the image is subjected to five convolution layers and four residual modules to obtain a rich feature layer, the feature distribution of the feature layer is scattered and not prominent enough, leading to slow convergence or non-convergence during subsequent learning. Therefore, a GAM hybrid attention mechanism is introduced before the fifth CBS convolution layer in the middle of the shared bottom network. The channel attention mechanism adjusts the information weight of each channel of the feature layer based on different tasks, thereby focusing on learning features with high weights, highlighting and gathering scattered features in the feature layer, and adjusting the feature direction for subsequent task-specific learning of the model, thereby improving the learning efficiency and learning effect of the model.

[0054] The image data is subjected to the entire bottom shared network to output three different scales of shared bottom features, corresponding to the large-scale feature layer in yellow, the medium-scale feature layer in blue, and the small-scale feature layer in purple in the figure. These features can be used for different task-specific learning of each module in the subsequent model, providing the possibility of independent training of each module.

[0055] For the image classification module and the target detection module, in the classification module, compared with the input image size 224*224*3 or 227*227*3 of the traditional CNN classification model, the input image size of the network is relatively large (640*640*3), and the feature layer size output by the shared bottom network is still large. Inspired by the ResNet-50 model and the YoloV5 model, in order to fully utilize the feature information of the feature layer, three scale feature layers from the shared network are fused by using the residual structure to further extract and learn the shared features. When the feature layer is 5*5, it is no longer extracted, and average pooling and full connection are performed to output the image classification result of the image classification module.

[0056] As shown in Figure 5 , this structure is an ultrasound section classifier of the GBI-YoloV5 network. This structure not only uses two residual structure feature extraction, but also fuses shared bottom features of other sizes, not only deepens the model network, but also enriches the feature map of the classifier, increases the classification effect of the model, and speeds up the convergence speed of the classifier. As shown in Figure 6 , the Figure 6 is a schematic diagram of a fetal heart ultrasound key structure detection module. In the model target classification module, inspired by YoloV5, in order to obtain the accurate position of the fetal heart anatomical structure and the key anatomical structure, after the shared bottom network, the multi-scale features are further fused by combining SPPF, FPN and PAN, and 9 groups of anchor boxes provided by YoloV5 are used to generate more accurate key anatomical regions.

[0057] After generating the shared bottom network, the multi-scale features are further fused by combining SPPF, FPN and PAN, and the key anatomical structure detector is generated by using 9 groups of anchor boxes provided by YoloV5; wherein, there is a transition layer of spatial pyramid pooling layer (SPPF) between the shared bottom network and the key anatomical structure detector, and the purpose is to fix the size of the shared feature layer and enhance the feature extraction, so as to provide rich feature layers for the subsequent FPN.

[0058] The FPN comprises a convolutional layer, a residual layer, an up-sampling layer and a feature fusion layer, the convolutional layer and the up-sampling layer are used to adjust the channel number and size of the feature layer from the SPPF layer respectively, the different scale features from the shared network are fused, and the residual layer is used for feature extraction of the fused feature layer, so as to realize the construction of the FPN; wherein, three feature layers of different scales are still output in the FPN, the feature layer of the largest scale of the FPN is used as the input feature of the PAN and the output of the target detection network, and the medium and small scale feature layers are used as the input feature layers of the PAN; the PAN comprises a convolutional layer, a residual layer and a feature fusion layer, the convolutional layer is used to adjust the channel number of the largest scale feature from the SPPF layer, the largest scale feature is fused with the features of other scales from the PAN, the residual layer is used for extracting the fused feature, so as to construct the PAN; wherein, two feature layers of different scales are output in the PAN structure for target detection.

[0059] The GBI-YoloV5 network model of the embodiment adopts a transfer learning segmented alternating joint optimization training mode. Since the convergence speed of the key region detection module is faster than that of the image type classification module in the entire multi-task network, in order to prevent the key anatomical structure detection module and the ultrasound section classification module from affecting each other during training, and to achieve better performance of network training, specifically, the multi-task learning network model is trained by using a transfer learning mode, including:

[0060] Step S11, initializing the shared bottom network of the GBI-YoloV5 using the excellent model parameters of the YoloV5 network open source, freezing the backbone network parameters, and training the anatomical key region detection module alone; step S12, initializing the shared bottom network of the network using the network model trained in step S11, continuing to train the anatomical key structure detector alone, and this time the shared bottom network is not frozen; step S13, initializing the shared bottom network using the network model trained in step S12, freezing the shared bottom network and the key anatomical structure detector, and training the ultrasound section classifier alone; step S14, the model trained in step S13 is the model training result of the multi-task learning network model, and the entire network can be initialized using this result to analyze the fetal heart ultrasound image quality detection.

[0061] The multi-task learning network model GBI-YoloV5 designed in the embodiment has the structure as shown in Figure 7The Residual block module is an ultrasound section classifier designed in this embodiment. The a, b, and c three-layer feature maps shared from the shared bottom network are input into the Residual block classifier. After the shared feature maps are input into the ultrasound section classifier, the ultrasound section classifier uses the Conv Block residual network and the Identity Block residual network to learn the features and adjust the size of the features so as to be equal to other scale feature sizes. The adjusted feature maps and the size are added to the scale feature maps according to the feature points for feature fusion. After the three scale features are fused, the average pooling is used to flatten the feature maps. Finally, a fully connected layer is used to output the four classification results of this embodiment.

[0062] Two ultrasound section classifiers are designed in this embodiment. The ultrasound section classifier in the GBI-YoloV5 multi-task network not only uses the residual network structure, but also combines the multi-scale feature fusion theory. The residual structure in the ultrasound section classifier is more abundant, deeper, and has stronger network expression ability than the ultrasound section classifier structure in the Our-1 multi-task network. The Our-1 multi-task network model only uses the feature map output by the last layer of the shared network as input data, and directly outputs four ultrasound section classification results after the ultrasound section classifier.

[0063] The structure of another end-to-end multi-task learning network model Our-1 designed in this embodiment is shown in Figure 8 The biggest difference between the Our-1 multi-task network model and the GBI-YoloV5 multi-task network model is the structure and depth of the ultrasound section classifier. The rest of the shared bottom network and the key anatomical structure detector are basically the same. The input of the ultrasound section classifier of the Our-1 multi-task network model is only the feature map of the smallest scale of the shared bottom network. Therefore, the feature richness is not as rich as the input feature of the ultrasound section classifier of the GBI-YoloV5 network.

[0064] In the classification module, three different scale feature map data from the shared network are fused by using the residual structure to further extract and learn the shared features. Two residual structures are used in this embodiment to extract features. After multi-layer residual feature fusion, average pooling and full connection operations are performed on the feature maps, and finally four ultrasound section types such as non-standard section, standard apical four-chamber heart section, standard left ventricular outflow tract section, and standard right ventricular outflow tract section are obtained. The classifier of this embodiment combines the ideas of residual network and multi-branch feature fusion, not only deepens the depth of the model, but also enriches the feature extraction of the classifier, increases the classification effect of the model, and speeds up the convergence time of the model classifier.

[0065] The ultrasound image dataset of the embodiment is mainly derived from the extraction of fetal heart ultrasound video, therefore, the standard ultrasound section sample data is obviously more than the non-standard ultrasound plane sample, and the standard ultrasound section and the non-standard section of the embodiment have strong similarity, which significantly increases the difficulty of classification. Since the Softmax cross-entropy loss calculation considers the influence of the classification category in the whole calculation process, it is equivalent to a classifier calculating the probability of multiple classification categories, which has higher robustness and stability for the classification task of the embodiment. When the ultrasound section classifier performs the classification task, the Softmax cross-entropy loss function composed of softmax and cross-entropy loss function is used to replace the cross-entropy loss function, for softmax, it is represented as formula (3-2):

[0066] (3-2)

[0067] wherein j represents the category index, z is the discrete probability distribution result of the full connection layer output, is the kth value of the full connection layer output;

[0068] For the cross-entropy loss function, it is represented as formula (3-3):

[0069] (3-3)

[0070] wherein the of formula (3-3) is the output of formula (3-2), is the true value of the classification, and the output of formula (3-3) is the final loss of the classification.

[0071] In the key anatomical structure detector, the fetal heart ultrasound image tissue structure detection task of the embodiment is mainly improved on the YoloV5 target detection model, therefore, the loss composition of the whole target detection includes classification loss, target confidence loss and positioning loss, therefore, the overall loss function definition of the detector is formula (3-4):

[0072] (3-4)

[0073] wherein N is the number of detection layers, B is the target number of anchor boxes allocated to a certain label, SxS is the number of grids divided by the current scale, L box , L obj and L cls represent target confidence loss, confidence loss and classification loss respectively, , and They represent the weights of the above three losses respectively. According to experience, in this embodiment, 0.4, 0.3 and 0.3 are usually taken as the weight values ​​of the three losses respectively.

[0074] For target confidence loss, CIOU rectangular box loss is mainly used, and the calculation principle is expressed as formula (3-5):

[0075] (3-5)

[0076] The confidence loss and classification loss are both solved using the binary cross entropy loss function. The calculation principle of this loss function is expressed as formula (3-6):

[0077] (3-6)

[0078] Among them, n is the number of detection target categories, The label for the target classification, for The probabilities of the labels, b and b gt Represent the anchor box and the real box during detection, W, H, w gt and h gt Represent the width and height of the anchor box and the real box respectively, ρ is the distance between the center points of the anchor box and the real box, d and c are the farthest distances between the boundaries of the two boxes, and α is the weight coefficient.

[0079] Therefore, the multi-task learning model of this embodiment is mainly composed of classification tasks and detection tasks. Therefore, the total loss of the multi-task learning model of this embodiment can be expressed as follows:

[0080] (3-7)

[0081] It can be seen that the multi-task learning network model design is mainly composed of a shared underlying network block, a key anatomical structure detection block, and an ultrasound section classification block. It makes full use of the underlying shared features to realize the independent training of the key anatomical structure detection module and the ultrasound section classification module.

[0082] In this embodiment, the ultrasound image quality analysis system interface is constructed, including:

[0083] PyQt5 is a cross-platform toolkit that integrates Python and Qt for creating graphical user interface (GUI) applications. It is compatible with both Python 2.x and Python 3.x. This example uses Python 3.8.15 for this design. PyQt5 consists of Python modules and boasts over 620 classes and over 6,000 functions and methods, significantly improving GUI development efficiency. The created GUIs can run on operating systems such as Linux, Windows, and macOS. This example is designed to run on Windows.

[0084] Use the Anaconda software source to install the corresponding virtual environment required by PyQt5 for the Python environment, and configure the corresponding PyQt5 components through the PyCharm software, which include QT Design and Pyuic components, such as Figure 9 The QT Design and Pyuic components in the PyQt5 component diagram are PyQt5 development components;

[0085] By running the QT Design component, select the appropriate window to layout the window components, such as Figure 10 As shown in the QTDesign component, after QT Design completes the design, it is saved as a ui file. The Pyuic component is used to convert the ui directly into a Python script. The script is stored in the form of a class, so Python can call this script using the class method, and use the Python programming language to complete distributed parallel processing, various control functions, system initialization and other functions, thereby realizing PyQt5 design visualization system design.

[0086] like Figure 11 As shown, Figure 11 The GUI for the multi-task automatic detection system shows a schematic diagram of the automatic detection system's operation results. Use buttons 1 or 2 to select input images or folder data. After completing the selection, the input image is displayed in control 3. Click control 4 to run image analysis. The analyzed image is displayed in palette 5, the analysis log is displayed in analysis console 10, and the run detection time is displayed in control 9. Controls 7 and 8 are used to view historical operation records. The analysis log, run time, and analyzed images are finally saved in the folder displayed in control 6. Each time the system is started, the last run result is read and displayed in the system.

[0087] The embodiment can also use PyQt5 to design a distributed multi-threaded task execution GUI interface. The system can save and read parsed information, select a data saving address, and the like, and call a trained multi-task model at the bottom of the interface to effectively improve the use effect of the system and enhance the usability of the system. Through the system, the fetal heart ultrasound image detection efficiency can be effectively improved. The fetal heart ultrasound image detection can be completed by only controlling the mouse. Moreover, the software can run and browse the access history system running result, greatly shortening the operation image detection time and tedious manual call multi-task model detection, and further improving the speed of the multi-task model detection auxiliary task.

[0088] The non-standard ultrasound section data set is relatively rare compared with the standard section image, and the non-standard section is very similar to the standard section. Therefore, the training and detection of the non-standard section need to be considered during model training, otherwise the trained model may not be able to detect the non-standard type of ultrasound section or the detection accuracy is not high, thereby causing the performance of the entire model to decrease. For this, the best way is to increase the number of non-standard sections, and when there is no condition for supplement, the data set should be enriched by fully utilizing image preprocessing.

[0089] It can be seen that the application is inspired by the ResNet-50 classification model, based on the YoloV5 target detection framework, using convolutional shared networks and attention mechanisms, a multi-task machine learning framework for fetal heart disease ultrasound image classification and anatomical structure key region target detection is designed, called GBI-YoloV5 model, for automatic quality evaluation of ultrasound images. The multi-task learning model designed in this paper can identify more than ten key anatomical structures such as left ventricle, right ventricle and right atrium, and at the same time complete ultrasound image quality evaluation. If the ultrasound image is a standard quality ultrasound image, then the ultrasound image classification detection is performed to determine the organ position area to which the image belongs. First, the backbone network of YoloV5 is used as the shared bottom network of the multi-task learning model of the embodiment, and the third, fourth and fifth shared feature maps in the bottom network are shared. The three scale feature maps are input into the model classifier and detector respectively for deep independent learning. Secondly, inspired by the ResNet-50 network and the YoloV5 network, the same residual structure as the ResNet-50 network is used in the model image classifier to fuse multi-scale shared features, and finally the image type is predicted through full connection to complete the classifier construction; in the model detector, the structure of the YoloV5 model is used to detect the key region. Finally, the image and the analysis structure result are analyzed by using the visual interface to complete the auxiliary ultrasound image analysis.

[0090] Specifically, the related network architecture involved in the multi-task learning model proposed in the embodiment is introduced in detail, including the theory of multi-task model in the model, convolutional neural network, residual network, attention mechanism, and spatial pyramid pooling (SPPF), etc. These modules are the most core components of the entire multi-task model and directly affect the performance of the multi-task model.

[0091] For the multi-task learning model, multi-task learning (Multi-task learning) and single-task learning (Single-task learning) are a relative concept. Multi-task machine learning, as the name implies, is a deep learning method that uses only one data set and one model network to achieve multiple different task requirements. It solves complex problems by decomposing complex multi-task into multiple simple and independent single tasks and then combining the results. Multi-task machine learning is a deep learning framework that induces feature sharing and transfer learning mechanism, belonging to the joint learning framework and multiple task parallel learning framework. Compared with using multiple single-task machine learning models to complete multiple machine learning tasks, multi-task machine learning not only reduces the cumbersome model training, but also reduces the large amount of calculation in the training process, greatly shortens the model training time and detection time, and expands the applicable scenarios.

[0092] Based on the outstanding advantages of the multi-task learning network, the embodiment designs a YoloV5 single-task neural network based on ResNet-50 image classifier, uses a shared bottom network to reasonably combine the two to design a multi-task machine learning network structure, and realizes fetal heart ultrasound image quality control and fetal heart ultrasound image detection and analysis with one data set and one model network.

[0093] For spatial pyramid pooling, there are SPP and SPPF structures. The embodiment uses the SPPF spatial pyramid pooling structure improved and upgraded based on the SPP structure. SPPF is improved based on SPP structure. The purpose of both is to control the input features to have fixed feature vectors. SPPF specifies the initial convolution kernel pooling after one convolution of the input features, and then the result of each pooling is used as the input of the next pooling layer. Finally, the results of each pooling and the results of the first convolution are fused, and then a convolution operation is performed to output. In short, SPPF only specifies the convolution kernel once, while SPP specifies the convolution kernel three times. Therefore, SPPF is faster than SPP in speed and has less computational complexity than SPP. Therefore, the SPPF used in the embodiment is used as the structure of spatial pyramid pooling.

[0094] For the multi-task network model GBI-YoloV5, multi-task learning can complete two or more tasks at a time, which is more time-saving and labor-saving than single-task learning. In this embodiment, the VGG16, ResNet-50 and YoloV5 networks are combined, and the convolutional neural network feature sharing feature is used to combine the two models reasonably to form an end-to-end network model for multi-task learning. The implementation principle is not described in detail here.

[0095] The YoloV5 neck of this embodiment is mainly composed of FPN and PAN, which realizes multi-scale feature fusion. The FPN structure is similar to an upsampling process, mainly composed of CBS structure, upsample, CPS structure and Concat. The small-scale feature map output in the backbone network first enters a CBS structure to adjust the channel number, and its output is used as the input of the upsample structure and also as the input of the PAN small-scale feature fusion; then, the output feature map is adjusted in size by the upsample structure to be consistent with the size of the medium-scale feature map output by the backbone network; then, the feature map after adjusting the size is fused with the medium-scale feature map output by the backbone network; then, the fused feature map passes through a CPS2_1 and a CBS structure, and the output feature map is used as the input of the PAN medium-scale feature fusion and also as the input of the upsample structure; then, the feature map after adjusting the size by the upsample is fused with the large-scale feature map output by the backbone network; finally, the fused feature map passes through a CPS2_1 structure to output a feature map, which is used as the input of the yolo Head and also as the input feature map of the PAN. Thus, the construction of the feature pyramid is realized.

[0096] The PAN structure is similar to a down-sampling process, mainly composed of CBS structure, CPS structure and Concat. First, the feature map output by the FAN is adjusted in size by the CBS structure; then, it is fused with the medium-scale feature output by the FAN; then, the fused feature map passes through a CPS2_1, which is used as the output feature map of the yolo Head and also as the input of the next layer of the PAN; finally, the output feature map is adjusted in size by the CBS and fused with the small-scale feature map output by the FAN, and after passing through a CPS2_1 structure, it is used as the output of the yolo Head. Thus, the construction of the entire YoloV5 model is completed.

[0097] Finally, the trained end-to-end multi-task learning model is used to complete the fetal ultrasound image quality control, and the YoloV5 target detection model and the Resnet50 image classification model are used to complete the fetal heart disease ultrasound image quality control. The experimental data testing is performed on a unified device, and the visual results are obtained using the PyQt5 interface.

[0098] In practical applications, the multi-task learning model of the embodiment is improved on the YoloV5 network structure, and the shared bottom layer features from the backbone network are deepened through the CNN learning method to further enhance the feature layer receptive field and provide richer feature maps for the ultrasound image multi-task network learning. The embodiment designs two kinds of multi-task learning network models. One is an ultrasound section classifier without multi-scale fusion, referred to as Our-1 model; the other is an ultrasound section classifier with multi-scale fusion, referred to as GBI-YoloV5 model, which is also the multi-task learning network model designed in the core research of the embodiment. The total process of the multi-task model design: the design principle diagram of the fetal heart disease ultrasound image quality control end-to-end multi-task network designed in the embodiment is as shown in Figure 2

[0099] In order to more suitably build and learn the machine learning model, the key region detection module and the ultrasound section classification module in the multi-task learning are regarded as independent learning tasks in the embodiment, that is, the three different scale feature maps output by the shared bottom layer network should be the feature inputs of the detection module and the classification module. The input feature maps are processed by the respective internal network to make the respective internal feature maps more conform to the respective network learning.

[0100] ​Firstly, the non-standard ultrasonic section, the standard right ventricular outflow tract section, the standard apical four-chamber heart section and the standard left ventricular outflow tract section are input into the GBI-YoloV5 multi-task learning network, and after the feature extraction of the shared bottom network, the feature map after the feature extraction is shared by the shared network principle, and the shared features are input into the key anatomical structure detector and the ultrasonic section classifier respectively to further specialize in specific tasks to learn specific features, and then the corresponding detection tasks are realized. In the key anatomical structure detector, the feature pyramid pooling is used for enhanced feature extraction, and the multi-scale feature fusion is realized by combining the feature pyramid and the feature pyramid attention mechanism structure, and finally three kinds of scale detection heads are output. In the ultrasonic section detector, the entire ultrasonic section classifier is called the Residual block module, and the Residual block module is mainly composed of a residual network structure. The three scale features shared to the Residual block module are first learned, and the feature scale is adjusted at the same time to be fused with the medium scale feature, the fused feature is further adjusted in scale to be fused with the small scale feature, the fused feature is further learned, and the learned feature is output by an average pooling layer and a full connection. Four ultrasonic section classification results. In the GBI-YoloV5 network, 10 kinds of anatomical key structures and ultrasonic section classification tasks can be detected at the same time, and finally the network outputs the anatomical key structure and the ultrasonic section type. At the same time, the multi-scale lightweight GAM channel and the spatial mixed attention mechanism which almost do not increase the network calculation amount are added in the shared bottom layer of the multi-task network, which enhances the expression ability of the feature map and reduces the interference of redundant information to the network learning.

[0101] Therefore, the embodiment is designed for the ultrasonic sections of the fetal apical four-chamber heart, left ventricular outflow tract and right ventricular outflow tract anatomical structure, and the YoloV5 target detection network is fused with the attention mechanism, combined with the ResNet-50 network, to construct an end-to-end multi-task learning network model for fetal heart disease ultrasonic image quality control, to complete the construction of the non-standard section, the standard apical four-chamber heart section, the standard left ventricular outflow tract section and the standard right ventricular outflow tract section dataset, and the section classification, key anatomical structure recognition, key anatomical structure labeling and key anatomical structure frame selection of the ultrasonic section. The main work arrangement of the embodiment is as follows:

[0102] The shared bottom layer feature principle of convolutional neural network is used to realize the fusion of the main network of YoloV5 multi-scale single-stage target detection and attention mechanism to form a shared bottom layer network. For the shared features, a feature pyramid is used to fuse one-step multi-scale features and further deep learning features to form a detector of key region anatomical structure of fetal heart ultrasound images in a multi-task learning network; a convolutional neural network is used to fuse multi-scale features to fully learn the shared features to form a fetal heart ultrasound section classifier in a multi-task learning network, thereby completing the construction of an end-to-end multi-task learning network for fetal heart disease ultrasound image quality control, and the performance of the related data set deep learning, network model training and network model testing, and the entire model network and each detection classifier are evaluated by using targeted indicators, and finally, a fetal heart disease ultrasound image analysis auxiliary result visualization system is realized by using PyQt5, which can realize fetal heart ultrasound section analysis and review by only controlling the mouse, so as to assist medical staff in diagnosing the patient's condition, improve the diagnosis efficiency, relieve the working pressure of ultrasound doctors, and relieve the medical pressure of patients.

[0103] The above-mentioned embodiments are only preferred embodiments of the present application, and cannot be used to limit the scope of protection of the present application. Any non-essential changes and replacements made by those skilled in the art on the basis of the present application shall fall within the scope of protection of the present application.

Claims

1. A multi-task learning based fetal heart disease ultrasound image quality detection method, characterized in that, The method comprises the following steps: constructing a multi-task learning network model; wherein the multi-task learning network model comprises a shared bottom network, a key anatomical structure detector, and an ultrasound section classifier; inputting the fetal ultrasound data set into the multi-task learning network model for detection, and performing feature extraction through the shared bottom network; sharing the feature maps after completing feature extraction through the shared network characteristics, and inputting the shared feature maps into the key anatomical structure detector and the ultrasound section classifier respectively to perform detection of the corresponding tasks; after the key anatomical structure detector outputs the detected and framed target region image, the image quality and image type are output in the ultrasound section classifier, and the two outputs are combined to analyze the output image; using an ultrasound image quality analysis system interface to analyze the output image and the analysis structure result to complete auxiliary ultrasound image detection; wherein the shared bottom network comprises: an input layer, the image input into the input layer is a data set image after mosaic enhancement and image adaptive preprocessing of the original image; a convolution layer structure, the shared bottom network is composed of five convolution layers with a convolution kernel of 3 and one convolution layer with a convolution kernel of 6; a residual network structure, the entire shared bottom network is composed of 4 residual structure blocks with a convolution kernel of 1 and a step of 1, wherein the main part of the residual structure block is composed of N residual blocks without residual structure which are nested continuously; an attention mechanism block, a GAM channel and spatial mixed attention mechanism are introduced before the fifth convolution layer of the YoloV5 main network to enhance feature extraction, and the shared feature map is output after the fifth convolution layer and the residual layer; wherein the activation function of each layer of the shared bottom network is a SiLU activation function; the attention mechanism block is a GAM mixed attention mechanism, which integrates a multi-layer perceptron module; after the input feature map is output through the GAM channel, it is multiplied with the original feature map according to the feature points, and the obtained feature map keeps the same size as the original feature map, and the obtained feature map is used as the input of the GAM mixed attention mechanism; in the GAM mixed attention mechanism, a convolution layer with a convolution kernel of 7 is used to increase and decrease the dimension of the feature map, i.e. using dilated convolution instead of original convolution to improve the convolution speed.

2. The method of claim 1, wherein: after generating the shared bottom network, combine SPPF, FPN and PAN to fuse multi-scale features, and use the 9 groups of anchor boxes provided by YoloV5 to generate the key anatomical structure detector; wherein there is a spatial pyramid pooling layer (SPPF) transition layer between the shared bottom network and the key anatomical structure detector, which is used to fix the size of the shared feature layer and enhance feature extraction, and provide rich feature layers for the subsequent FPN.

3. The method of claim 2, wherein: The FPN comprises a convolutional layer, a residual layer, an up-sampling layer and a feature fusion layer, the channel number and size of the feature layer from the SPPF layer are respectively adjusted by the convolutional layer and the up-sampling layer, different scale features from the shared network are fused, and the feature extraction of the fused feature layer is performed by the residual layer, so as to realize the construction of the FPN; wherein, three feature layers of different scales are still output in the FPN, the feature layer of the largest scale of the FPN is used as the input feature of the PAN and the output of the target detection network, and the medium and small scale feature layers are used as the input feature layers of the PAN; The PAN comprises a convolutional layer, a residual layer and a feature fusion layer, the channel number of the largest scale feature from the SPPF layer is adjusted by the convolutional layer, the feature is fused with the features of other scales from the PAN, the fused feature is extracted by the residual layer, and thus the PAN is constructed; wherein, two feature layers of different scales are output in the PAN structure for target detection.

4. The method of claim 1, wherein, The multi-task learning network model is trained, comprising: Step S11, initializing the shared bottom network of the GBI-YoloV5 by using the excellent model parameter of the YoloV5 network, freezing the backbone network parameter, and training the key anatomical region detection module alone; Step S12, initializing the shared bottom network of the network by using the network model trained in step S11, continuing to train the key anatomical structure detector alone, and not freezing the shared bottom network in this training; Step S13, initializing the shared bottom network by using the network model trained in step S12, freezing the shared bottom network and the key anatomical structure detector, and training the ultrasound section classifier alone; Step S14, the model trained in step S13 is the model training result of the multi-task learning network model, and the entire network can be initialized by using the result to parse the fetal heart ultrasound image quality detection.

5. The method of claim 1, wherein: after inputting the shared feature map into the ultrasound section classifier, the ultrasound section classifier uses the Conv Block residual network and the Identity Block residual network to learn the feature and adjust the size of the feature so as to be equal to the size of the feature of other scales, the adjusted feature map and the size are added to the scale feature map according to the feature points for feature fusion, the feature map is flattened after the three scale features are fused by using the average pooling, and finally a fully connected layer is used to output the classification result.

6. The method of claim 1, wherein: when the ultrasound section classifier performs the classification task, the Softmax cross-entropy loss function composed of the softmax and the cross-entropy loss function is used to replace the cross-entropy loss function, and the softmax is represented as formula (3-2): (3-2) wherein j represents a class index, and z is a discrete probability distribution result of the full connection layer output, is the kth value of the full connection layer output. the cross-entropy loss function is represented as formula (3-3): (3-3) wherein the output of formula (3-3) is the output of formula (3-2), is the true value of the classification, and the output of formula (3-3) is the final loss of the classification.

7. The method of claim 6, wherein: in the key anatomical structure detector, the loss of the entire target detection comprises a classification loss, a target confidence loss and a positioning loss, therefore, the loss function of the detector is defined as formula (3-4): (3-4) Wherein, N is the number of detection layers, B is the target number of anchor boxes assigned to a certain type of label, SxS is the number of grids into which the current scale is divided, L box , L obj , and L cls represent target confidence loss, confidence loss, and classification loss respectively, , , and represent the weights of the above three losses respectively; For the target confidence loss, the CIOU rectangular frame loss is mainly used, and the calculation principle is represented as formula (3-5): (3-5) For the confidence loss and the classification loss, the binary cross-entropy loss function is used to solve, and the calculation principle of the loss function is represented as formula (3-6): (3-6) where n is the number of detection target categories, is the label of the target category, is the probability of the label, b and b gt respectively represent the anchor box and the real box at the time of detection, W, H, w gt and h gt represent the width and height of the anchor box and the real box, respectively, and p is the distance between the center points of the anchor box and the real box, d and c are the farthest distances of the two box boundaries, and a is a weight coefficient.

8. The method of claim 1, wherein, The interface of the ultrasonic image quality analysis system is built, including: Use Anaconda software source to install the virtual environment required by PyQt5 for Python environment, configure the corresponding PyQt5 components through PyCharm software, and the components include QT Design and Pyuic components, which are PyQt5 development components; By running the QT Design component, select the appropriate window for window component layout, save it as a ui file after designing in QT Design, and use the Pyuic component to directly convert the ui into a Python script, which is stored in the form of a class in the script. Therefore, Python can call this script using the class method, and use Python programming language to complete distributed parallel processing, control function, system initialization function, so as to realize the design of PyQt5 design visualization system.

Citation Information

Patent Citations

  • Fetal heart ultrasound image recognition method based on multiple granularities

    CN113642611A

  • Remote sensing image ship small target detection method based on target perception

    CN113780152A