Ultrasonic image enhancement system

By introducing pre-training and migration units, adversarial generation network units and fusion network construction units in ultrasonic image analysis technology, the problem of insufficient accuracy of ultrasonic image analysis in the prior art is solved, and more efficient lesion degree assessment and more reliable clinical diagnosis are achieved.

CN119991448APending Publication Date: 2025-05-13JINHUA MUNICIPAL CENT HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510056336.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing ultrasound image analysis technology has insufficient accuracy in evaluating the degree of lesions, mainly because the feature extraction is too single and the model structure is simple, so it is impossible to effectively handle nonlinear relationships and complex feature interactions in ultrasound images.

Method used

An ultrasonic image enhancement system is provided, including a pre-training and migration unit, an adversarial generation network unit and a convergent network construction unit. The pre-training and migration unit uses the medical image data set to pre-train the neural network model and migrate to the ultrasonic image analysis task. The adversarial generation network unit generates high-quality simulated ultrasonic images through adversarial training of the generator and discriminator. The fusion network construction unit uses a dual-branch structure to extract original pixel features, texture and spectrum features, and fuses to evaluate the degree of lesion.

Benefits of technology

By expanding data diversity and integrating multi-feature information, the accuracy of ultrasound image analysis system in evaluating the degree of lesions is improved, the system's generalization ability is enhanced, and a more reliable clinical diagnosis basis is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991448A_ABST
    Figure CN119991448A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image analysis, in particular to an ultrasonic image enhancement system, which comprises a pre-training and migration unit, an adversarial generative network unit and a fusion network construction unit, and is characterized in that the pre-training and migration unit utilizes a medical image data set to pre-train a model and then migrate to an ultrasonic image analysis task; a generator in the generative adversarial network unit generates a simulated ultrasonic image, a discriminator distinguishes a real image from a generated image, the generator is optimized through adversarial training, the fusion network construction unit adopts a double-branch structure, original pixel features and texture and spectrum features are extracted respectively, and comprehensive features used for evaluating the lesion degree are obtained through fusion processing. All the units work cooperatively, data are effectively expanded, multi-feature information is integrated, the accuracy of evaluating the lesion degree by the ultrasonic image analysis system is improved, and a more reliable basis is provided for clinical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image analysis, and in particular to an ultrasonic image enhancement system. Background Art

[0002] Medical image analysis is an important technology. Medical images play a key role in disease diagnosis and condition monitoring. Ultrasound images are widely used due to their advantages such as being non-invasive, real-time, convenient and low-cost. However, ultrasound image analysis still faces many challenges in assessing the extent of lesions, and its accuracy needs to be improved.

[0003] Although some computer-aided diagnosis technologies have been put into use, they all have obvious shortcomings. Many technologies are too single-minded in feature extraction, focusing only on the original pixel features or simple texture features of ultrasound images, ignoring other feature dimensions in the image that contain rich information. In model construction, most of them use simple linear models or shallow neural networks. The model structure is simple and cannot effectively handle the nonlinear relationships and complex feature interactions in ultrasound images. It is difficult to fully grasp the rich information in ultrasound images and cannot deeply explore the inherent connections and potential laws between different features, making it difficult to achieve high-precision assessment of the degree of lesions. In order to solve this technical problem, we provide an ultrasound image enhancement system. Summary of the invention

[0004] The object of the present invention is to provide an ultrasonic image enhancement system to solve the problems raised in the above background technology.

[0005] To achieve the above object, an ultrasound image enhancement system is provided, comprising a pre-training and transfer unit, an adversarial generation network unit and a fusion network construction unit;

[0006] The pre-training and migration unit pre-trains the basic neural network model using the medical image data set to complete medical image feature extraction and pattern recognition, and migrates the pre-trained model to the ultrasound image analysis task;

[0007] The adversarial generative network unit includes a generator and a discriminator, wherein the generator generates a simulated ultrasound image according to the distribution characteristics of the ultrasound image, and the discriminator is used to distinguish between a real image and a generated image, and optimize the quality of the image generated by the generator through adversarial training;

[0008] The fusion network construction unit constructs a dual-branch feature fusion network, wherein the first branch is used to extract original pixel features of the ultrasound image, and the second branch uses a preprocessing module based on morphological operations and frequency domain analysis to extract texture and spectrum features of the ultrasound image, and fuses the features of the two branches to obtain comprehensive features, and uses the comprehensive features to evaluate the degree of lesions in the medical image.

[0009] As a further improvement of the technical solution, the method for pre-training the basic neural network model in the pre-training and migration unit is as follows:

[0010] A neural network model based on a deep neural network architecture includes a plurality of residual blocks, wherein the residual blocks are used to learn the residual between input and output;

[0011] Medical imaging datasets are used for pre-training. During the pre-training process, image data is input into a deep neural network. For each pixel in the image, the model extracts local features through a convolutional layer, and as the number of network layers increases, the resolution of the feature map is reduced through a pooling layer.

[0012] The cross entropy loss function is used to measure the difference between the model prediction results and the true label. Through the back propagation algorithm, the weight parameters in the model are adjusted according to the gradient calculated by the loss function to minimize the loss function. After completing the pre-training, the weight parameters of the model are migrated to the ultrasound image analysis task.

[0013] As a further improvement of the technical solution, the network structure of the generator in the adversarial generative network unit is as follows:

[0014] The generator adopts a multi-layer deconvolution network structure, in which the starting layer is a fully connected layer, which maps the random noise vector to a low-dimensional feature space to obtain an initial feature map;

[0015] Then, multiple deconvolution layers are connected in sequence. The first deconvolution layer uses a 4×4 deconvolution kernel to expand the size of the feature map to 8×8. Subsequent deconvolution layers continue to gradually expand the size of the feature map and adjust the number of channels, where the number of channels is the number of specific dimensions of the image. The last deconvolution layer uses a 4×4 deconvolution kernel.

[0016] As a further improvement of the technical solution, the network structure of the discriminator in the adversarial generation network unit is as follows:

[0017] The discriminator adopts a convolutional neural network structure, and the input is an ultrasound image, which includes a real ultrasound image and a simulated ultrasound image generated by the generator. The image first passes through a convolution layer to extract the preliminary features of the image and output the number of channels;

[0018] Then, multiple convolutional layers are connected, and the last layer is a fully connected layer, which maps the extracted features to a single output node and compresses the output value through an activation function. The output value represents the probability that the input image is a real image.

[0019] As a further improvement of the technical solution, the process of adversarial training in the adversarial generative network unit is as follows:

[0020] First, initialize the parameters of the generator and discriminator. During the training process, each iteration is divided into two steps:

[0021] S1. Fix the parameters of the generator, input the real ultrasound image and the simulated ultrasound image generated by the generator into the discriminator, and calculate the loss function of the discriminator, where for the real image, the discriminator outputs 1, indicating that it is correctly identified as a real image, and for the generated image, the discriminator outputs 0, indicating that it is correctly identified as a generated image;

[0022] S2. Fix the parameters of the discriminator, train only the generator, input a random noise vector into the generator to generate simulated ultrasound images, and then input these generated images into the discriminator to calculate the loss function of the generator, wherein the random noise vector is a vector composed of randomly generated arrays;

[0023] S3. Repeat the above two steps to iterate. During the iteration process, adjust the parameters of the generator and the discriminator proportionally to obtain a simulated ultrasound image.

[0024] As a further improvement of the technical solution, the method for extracting original pixel features of ultrasound images by the first branch in the fusion network construction unit is as follows:

[0025] The convolutional neural network structure is used to input the simulated ultrasound image, which first passes through a convolution layer to extract the local features of the image and output the number of channels;

[0026] Then, multiple convolutional layers are connected. The convolution kernel size, step size, and padding of each convolutional layer can be adaptively adjusted to obtain the feature map. The resolution of the feature map is then reduced through the pooling layer. Finally, the feature map processed by multiple convolutional layers and pooling layers is flattened to obtain a one-dimensional feature vector, which is the original pixel feature of the ultrasound image.

[0027] As a further improvement of the technical solution, the method for processing ultrasound images by the preprocessing module based on morphological operation and frequency domain analysis in the fusion network construction unit is as follows:

[0028] First, the ultrasound image is corroded using a structural element, which is a template used to define the shape and size of the corrosion operation. The formula for the corrosion operation is: Where E(x, y) is the pixel value of the image after corrosion, I(x, y) is the pixel value of the original image, and B is the structural element;

[0029] Then the same structural element is used to perform a dilation operation, and the image after the morphological operation is subjected to Fourier transform, the image is converted from the spatial domain to the frequency domain, the frequency domain image is obtained, and the features of the frequency domain image are extracted.

[0030] As a further improvement of the technical solution, the method for extracting ultrasound image texture and spectrum features by using a preprocessing module in the second branch of the fusion network construction unit is as follows:

[0031] For the image preprocessed by morphological operation and frequency domain analysis, the texture feature extraction algorithm is used to calculate the gray level co-occurrence matrix of the image in different directions and distances, and the texture feature parameters are calculated based on the gray level co-occurrence matrix;

[0032] At the same time, the spectrum features obtained by frequency domain analysis are used to extract the statistical features of the spectrum, and the extracted texture features and spectrum features are combined into a feature vector, which is the texture and spectrum features of the ultrasound image.

[0033] As a further improvement of the technical solution, the method for feature fusion in the fusion network construction unit is as follows:

[0034] The original pixel feature vector extracted by the first branch and the texture and spectrum feature vector extracted by the second branch are merged by splicing and fusion;

[0035] The concatenated feature vector is then input into a fully connected layer, the number of nodes of which is determined according to the dimension of the fused features. The fused features are transformed and integrated through the fully connected layer, and finally, the fused features are mapped to a single output value through an output layer, which is the comprehensive feature used to evaluate the degree of lesions.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] In an ultrasound image enhancement system, a pre-training and transfer unit uses a medical imaging data set to pre-train a deep neural network, extract common features and transfer them to ultrasound image analysis, avoiding the limitations of single ultrasound image data training and laying the foundation for subsequent accurate evaluation. The adversarial generative network unit generates high-quality simulated ultrasound images through adversarial training of generators and discriminators, expands data diversity, solves the problem of small sample size, and enhances the system's generalization ability. The fusion network construction unit adopts a dual-branch structure. The first branch extracts original pixel features, and the second branch extracts texture and spectrum features based on morphology and frequency domain analysis, avoiding the one-sidedness of single feature extraction. Through splicing, fusion and fully connected layer processing, it can integrate multi-feature information, effectively handle nonlinear relationships and complex feature interactions, and improve the accuracy of lesion degree recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is an overall block diagram of the present invention.

[0039] The meaning of each number in the figure is:

[0040] 1. Pre-training and transfer unit; 2. Adversarial generation network unit; 3. Fusion network construction unit. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] The present invention provides an ultrasonic image enhancement system, see Figure 1 As shown, it includes a pre-training and migration unit 1, an adversarial generation network unit 2 and a fusion network construction unit 3.

[0043] The pre-training and migration unit 1 uses the medical image data set to pre-train the basic neural network model, completes the medical image feature extraction and pattern recognition, and migrates the pre-trained model to the ultrasound image analysis task.

[0044] The method for pre-training the basic neural network model in the pre-training and migration unit 1 is as follows:

[0045] A deep neural network architecture is selected, which contains multiple residual blocks. The residual blocks learn the residual between the input and output. The residual connection can effectively solve the gradient vanishing problem of the deep network, greatly increase the network trainable depth, and can extract more complex and deep image features. Medical imaging datasets are used for pre-training. During the pre-training process, the image data is input into the deep neural network. For each pixel in the image, the model extracts local features through the convolution layer. The convolution layer can automatically learn the local feature pattern of the image, which can effectively capture the image detail information and provide a rich feature basis for subsequent image analysis.

[0046] As the number of network layers increases, the resolution of the feature map is reduced through the pooling layer, and the cross-entropy loss function is used to measure the difference between the model prediction results and the true label. The cross-entropy can well quantify the difference between the predicted and true distributions, provide a clear optimization goal for model training, and guide the model to learn the correct classification mode. The weight parameters in the model are adjusted according to the gradient calculated by the loss function through the back-propagation algorithm to minimize the loss function. After pre-training, the weight parameters of the model are transferred to the ultrasound image analysis task, and only the last few fully connected layers are fine-tuned to adapt to the specific features and analysis requirements of ultrasound images. Then, these layers are further trained using the ultrasound image dataset to generate random noise vectors. The pre-trained model has learned common features, and fine-tuning can quickly adapt to new tasks, saving training time and resources and avoiding overfitting, so that the model can play an efficient role in ultrasound image analysis tasks.

[0047] The adversarial generative network unit 2 includes a generator and a discriminator. The generator generates a simulated ultrasound image based on the distribution characteristics of the ultrasound image. The discriminator is used to distinguish between the real image and the generated image, and optimizes the image quality generated by the generator through adversarial training.

[0048] The network structure of the generator in the adversarial generation network unit 2 is as follows:

[0049] The fully connected layer can convert random noise into an initial feature representation with certain structural information, providing basic features for subsequent deconvolution operations. The generator adopts a multi-layer deconvolution network structure. The starting layer is a fully connected layer, which maps the random noise vector to a low-dimensional feature space to obtain an initial feature map, starting the initial construction process of image generation.

[0050] Then, multiple deconvolution layers are connected in sequence. The first deconvolution layer uses a 4×4 deconvolution kernel with a step size of 2 and a padding of 1. The size of the feature map is expanded to 8×8, and the number of channels is reduced to 256. The nonlinear characteristics are introduced through the ReLUj activation function. The deconvolution operation can gradually restore the image size and details. The ReLU activation function increases the expressiveness of the model, effectively improves the quality and realism of the generated image, and preliminarily constructs a feature map with more image information.

[0051] Subsequent deconvolution layers continue to gradually expand the size of the feature map and adjust the number of channels. As the number of network layers advances, the image features and structures are continuously refined, so that the generated image gradually approaches the complex characteristics of the real ultrasound image, and gradually constructs a more realistic image feature representation.

[0052] The last deconvolution layer uses a 4×4 deconvolution kernel with a stride of 2 and a padding of 1 to expand the feature map size to the same size as the ultrasound image. The number of channels is set to 1, and an activation function is used to map the output value to the interval [-1, 1] to generate a simulated ultrasound image. The activation function can limit the output to a suitable interval to meet the image pixel value range requirements, ensuring that the pixel value of the generated image is within a reasonable range, and finally generating a simulated ultrasound image that can be used for subsequent analysis and identification tasks.

[0053] The network structure of the discriminator in the adversarial generation network unit 2 is as follows:

[0054] The discriminator adopts a convolutional neural network structure. The input is an ultrasound image, including a real ultrasound image and a simulated ultrasound image generated by the generator. The image first passes through a convolution layer, using a 4×4 convolution kernel, a step size of 2, and a padding of 1 to extract the preliminary features of the image. The number of output channels is 64. Then, nonlinearity is introduced through the activation function. The convolution layer can preliminarily capture the local features of the image. The activation function can alleviate the gradient vanishing problem when processing gradients. It can not only effectively extract features but also ensure the gradient transfer in the depth direction of the network, and obtain a preliminary feature representation of the image with a certain degree of abstraction.

[0055] Then, multiple convolutional layers are connected. The convolution kernel size, step size, and padding of each convolutional layer are gradually adjusted according to the network structure design. As the number of network layers increases, the size of the feature map gradually decreases, while the degree of abstraction of the features gradually increases. Through continuous convolution and downsampling operations, the high-level semantic features of the image are gradually focused on, so that the discriminator can learn the key features to distinguish between real and generated images, thereby improving the accuracy of the discriminator's judgment on the authenticity of the image.

[0056] The last layer is a fully connected layer, which maps the extracted features to a single output node and compresses the output value to the [0, 1] interval through a function. The formula is: The output value represents the probability that the input image is a real image. The function can map any real number to an interval, which meets the requirements of probability representation. The output of the discriminator is standardized as the probability value of the image authenticity, which is convenient for the subsequent calculation of the loss function and training optimization. It can intuitively judge the possibility of whether the input image is a real ultrasound image or a simulated ultrasound image generated by the generator, providing an effective judgment basis for adversarial training.

[0057] The process of adversarial training in adversarial generative network unit 2 is as follows:

[0058] First, initialize the parameters of the generator and discriminator. Random initialization allows the model parameters to be trained from different starting points to avoid falling into local optimality. During the training process, each iteration is divided into two steps:

[0059] S1. Fix the parameters of the generator, input the real ultrasound image and the simulated ultrasound image generated by the generator into the discriminator, and calculate the loss function of the discriminator. For the real image, it is hoped that the discriminator outputs 1, indicating that it is correctly identified as a real image; for the generated image, it is hoped that the discriminator outputs 0, indicating that it is correctly identified as a generated image. The loss function of the discriminator adopts the binary cross entropy loss function, and the formula is Where n is the number of samples, y i is the true label, for the real image y i =1, for the generated image y i =0, p i The probability that the discriminator predicts a real image is obtained. Through the back propagation algorithm, the weight parameters of the discriminator are adjusted according to the gradient calculated by the loss function to minimize the loss function of the discriminator. The binary cross entropy loss function can effectively measure the difference between the discriminator prediction and the real label. Back propagation can accurately update the discriminator parameters based on the gradient, so that the discriminator can quickly learn to distinguish the features of real and generated images, and improve the accuracy and reliability of the discriminator's judgment of image authenticity.

[0060] S2. Fix the parameters of the discriminator and train only the generator. Input the random noise vector into the generator to generate simulated ultrasound images. Then input these generated images into the discriminator and calculate the loss function of the generator. The loss function of the generator also uses the binary cross entropy loss function, but the calculation method is slightly different from that of the discriminator. The formula is: Where n is the number of samples, p i In order to increase the probability that the discriminator predicts that the generated image is a real image, the weight parameters of the generator are adjusted through the back-propagation algorithm according to the gradient calculated by the loss function to minimize the loss function of the generator. This loss function can prompt the generator to generate more realistic images to confuse the discriminator. Back-propagation can guide the optimization of the generator parameters, so that the generator can continuously improve the quality of the generated images and make the generated images more similar to the real ultrasound images in feature distribution.

[0061] S3. Repeat the above two steps for multiple iterations. During the iteration process, gradually adjust the parameters of the generator and the discriminator. Reasonable adjustment of the parameters can balance the training effects of the generator and the discriminator, avoid over-training of one side and stagnation of the other side, and make the two evolve together in adversarial training and finally reach a balance.

[0062] The generator can generate simulated ultrasound images that are similar in distribution to real ultrasound images and difficult to be distinguished by the discriminator, thereby improving the data diversity and generalization ability of the ultrasound image analysis system and providing richer and more effective data support for subsequent ultrasound image analysis tasks.

[0063] The fusion network construction unit 3 constructs a dual-branch feature fusion network, the first branch is used to extract the original pixel features of the ultrasound image, and the second branch uses a preprocessing module based on morphological operations and frequency domain analysis to extract the texture and spectrum features of the ultrasound image, and fuses the features of the two branches to obtain comprehensive features, which are used to evaluate the degree of lesions.

[0064] The method for extracting original pixel features of ultrasound images by the first branch in the fusion network construction unit 3 is as follows:

[0065] The convolutional neural network structure is adopted. The simulated ultrasound image is input and first passes through a convolution layer. A 3×3 convolution kernel is used with a step size of 1 and a padding of 1 to extract the local features of the image. The number of output channels is set to 32. The 3×3 convolution kernel can effectively capture the local details of the image. The step size of 1 and padding of 1 can keep the size of the feature map unchanged. The output of 32 channels can preliminarily extract multi-dimensional features, accurately obtain the local texture, edge and other information of the image and increase the feature richness, providing input with certain feature representation for the subsequent convolution layer. The formula is Where F1(x, y) is the pixel value of the convolutional layer output feature map at position (x, y), K(i, j) is the weight of the convolution kernel at position (i, j), and I(x, y) is the pixel value of the input image at position (x, y).

[0066] Then connect multiple convolutional layers. The convolution kernel size, step size and padding of each convolution layer can be adaptively adjusted. As the number of network layers increases, the adjustment of the convolution kernel size, step size and padding is to gradually expand the receptive field and extract more abstract features. It can learn the feature representation of different levels of the image, from local details to more global structural information, and enrich the semantic information of the feature map. The calculation formula of the subsequent convolutional layer is similar to the above convolutional layer formula, except that the convolution kernel weights are different and the relevant parameters are adjusted according to the changes in the network structure.

[0067] After obtaining the feature map, the resolution of the feature map is reduced through the pooling layer. Pooling can reduce the amount of data, reduce the computational complexity and enhance the translation invariance of the feature, so that the network pays more attention to the overall features and large-scale structures of the image, and simplifies the feature map while retaining important information. The formula is Where P(x, y) is the pixel value at the (x, y) position of the pooling layer output, F(x, y) is the pixel value at the (x, y) position of the input feature map, and R is the area covered by the pooling kernel.

[0068] Finally, the feature map processed by multiple convolutional layers and pooling layers is flattened to obtain a one-dimensional feature vector, which is the original pixel feature of the ultrasound image. The flattening operation can convert the multi-dimensional feature map into a one-dimensional vector form suitable for subsequent fully connected layer processing, which is convenient for fusion with other branch features and the final lesion degree assessment task. The original pixel features of the image are organized into a unified vector representation, laying the foundation for the generation of comprehensive features.

[0069] The method for processing the ultrasound image by the preprocessing module based on morphological operation and frequency domain analysis in the fusion network construction unit 3 is as follows:

[0070] First, a structural element is used to perform corrosion operation on the ultrasound image. The formula of the corrosion operation is:

[0071]

[0072] Among them, E(x, y) is the pixel value of the image after corrosion, I(x, y) is the pixel value of the original image, and B is the structural element. The corrosion operation can remove small noise points and fine edge structures in the image, making the image smoother, reducing noise interference, and highlighting the larger target area contour in the image.

[0073] Then the same structural element is used to perform the dilation operation. The formula for the dilation operation is:

[0074]

[0075] The dilation operation can fill the holes in the image, connect adjacent areas, enhance the target area in the image, and restore some target information that may be lost due to the corrosion operation. On the basis of maintaining the integrity of the main target area of ​​the image, it further optimizes the image structure and facilitates subsequent texture feature extraction.

[0076] Perform Fourier transform on the image after morphological operation to convert the image from spatial domain to frequency domain. The formula is: Among them, F(u, v) is the frequency domain image, f(x, y) is the spatial domain image, M and N are the image sizes, u and v are the frequency domain coordinates, and Fourier transform can convert the image from the pixel representation in the spatial domain to the spectrum representation in the frequency domain. The frequency components of the image can be analyzed from the frequency domain perspective. Different frequency components correspond to different feature information of the image, and the feature distribution of the image in the frequency domain can be obtained.

[0077] After acquiring the frequency domain image, the features of the frequency domain image are extracted. These statistical features can reflect the energy distribution of the image in the frequency domain, concisely summarize the characteristics of the frequency domain image, and complement the texture and structural information of the image, providing a more comprehensive image information description for subsequent fusion with other features.

[0078] The method for extracting the texture and spectrum features of the ultrasound image by the second branch in the fusion network construction unit 3 using the preprocessing module is as follows:

[0079] For images that have been preprocessed by morphological operations and frequency domain analysis, a texture feature extraction algorithm is used to calculate the grayscale co-occurrence matrix of the image in different directions and distances. The grayscale co-occurrence matrix can reflect the spatial distribution relationship of the grayscale values ​​in the image. The matrices of different directions and distances can capture texture information of different scales and directions, comprehensively describe the image texture characteristics, and provide a data basis for the subsequent calculation of texture feature parameters.

[0080] The texture feature parameter, namely contrast, is calculated based on the gray-level co-occurrence matrix, which reflects the degree of difference in the grayscale values ​​of pixels in the image. A higher contrast indicates that the image texture changes dramatically, which provides a quantifiable basis for image classification or lesion assessment and improves the accuracy and effectiveness of texture feature analysis of ultrasound images.

[0081] At the same time, the spectrum features obtained by frequency domain analysis are used to extract the statistical characteristics of the spectrum. The spectrum features reflect the energy distribution of the image in the frequency domain, and complement the texture and structural information of the image. These features can assist in determining the type and extent of the lesion. The spectrum statistical features can summarize the image characteristics from the frequency domain perspective, and combined with texture features can more comprehensively describe the image, providing richer feature information for ultrasound image analysis.

[0082] The extracted texture features and spectral features are combined into a feature vector, which is the texture and spectral features of the ultrasound image. The combined feature vector facilitates subsequent fusion with other branch features and unified lesion severity assessment. It integrates various image feature information and provides a more comprehensive and representative feature representation for ultrasound image analysis, which is beneficial to improving the accuracy of detection and assessment of lesions in ultrasound images.

[0083] The method of feature fusion in the fusion network construction unit 3 is as follows:

[0084] The original pixel feature vector extracted by the first branch and the texture and spectrum feature vector extracted by the second branch are fused by splicing fusion. Suppose the feature vector obtained by the first branch is The eigenvector obtained by the second branch is Concatenate them into a new feature vector, the expression of the new feature vector is The stitching operation is simple and direct, and can retain the complete information of the two branch features, fully integrate the features from different sources, and not lose the feature details of either side. It provides richer and more comprehensive input features for the subsequent fully connected layer, and enables the model to comprehensively consider the original pixel information and texture spectrum information to evaluate the degree of lesions.

[0085] The concatenated feature vector is then input into a fully connected layer. The number of nodes in the fully connected layer is determined according to the dimension of the fused features and the requirements of subsequent tasks. Each node in the fully connected layer has a connection weight with each element of the input feature vector. Suppose the input feature vector Fully connected layer weight matrix W, bias vector Ze fully connected layer output Where σ is the activation function. The fully connected layer can perform nonlinear transformation and information integration on the concatenated features. By adjusting the weight matrix to learn the complex relationship between features, the model's ability to express fused features is enhanced, and more advanced and abstract feature representations are extracted, making the features more suitable for the task of assessing the severity of lesions.

[0086] Finally, an output layer is used to map the fused features to a single output value, which is the comprehensive feature used to assess the severity of the lesion. The output layer is also a fully connected layer with only one node. The weight vector of the fully connected layer is Bias b out , then the output value where σ out The activation function of the output layer compresses the multi-dimensional fusion features into a single value to facilitate intuitive assessment of the degree of lesions. The selection of the activation function makes the output conform to the semantic range of the degree of lesions, providing a concise and intuitive quantitative indicator of the degree of lesions, which can directly perform quantitative assessment of the degree of lesions in ultrasound images based on the comprehensive features.

[0087] In the present invention, the pre-training and migration unit 1 uses the medical image data set to pre-train the model and then migrates it to the ultrasound image analysis task. The generator in the adversarial generative network unit 2 generates a simulated ultrasound image, and the discriminator distinguishes between real and generated images. The generator is optimized through adversarial training. The fusion network construction unit 3 adopts a dual-branch structure to extract original pixel features and texture and spectrum features respectively, and obtains comprehensive features for evaluating the degree of lesions through fusion processing. The units work together to effectively expand data and integrate multi-feature information, thereby improving the accuracy of the ultrasound image analysis system in evaluating the degree of lesions and providing a more reliable basis for clinical diagnosis.

[0088] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. An ultrasonic image enhancement system, characterized in that: It includes a pre-training and transfer unit (1), an adversarial generation network unit (2), and a fusion network construction unit (3); The pre-training and migration unit (1) pre-trains the basic neural network model using the medical image data set to complete the medical image feature extraction and pattern recognition, and migrates the pre-trained model to the ultrasound image analysis task; The adversarial generative network unit (2) comprises a generator and a discriminator, wherein the generator generates a simulated ultrasound image according to the distribution characteristics of the ultrasound image, and the discriminator is used to distinguish between a real image and a generated image, and optimizes the quality of the image generated by the generator through adversarial training; The fusion network construction unit (3) constructs a dual-branch feature fusion network, wherein the first branch is used to extract original pixel features of the ultrasound image, and the second branch uses a preprocessing module based on morphological operations and frequency domain analysis to extract texture and spectrum features of the ultrasound image, and fuses the features of the two branches to obtain comprehensive features, and uses the comprehensive features to evaluate the degree of lesions in the medical image.

2. The ultrasonic image enhancement system according to claim 1, characterized in that: The method for pre-training the basic neural network model in the pre-training and migration unit (1) is specifically as follows: A neural network model based on a deep neural network architecture includes a plurality of residual blocks, wherein the residual blocks are used to learn the residual between input and output; Medical imaging datasets are used for pre-training. During the pre-training process, image data is input into a deep neural network. For each pixel in the image, the model extracts local features through a convolutional layer, and as the number of network layers increases, the resolution of the feature map is reduced through a pooling layer. The cross entropy loss function is used to measure the difference between the model prediction results and the true label. Through the back propagation algorithm, the weight parameters in the model are adjusted according to the gradient calculated by the loss function to minimize the loss function. After completing the pre-training, the weight parameters of the model are migrated to the ultrasound image analysis task.

3. The ultrasonic image enhancement system according to claim 1, characterized in that: The network structure of the generator in the adversarial generation network unit (2) is as follows: The generator adopts a multi-layer deconvolution network structure, in which the starting layer is a fully connected layer, which maps the random noise vector to a low-dimensional feature space to obtain an initial feature map; Then, multiple deconvolution layers are connected in sequence. The first deconvolution layer uses a 4×4 deconvolution kernel to expand the size of the feature map to 8×8. Subsequent deconvolution layers continue to gradually expand the size of the feature map and adjust the number of channels, where the number of channels is the number of specific dimensions of the image. The last deconvolution layer uses a 4×4 deconvolution kernel.

4. The ultrasonic image enhancement system according to claim 3, characterized in that: The network structure of the discriminator in the adversarial generation network unit (2) is as follows: The discriminator adopts a convolutional neural network structure, and the input is an ultrasound image, which includes a real ultrasound image and a simulated ultrasound image generated by the generator. The image first passes through a convolution layer to extract the preliminary features of the image and output the number of channels; Then, multiple convolutional layers are connected, and the last layer is a fully connected layer, which maps the extracted features to a single output node and compresses the output value through an activation function. The output value represents the probability that the input image is a real image.

5. The ultrasonic image enhancement system according to claim 4, characterized in that: The adversarial training process in the adversarial generative network unit (2) is specifically as follows: First, initialize the parameters of the generator and discriminator. During the training process, each iteration is divided into two steps: S1. Fix the parameters of the generator, input the real ultrasound image and the simulated ultrasound image generated by the generator into the discriminator, and calculate the loss function of the discriminator, where for the real image, the discriminator outputs 1, indicating that it is correctly identified as a real image, and for the generated image, the discriminator outputs 0, indicating that it is correctly identified as a generated image; S2. Fix the parameters of the discriminator, train only the generator, input a random noise vector into the generator to generate simulated ultrasound images, and then input these generated images into the discriminator to calculate the loss function of the generator, wherein the random noise vector is a vector composed of randomly generated arrays; S3. Repeat the above two steps to iterate. During the iteration process, adjust the parameters of the generator and the discriminator proportionally to obtain a simulated ultrasound image.

6. The ultrasonic image enhancement system according to claim 5, characterized in that: The method for extracting original pixel features of ultrasound images by the first branch in the fusion network construction unit (3) is specifically as follows: The convolutional neural network structure is used to input the simulated ultrasound image, which first passes through a convolution layer to extract the local features of the image and output the number of channels; Then, multiple convolutional layers are connected. The convolution kernel size, step size, and padding of each convolutional layer can be adaptively adjusted to obtain the feature map. The resolution of the feature map is then reduced through the pooling layer. Finally, the feature map processed by multiple convolutional layers and pooling layers is flattened to obtain a one-dimensional feature vector, which is the original pixel feature of the ultrasound image.

7. The ultrasonic image enhancement system according to claim 6, characterized in that: The method for processing the ultrasound image by the preprocessing module based on morphological operation and frequency domain analysis in the fusion network construction unit (3) is as follows: First, the ultrasound image is corroded using a structural element, which is a template used to define the shape and size of the corrosion operation. The formula for the corrosion operation is: Where E(x, y) is the pixel value of the eroded image, I(x, y) is the pixel value of the original image, and B is the structural element; Then the same structural element is used to perform a dilation operation, and the image after the morphological operation is subjected to Fourier transform, the image is converted from the spatial domain to the frequency domain, the frequency domain image is obtained, and the features of the frequency domain image are extracted.

8. The ultrasonic image enhancement system according to claim 7, characterized in that: The second branch in the fusion network construction unit (3) uses a preprocessing module to extract ultrasound image texture and spectrum features, which is specifically as follows: For the image preprocessed by morphological operation and frequency domain analysis, the texture feature extraction algorithm is used to calculate the gray level co-occurrence matrix of the image in different directions and distances, and the texture feature parameters are calculated based on the gray level co-occurrence matrix; At the same time, the spectrum features obtained by frequency domain analysis are used to extract the statistical features of the spectrum, and the extracted texture features and spectrum features are combined into a feature vector, which is the texture and spectrum features of the ultrasound image.

9. The ultrasonic image enhancement system according to claim 8, characterized in that: The method for feature fusion in the fusion network construction unit (3) is specifically as follows: The original pixel feature vector extracted by the first branch and the texture and spectrum feature vector extracted by the second branch are merged by splicing and fusion; The concatenated feature vector is then input into a fully connected layer, the number of nodes of which is determined according to the dimension of the fused features. The fused features are transformed and integrated through the fully connected layer, and finally, the fused features are mapped to a single output value through an output layer, which is the comprehensive feature used to evaluate the degree of lesions.

Citation Information

Cited By

  • Ultrasonic image enhancement method and system based on multi-modal image migration and medium

    CN120852200A