Intelligent Processing Method, Electronic Device and Storage Medium for Ultrasonic Images of Hepatic Echinococcosis Based on Deep Learning

Through deep learning-based methods, intelligent processing of echinococcosis ultrasound images has been solved, and the problems of low accuracy and frequent misdiagnosis of traditional ultrasound diagnosis have been solved, achieving more efficient and accurate diagnosis.

CN119624959BActive Publication Date: 2025-06-10JIANGSU SHIYU INTELLIGENT MEDICAL TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510154559.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-10
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Traditional ultrasound diagnosis has low accuracy for echinococcosis and is prone to misdiagnosis.

Method used

The intelligent ultrasonic image processing method of hepatic echinococcosis based on deep learning is adopted, including data collection, image preprocessing and classification model training. The specific steps include the adaptive denoising and enhancement model preprocessing the ultrasonic image data, establishing a classification data set, and training the classification model DFEV-VIT.

Benefits of technology

It significantly improves the diagnostic efficiency and accuracy of echinococcosis and reduces misdiagnosis and misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624959B_ABST
    Figure CN119624959B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent processing method, an electronic device and a storage medium for ultrasound images of hepatic echinococcosis based on deep learning, including the following steps: Data collection: Collect hepatic ultrasound image data; Use an adaptive denoising and enhancement model to perform image preprocessing on the collected hepatic ultrasound image data, and classify and label the data after image preprocessing to establish a classification data set, and label the images as cystic echinococcosis, alveolar echinococcosis, and other liver lesions; Based on the classification data set, train an optimized VIT classification model DFEV-VIT. Beneficial effects of the present invention: By using the optimized CLAHE algorithm ADEN to perform image preprocessing on the collected data, and then training the classification model DFEV-VIT based on the classification data set after image preprocessing, the diagnostic efficiency and accuracy can be greatly improved, and the missed diagnosis and misdiagnosis of echinococcosis can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to an intelligent processing method, an electronic device and a storage medium for ultrasonic images of hepatic echinococcosis based on deep learning. Background Technique

[0002] Echinococcosis, also known as hydatid disease, is a disease caused by parasites. The pathogen of this disease is Echinococcus granulosus, which is mainly transmitted to humans through contact infection or consumption of undercooked animal products infected with this parasite. Echinococcosis can affect organs such as the liver and lungs, and can lead to death in severe cases. Among them, hepatic echinococcosis mainly includes cystic hepatic echinococcosis and alveolar hepatic echinococcosis. Almost all primary alveolar echinococcosis lesions and 70% of cystic echinococcosis lesions originate from the liver. The early stage of echinococcosis can be asymptomatic for 10 - 15 years. The long latency period and infiltrative growth result in many patients with alveolar echinococcosis being inoperable when the lesions are finally diagnosed. Some patients with alveolar echinococcosis receive incorrect treatment due to misdiagnosis, and patients with cystic echinococcosis are often over-treated. With the application of big data and artificial intelligence in the medical field, classifying the ultrasonic images of patients into gold-standard cases for model training to assist doctors in the clinical diagnosis and treatment of echinococcosis has become a trend. However, the following problems exist in the existing image processing of echinococcosis:

[0003] 1. The ultrasonic features of echinococcosis overlap greatly with those of many other focal liver lesions, but there are important differences in the treatment strategies for these diseases. Specifically, for alveolar echinococcosis, the ultrasonic manifestations vary greatly and may be misdiagnosed as hepatic hemangioma, liver abscess, intrahepatic cholangiocarcinoma or liver metastasis;

[0004] 2. In a prospective study conducted in Germany in 2000, a total of 55 echinococcosis lesions in 30 patients were detected by CT, but only 31 of these lesions were detected by ultrasonic imaging, with a low accuracy and prone to misdiagnosis. Summary of the Invention

[0005] In view of this, the present invention aims to propose an intelligent processing method, an electronic device and a storage medium for ultrasonic images of hepatic echinococcosis based on deep learning to address the challenges of low accuracy and misdiagnosis of traditional ultrasonic diagnosis for echinococcosis.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows:

[0007] An intelligent processing method for ultrasonic images of hepatic echinococcosis based on deep learning, comprising the following steps:

[0008] S1. Data collection: Collect hepatic ultrasonic image data;

[0009] S2. Use the adaptive denoising and enhancement model to perform image preprocessing on the liver ultrasound image data collected in step S1, and classify and label the data pairs after image preprocessing to establish a classification dataset, and label the images as cystic echinococcosis, alveolar echinococcosis, and other liver lesions;

[0010] S3. Based on the classification dataset, train the classification model DFEV-VIT;

[0011] In step S2, use the adaptive denoising and enhancement model to perform image preprocessing on the liver ultrasound image data collected in step S1, including:

[0012] S21. Evaluate the signal-to-noise ratio of the image and suppress noise;

[0013] S22. Perform adaptive wavelet denoising on the image processed in step S21;

[0014] S23. Perform edge-preserving filtering and non-local means denoising on the image processed in step S22;

[0015] S24. Perform gray-scale dynamic expansion on the image processed in step S23;

[0016] S25. Perform dynamic block processing on the image processed in step S24;

[0017] S26. Set an inter-block overlap module for the image processed in step S25;

[0018] S27. Perform contrast-limited histogram equalization on the image processed in step S26;

[0019] S28. Perform image reconstruction on the image processed in step S27;

[0020] S29. Perform image post-processing and optimization on the image processed in step S28;

[0021] In step S3, based on the classification dataset, train the classification model DFEV-VIT, including:

[0022] S31. Perform image preprocessing and batching on the classification dataset;

[0023] S32. Embed the image blocks for the data processed in step S31;

[0024] S33. Perform position encoding and class marking on the data processed in step S32;

[0025] S34. Process the data processed in step S33 using a Transformer encoder;

[0026] S35. Process the features of the data processed in step S34 using a feature rejection mechanism;

[0027] S36. Classify and predict the data processed in step S35;

[0028] S37. Calculate the loss and perform backpropagation on the data processed in step S36.

[0029] Furthermore, in step S21, the signal-to-noise ratio evaluation and noise suppression include:

[0030] Evaluate the signal-to-noise ratio of the image using the standard deviation method, and measure the speckle noise and signal intensity in the image; among them, the calculation of the signal intensity: calculate the average gray value of the image, and the average gray value represents the signal intensity. The expression of the average gray value:

[0031] ;

[0032] Among them, is the total number of pixels in the region, is the th pixel's gray value;

[0033] Calculation of the speckle noise intensity: Select a region with 1 / 10 of the current picture size in the upper right corner of the image as the background, and calculate the standard deviation of this region:

[0034] ;

[0035] Calculation of the signal-to-noise ratio:

[0036] ;

[0037] Among them, is the average value of this region.

[0038] Furthermore, in step S22, the adaptive wavelet denoising includes:

[0039] First, for the images with a signal-to-noise ratio lower than 20 in step S21, use adaptive wavelet denoising to remove noise;

[0040] In this process, the image needs to be transformed into the wavelet domain, and the part corresponding to the noise in the wavelet coefficients of the image is removed through threshold processing;

[0041] If the signal-to-noise ratio after denoising is higher than the set value, keep the current threshold;

[0042] If the signal-to-noise ratio does not increase or decreases, adjust the threshold size dynamically according to the feedback;

[0043] Finally, use the inverse wavelet transform to transform the denoised wavelet coefficients back to the spatial domain to generate an ultrasonic image.

[0044] Further, in step S23, edge-preserving filtering and non-local means denoising include:

[0045] Apply non-local means denoising to the image that has undergone wavelet denoising first. For each pixel in the image to be processed, calculate the similarity by comparing it with other pixels in its neighborhood:

[0046] ;

[0047] where is the pixel value, is the similarity control parameter, represents the difference of the corresponding pixel, is the normalization factor, and perform weighted averaging on the pixels according to the weights calculated by the similarity to obtain the denoised pixel value:

[0048] ;

[0049] where is the neighborhood set of pixel ;

[0050] After the non-local means denoising is completed, use an edge-preserving filter to process the denoised image to clarify the edge and texture features:

[0051] ;

[0052] where is the new gray value of the pixel at position x in the image after edge filtering processing, x and y are the coordinates of the pixel being processed, are the Gaussian functions in the spatial domain and the gray domain respectively, is the normalization factor.

[0053] Further, in step S24, gray-scale dynamic expansion includes:

[0054] Perform gray-scale dynamic expansion on the denoised image. Adopt the histogram equalization method to redistribute the gray values of the image. Use the linear mapping rule to gradually expand the pixels concentrated in the set gray range to the entire gray range; calculate the gray histogram of the original image to determine the gray value distribution, find the minimum and maximum gray values of the image, and convert them to the standard gray range through linear mapping, and apply the mapped values to all pixels in the image.

[0055] Further, in step S25, dynamic block processing includes:

[0056] Dynamically adjust the size and shape of the image blocks according to the texture of the image, and calculate the texture features of each region of the image using local gradients and gray-level co-occurrence matrices.

[0057] Further, in step S26, a block - to - block overlap module is set up, including:

[0058] An overlapping area is set between adjacent blocks, and the overlapping ratio ranges from 15% to 25%.

[0059] Further, in step S27, contrast - limited histogram equalization includes:

[0060] First, use an edge - detection algorithm to identify features in the image by analyzing the edge map of the image;

[0061] Apply a clustering algorithm to the edge map to divide the image into different regions;

[0062] The adaptive denoising and enhancement model calculates the gray - level histogram for each region and restricts the uniform distribution of the histogram according to the set dynamic truncation threshold. The formula for the histogram is:

[0063] ;

[0064] where, represents the number of pixels at gray - level , represents the gray - level value of the image at position , represents the width of the image, represents the height of the image, represents the Kronecker function, is the gray - level, regarded as a constant;

[0065] Then, process the histogram according to the dynamic truncation threshold. The histogram of each region is restricted as follows:

[0066] ;

[0067] where, represents the total number of pixels in the specific image region currently being processed, represents the frequency threshold;

[0068] After completing the histogram - equalization process, the adaptive denoising and enhancement model performs a region - merging operation. The merging process is represented by the weighted - average method:

[0069] ;

[0070] where, and are the equalized images of the left - hand and right - hand regions respectively, and the weight factor ranges from 0 to 1.

[0071] Further, in step S28, image reconstruction includes:

[0072] Adopt a gradient-guided interpolation method to eliminate the artifacts generated by block processing;

[0073] First, analyze the gradient changes between adjacent blocks, and then use an edge-preserving interpolation method to combine the pixel values of adjacent blocks through weights to preserve the edge features of the image during the reconstruction process;

[0074] Merge the processed blocks into a complete ultrasound image.

[0075] Further, in step S29, image post-processing and optimization include:

[0076] Apply non-local means denoising, find similar pixels in the same region based on similarity, and remove the remaining noise:

[0077] ;

[0078] Where, is the pixel value after denoising, : normalization coefficient, represents all pixel indices within the search window, represents the weight, calculated through similarity measurement , is the smoothing parameter, controlling the influence range of the weight;

[0079] Apply an adaptive sharpening filter to only sharpen the edge region of the image. The adaptive sharpening of the image is expressed as:

[0080] ;

[0081] Where, represents the pixel value after sharpening, represents the gradient value calculated at point ; represents the parameter controlling the sharpening intensity; Edge detection is achieved by calculating the gradient of the pixel. The gradient expression of the pixel is:

[0082] .

[0083] Further, in step S31, image preprocessing and blocking include:

[0084] First, perform mean normalization using the dynamic normalization method:

[0085] ;

[0086] Where, is the pixel value of the input image at the position , and are the mean and standard deviation of the local region respectively, is a constant;

[0087] Evaluate the complexity of each image block through information entropy and determine the priority of subsequent feature extraction. The information entropy expression of region A is:

[0088] ;

[0089] where is the probability distribution of the pixel values in region , is the number of possible pixel values;

[0090] By evaluating the information entropy value, the size and shape of the image block can be dynamically selected. The image set is . In this process, the chunking strategy follows the following principle: for regions with information entropy higher than the set value, select image blocks smaller than the contrast threshold; for regions with information entropy lower than the set value, select image blocks larger than the contrast threshold.

[0091] Furthermore, in step S32, image block embedding includes:

[0092] For each preprocessed and chunked image block , set an embedding conversion module. The embedding conversion module converts the original image block into a feature vector with semantic information through an attention mechanism:

[0093] First, convert the input image block into a query , key and value through a linear transformation:

[0094] ;

[0095] Calculate the similarity score by using the proposed query and all keys :

[0096] ;

[0097] where is the dimension of the key;

[0098] Normalize the score through softmax to obtain the attention weight:

[0099] ;

[0100] Use the attention weights to perform weighted summation on the values to obtain the final feature representation:

[0101] .

[0102] Furthermore, in step S33, the positional encoding and the class token include:

[0103] Set up a hybrid positional encoding model that includes learnable parameters and a periodic function:

[0104] ;

[0105] where corresponds to the two-dimensional spatial coordinates of the image patch, and is a non-linear mapping function that maps to other dimensions, is a bias composed of learning parameters, and outputs values with the same dimension as the input coordinates;

[0106] Introduce an adaptive spatial weight, and the adaptive spatial weight adjusts the influence of the positional encoding of different image patches based on information entropy:

[0107] ;

[0108] where, is the output of a multi-layer perceptron that processes the positional encoding; expand the original single class token into feature aggregation, and the feature aggregation is represented as , and use the multi-head self-attention mechanism of Vision Transformer to extract and map global features.

[0109] Furthermore, in step S34, the Transformer encoder processing includes:

[0110] The Transformer encoder includes multiple stacked encoder layers, and each encoder layer includes two key modules: the multi-head self-attention mechanism and the feed-forward neural network;

[0111] The multi-head self-attention mechanism allows each image patch to establish an interaction relationship with all other image patches, and add an interactive bidirectional mode in the multi-head self-attention mechanism:

[0112] ;

[0113] ;

[0114] The reverse attention mechanism swaps the roles of keys and queries;

[0115] Introduce the attention gating mechanism:

[0116] ;

[0117] Among them, and are learnable parameters, the sigmoid activation function, denotes element-wise multiplication; through the gating mechanism, the multi-head self-attention mechanism can adaptively select to retain or suppress information;

[0118] The feed-forward neural network then performs a non-linear transformation on the output of the multi-head self-attention mechanism, and the output of the multi-head self-attention mechanism will be fed into the feed-forward neural network:

[0119] ;

[0120] The feed-forward neural network includes multiple feed-forward layers and activation functions, and the operation of each feed-forward layer is expressed as:

[0121] ;

[0122] Among them, , is the weight matrix, , is the bias term;

[0123] Each encoder layer also includes layer normalization and residual connection, and layer normalization ensures that the input of each layer maintains statistical characteristics:

[0124] ;

[0125] Among them, and are the mean and standard deviation respectively, and are learnable parameters;

[0126] The residual connection is to jump-connect the input to the output.

[0127] Furthermore, in step S35, the feature rejection mechanism is used to process features, including:

[0128] By calculating the weights of each feature in the embedding vector evaluate the contribution of each feature using the standard deviation and information entropy metrics:

[0129] ;

[0130] Among them, is the The weight of a feature, is the value of the th feature;

[0131] Dynamically adjust the rejection threshold according to the batch size:

[0132] ;

[0133] For each element in the feature, if < , then the feature will be regarded as unimportant and set to zero.

[0134] Furthermore, in step S36, the classification prediction includes:

[0135] After being processed by the Transformer encoder, the added class token contains the global semantic information of the entire image. Pass the added class token as input to the multi-head self-attention mechanism. The multi-head self-attention mechanism includes two fully connected layers. A non-linear activation function is used between the two fully connected layers, and the final output layer uses Softmax activation to generate the probability distribution of each class.

[0136] Furthermore, in step S37, the loss calculation and backpropagation include:

[0137] Use the cross-entropy loss function to calculate the difference between the model prediction and the true label. The backpropagation algorithm calculates the gradient of the loss with respect to the model parameters and uses an optimizer to update the model parameters.

[0138] An electronic device includes a processor and a memory communicatively connected to the processor and used to store instructions executable by the processor. The memory stores instructions executable by the processor, and the instructions are executed by the processor. The processor is used to execute the above-mentioned intelligent processing method for ultrasound images of hepatic echinococcosis based on deep learning.

[0139] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned intelligent processing method for ultrasound images of hepatic echinococcosis based on deep learning.

[0140] Compared with the prior art, the intelligent processing method for ultrasound images of hepatic echinococcosis based on deep learning, the electronic device and the storage medium of the present invention have the following advantages:

[0141] The intelligent processing method, electronic device and storage medium for ultrasonic images of hepatic echinococcosis based on deep learning. In the present invention, the collected data is preprocessed by using the ADEN model, and then the classification model DFEV-VIT is trained based on the classification data set after image preprocessing, which can greatly improve the diagnostic efficiency and accuracy, and reduce the missed diagnosis and misdiagnosis of echinococcosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0142] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0143] Figure 1 It is a schematic flowchart of the processing method described in the embodiment of the present invention;

[0144] Figure 2 It is a schematic flowchart of the image preprocessing described in the embodiment of the present invention;

[0145] Figure 3 It is a schematic flowchart of the classification model described in the embodiment of the present invention

[0146] Figure 4 It is a schematic flowchart of training the classification model DFEV-VIT based on the classification data set described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0147] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0148] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.

[0149] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0150] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0151] As Figures 1 to 4 shown, an intelligent processing method for ultrasound images of hepatic echinococcosis based on deep learning includes the following steps:

[0152] S1. Data collection: Collect high-quality hepatic ultrasound image data to ensure the representativeness and diversity of the data

[0153] S2. Image preprocessing: The present invention uses an optimized CLAHE algorithm (ADEN, adaptive denoising and enhancement model) for image reconstruction processing. After the image reconstruction processing, classification and annotation are performed, clear pathological classification criteria are defined, a classification data set is established, and the images are labeled as cystic echinococcosis, alveolar echinococcosis, and other hepatic lesions;

[0154] S3. Based on the classification data set, train the classification model DFEV-VIT. Finally, the experimental result of this method is Top1 ACC = 95%.

[0155] In a preferred embodiment of the present invention, in step S2, the image preprocessing includes:

[0156] Adaptive denoising and enhancement model (ADEN):

[0157] Signal-to-noise ratio evaluation and noise suppression: First, use the standard deviation method to evaluate the signal-to-noise ratio of the image to measure the speckle noise and signal intensity in the image. Calculation of signal intensity: Calculate the average gray value of the image. This value represents the signal intensity:

[0158] ;

[0159] Among them, is the total number of pixels in the region, is the th pixel's gray value.

[0160] Calculation of noise intensity: Select a region with 1 / 10 of the current image size in the upper right corner of the image as the background, and calculate the standard deviation of this region:

[0161] ;

[0162] Among them, is the average value of this area.

[0163] Calculation of signal-to-noise ratio:

[0164] ;

[0165] The higher the value, the clearer the image and the less noise.

[0166] Adaptive wavelet denoising: Next, for images with SNR lower than 20, adaptive wavelet denoising is used to remove noise. In this process, the image is transformed into the wavelet domain, and the part corresponding to the noise in the wavelet coefficients is removed through threshold processing. A suitable threshold (the threshold is obtained by analyzing a large number of sample data through a random forest model to learn different types of noise and their corresponding optimal thresholds) is selected to achieve better noise suppression and edge preservation. This threshold is obtained by analyzing a large number of sample data through a random forest model to learn different types of noise and their corresponding optimal thresholds. In order to use the inverse wavelet transform to convert the denoised wavelet coefficients back to the spatial domain, a denoised ultrasound image is generated. If the SNR after denoising is significantly improved, it indicates that the processing effect is ideal, and the current threshold can be maintained at this time; if the SNR is not improved or decreased, the threshold is dynamically adjusted according to the feedback, for example, increasing the threshold (it can be 10%) to remove more high-frequency noise, or decreasing the threshold (it can be 10%, usually increasing first and then decreasing, and if it is still lower than 20, continue to increase or decrease until it reaches 20%) to restore more image details. Through this feedback mechanism, it is possible to effectively make adaptive adjustments for different image characteristics to ensure a more ideal denoising effect. Finally, the inverse wavelet transform is used to convert the denoised wavelet coefficients back to the spatial domain to generate a clear ultrasound image. This method realizes more efficient noise suppression by dynamically adjusting the threshold, while retaining important edge features and image details, significantly improving the image quality.

[0167] Edge-preserving filter and non-local means denoising: In order to further enhance the details and edge features of the image, the image after wavelet denoising is first applied with non-local means denoising to further suppress noise and preserve the details of the image. For each pixel in the image to be processed, the similarity is calculated by comparing other pixels in its neighborhood:

[0168] ;

[0169] Among them, is the pixel value, is the similarity control parameter, represents the difference of the corresponding pixel, is the normalization factor. The pixels are weighted and averaged according to the weights calculated by the similarity to obtain the denoised pixel values:

[0170] ;

[0171] where is the pixel neighborhood set.

[0172] After the non-local means denoising is completed, an edge-preserving filter (bilateral filter) is used to process the denoised image to clarify important edge and texture features:

[0173] ;

[0174] where, is the new gray value of the pixel at position x in the image after edge filtering processing, x, y are the coordinates of the pixel being processed, are the Gaussian functions in the spatial domain and the gray domain respectively, is the normalization factor.

[0175] Gray-scale dynamic expansion: The gray-scale of the denoised image is dynamically expanded to improve the contrast of the image. The histogram equalization method is used to redistribute the gray values of the image. Using the linear mapping rule, the pixels concentrated in a certain gray-scale range (calculate the gray-scale histogram of the original image to determine the gray-scale value distribution) are gradually expanded to the entire gray-scale range. Calculate the gray-scale histogram of the original image to determine the gray-scale value distribution. Find the minimum and maximum gray-scale values of the image and convert them to the standard gray-scale range (0 to 255) through linear mapping. Apply the mapped values to all pixels in the image to enhance the overall contrast.

[0176] Dynamic block processing: Dynamically adjust the size and shape of the blocks according to the texture complexity of the image. Calculate the texture features of each region using local gradients and gray-level co-occurrence matrices. Larger blocks are used in simple texture regions, while smaller blocks are used in complex regions (all sizes are designed for the contrast threshold, greater than 100 is a large block, and the rest are small blocks) to ensure capturing key details.

[0177] Set the inter - block overlap module: To improve the continuity of the image, the present invention designs an inter - block overlap module. Set the overlap ratio between 15% and 25%. This overlap design can not only improve the detail capture rate but also effectively reduce the boundary blurring caused by block - based processing. For example, for each block, its edge part can be covered into the adjacent block, so that more context information can be retained during splicing. In this embodiment, by setting an overlap area between adjacent blocks, the inter - block overlap module can effectively alleviate the edge blurring caused by block - based processing. During the actual processing, some detailed information of the image may be ignored, lost or become unclear during block - based processing. The overlap design helps to retain and strengthen these edge information, making the transition between different blocks smoother and enhancing the continuity of the image.

[0178] Contrast - limited histogram equalization: First, use an edge detection algorithm (such as Canny detection) to identify the key features in the image (the areas with significant contrast changes in the image, and these areas are called key features). This step can not only find the edges of the image but also provide important information for subsequent region division. By analyzing the edge map, the model can determine those key regions, which are usually the places with large contrast changes in the image. Next, a clustering algorithm (K - means) is applied to the edge map to divide the image into different regions. The feature similarity of these regions enables them to share the same random features during processing, thus maintaining the uniformity and natural effect between regions. After determining the regions of the image, the model calculates the gray - level histogram for each region and limits the uniform distribution of the histogram according to the set dynamic truncation threshold. The formula for the histogram is:

[0179] ;

[0180] where, represents the number of pixels with gray - level , represents the gray - level value of the image at position , represents the width of the image, represents the height of the image, represents the Kronecker function, is the gray - level, which can be regarded as a constant.

[0181] Then, process the histogram according to the dynamic truncation threshold, and the histogram of each region is limited in the following way:

[0182] ;

[0183] where, represents the total number of pixels in the specific image region being processed currently, Indicates the frequency threshold, which is 0.01 for this project.

[0184] After completing the histogram equalization process, the model also performs region merging operations. During this process, it pays special attention to the overlapping parts of adjacent regions to ensure that not only the details of the image can be maintained during merging, but also the visual seam phenomenon caused by region segmentation can be effectively reduced. The specific merging process can be represented by the weighted average method:

[0185] ;

[0186] Among them, and are the equalized images of the left and right regions respectively, and the weight factor usually ranges from 0 to 1. Adjusting this factor can ensure smooth gradients.

[0187] Image reconstruction: In the image reconstruction stage, a gradient-guided interpolation method is used to eliminate the artifacts generated by the block processing. First, the gradient changes between adjacent blocks are analyzed, and then an edge-preserving interpolation method is used to combine the pixel values of adjacent blocks through weights to ensure that the edge features of the image are retained during reconstruction. The processed blocks are merged into a complete ultrasound image to achieve clear edges and natural transitions.

[0188] Image post-processing and optimization: The quality of the final image is further improved by using non-local means denoising and adaptive sharpening techniques. Applying non-local means denoising, similar pixels within the same region are found based on similarity to effectively remove residual noise:

[0189] ;

[0190] Among them, is the pixel value after denoising, : the normalization coefficient. represents all pixel indices within the search window, represents the weight, usually calculated through similarity metrics , is the smoothing parameter that controls the influence range of the weight.

[0191] An adaptive sharpening filter is applied to sharpen only the edge regions, enhancing the edge features without over-enhancing other parts of the image. The adaptive sharpening of the image can be expressed as:

[0192] ;

[0193] represents the pixel value after sharpening.

[0194] Indicates the gradient value calculated at the point , usually obtained by using the Sobel operation.

[0195] Indicates the parameter that controls the sharpening intensity. Edge detection can be achieved by calculating the gradient of pixels. The gradient expression of pixels is:

[0196] ;

[0197] In a preferred embodiment of the present invention, in step S3, the dynamic feature extraction VIT classification model (i.e., the vision transformer model) includes:

[0198] Image preprocessing and chunking: The input image undergoes a strict preprocessing process. First, the dynamic normalization method is used to replace the traditional mean normalization:

[0199] ;

[0200] Among them, is the pixel value of the input image at the position , and are the mean and standard deviation of the local region respectively, is a small constant to avoid division by zero errors. Through such dynamic normalization, the model can adaptively evaluate the local information complexity of the image for subsequent chunking.

[0201] In the process of information processing and feature extraction, it is very important to evaluate the complexity of each image chunk through information entropy. Information entropy can not only be used as a quantitative index for the content and complexity of the image, but also be used to determine the priority of subsequent feature extraction.

[0202] ;

[0203] Among them, is the probability distribution of pixel values in the region , is the number of possible pixel values. By evaluating the entropy value, we can dynamically select the size and shape of the image chunks. The image set is . In this process, the chunking strategy follows the following principle: for regions with higher information entropy , smaller image chunks are selected. For regions with lower information entropy , larger image chunks are selected to perform effective computational resource allocation.

[0204] Image chunk embedding: For each preprocessed and chunked image chunk , we designed a learnable embedding transformation module. The main purpose of this module is to convert the original image patches into semantic-rich feature vectors through the attention mechanism, thereby improving the efficiency of subsequent feature extraction and classification:

[0205] First, the input image patch is linearly transformed into a query , a key and a value :

[0206] ;

[0207] The proposed query is used to calculate the similarity scores with all keys :

[0208] ;

[0209] where is the dimension of the key. Normalization is to prevent the scores from being too large, which may lead to the problem of gradient explosion.

[0210] The scores are normalized by softmax to obtain the attention weights:

[0211] ;

[0212] The attention weights are used to perform a weighted sum on the values to obtain the final feature representation:

[0213] ;

[0214] Position encoding and class token: Position encoding is a crucial component in Vision Transformer. Learnable position encoding vectors will be directly added to the embedding representation of the image patches. These position encoding vectors enable the model to distinguish image patches at different positions and retain the spatial structure information. To capture position features more effectively, we designed a hybrid position encoding model that includes learnable parameters and periodic functions:

[0215] ;

[0216] where corresponds to the two-dimensional spatial coordinates of the image patch. and are non-linear mapping functions that map to other dimensions. is a bias composed of learnable parameters, and the output is a value with the same dimension as the input coordinates.

[0217] To enhance the model's sensitivity to complex images, an adaptive spatial weight is introduced, which adjusts the influence of positional encoding for different image patches based on information entropy:

[0218] ;

[0219] where is the output of the multi-layer perceptron for processing positional encoding. To support class understanding in classification tasks, we expand the original single-class label into a rich feature aggregation representation , and use the multi-head self-attention mechanism of Vision Transformer to extract and map global features.

[0220] Transformer encoder processing: The Transformer encoder is composed of multiple identical encoder layers stacked together. Each encoder layer contains two key modules: the multi-head self-attention mechanism and the feed-forward neural network.

[0221] The multi-head self-attention mechanism allows each image patch to establish complex interaction relationships with all other image patches. In this study, an interactive bidirectional mode is added to the multi-head attention mechanism: ;

[0222] ;

[0223] The reverse attention mechanism swaps the roles of keys and queries, enabling the model to focus on the information of subsequent elements based on previous elements. This combination allows the model to fully utilize the information of all nodes in the sequence when calculating the representation of each feature, thereby enhancing the representation ability of the context. To further enhance the flexibility of information flow, an attention gating mechanism is introduced:

[0224] ;

[0225] and are learnable parameters, the sigmoid activation function, denotes element-wise multiplication. Through the gating mechanism, the multi-head attention mechanism can adaptively select to retain or suppress certain information.

[0226] The feed-forward neural network then performs a non-linear transformation on the output of the attention mechanism to further extract and process features. Next, the output of self-attention is fed into the feed-forward neural network:

[0227] ;

[0228] The feed-forward neural network usually contains multiple layers and activation functions (such as ReLU) to enable non-linear transformation. The operation of each feed-forward layer can be expressed as:

[0229] ;

[0230] Among them, , is the weight matrix, , is the bias term.

[0231] Each encoder layer also includes layer normalization and residual connections. Layer normalization ensures that the inputs to each layer maintain stable statistical properties:

[0232] ;

[0233] Among them, and are the mean and standard deviation respectively, and are learnable parameters. Layer normalization makes the distribution of the output more stable during each forward pass, improving the convergence speed.

[0234] The residual connection directly jumps the input to the output, alleviating the vanishing gradient problem in deep networks. The encoder layers are stacked continuously. In this study, 12 layers are stacked, and each layer progressively extracts more abstract and semantically rich feature representations.

[0235] Feature rejection mechanism: To effectively reduce the computational overhead and improve the expressive power of the model, we use a rejection mechanism. By calculating the weights of each feature in the embedding vector , determine which features are redundant in the task. Use metrics such as standard deviation and entropy to evaluate the contribution of each feature:

[0236] ;

[0237] Among them, is the weight of the th feature, is the value of the th feature, is the total number of features.

[0238] Dynamically adjust the rejection threshold according to the batch size:

[0239] ;

[0240] For each element in the feature, if , then the feature will be regarded as unimportant and set to zero.

[0241] Classification prediction: After being processed by the Transformer encoder, the initially added class token Now it contains the global semantic information of the entire image. This vector will be passed as input to a simple multi-layer perceptron (MLP) classification head. The classification head usually consists of two fully connected layers, with a non-linear activation function (GELU) in the middle. The final output layer uses Softmax activation to generate the probability distribution of each class.

[0242] Loss calculation and backpropagation: The cross-entropy loss function is used to calculate the difference between the model prediction and the true label. The backpropagation algorithm calculates the gradient of the loss with respect to the model parameters and updates the model parameters using the Adam optimizer.

[0243] The present invention also provides an electronic device, including a processor and a memory communicatively connected to the processor and used for storing executable instructions of the processor. The memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is used to execute the intelligent processing method for liver echinococcosis ultrasound images based on deep learning described above.

[0244] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the intelligent processing method for liver echinococcosis ultrasound images based on deep learning described above.

[0245] By using the ADEN model to preprocess the collected data and then training the DFEV-VIT model based on the classification data set after image preprocessing, the present invention can greatly improve the diagnostic efficiency and accuracy, and reduce the missed diagnosis and misdiagnosis of echinococcosis.

[0246] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Intelligent processing method of ultrasound images of hepatic echinococcosis based on deep learning, characterized by: The following steps are involved: S1. Data collection: Collect liver ultrasound imaging data; S2, using an adaptive denoising and enhancement model to perform image preprocessing on the liver ultrasound image data collected in step S1, and classifying and labeling the preprocessed image data to establish a classification data set, labeling the images as cystic echinococcosis, alveolar echinococcosis, and other liver lesions; S3, based on the classification data set, train the optimized VIT classification model DFEV-VIT; In step S2, the liver ultrasound image data collected in step S1 is preprocessed using an adaptive denoising and enhancement model, including: S21, evaluating the signal-to-noise ratio and suppressing noise on the image; S22, performing adaptive wavelet denoising on the image processed in step S21; S23, performing edge-preserving filter and non-local mean denoising on the image processed in step S22; S24, dynamically expanding the grayscale of the image processed in step S23; S25, performing dynamic block processing on the image processed in step S24; S26, setting an inter-block overlap module for the image processed in step S25; S27, performing contrast-limited histogram equalization on the image processed in step S26; S28, reconstructing the image processed in step S27; S29, performing image post-processing and optimization on the image processed in step S28; In step S3, based on the classification data set, the classification model DFEV-VIT is trained, including: S31, performing image preprocessing and segmentation on the classification data set; S32, embedding image blocks into the data processed in step S31; S33, performing position coding and category marking on the data processed in step S32; In step S33, position coding and category labeling include: Set up a hybrid positional encoding model with learnable parameters and a periodic function: PE(x,y)=sin(f1(x))·cos(f2(y))+Learnable(x,y); Among them, PE(x,y) is a hybrid position encoding model function, (x,y) corresponds to the two-dimensional spatial coordinates of the image block, f1 and f2 are nonlinear mapping functions, mapping (x,y) to other dimensions, Learnable(x,y) is a bias composed of learning parameters, and the output is a value of the same dimension as the input coordinates; Adaptive spatial weights are introduced to adjust the position encoding influence of different image blocks based on information entropy: Among them, W spatial is the adaptive spatial weight, MLP (PE (x, y)) is the multi-layer perceptron output that processes position encoding; the original single category labeling is expanded to feature aggregation, feature aggregation is represented as CLS, and the multi-head self-attention mechanism of Vision Transformer is used to extract and map global features; S34, using a Transformer encoder to process the data processed in step S33; S35, using a feature discarding mechanism to process features of the data processed in step S34; In step S35, the feature rejection mechanism is used to process the feature, including: By calculating the embedding vector E(B i ), and use standard deviation and information entropy indicators to evaluate the contribution of each feature: Among them, w j is the weight of the jth feature, E j is the value of the jth feature, and m is the total number of features; Dynamically adjust the rejection threshold based on batch size: θ bq =k·mean(E(B i )); For each element in the feature, if w j <θ bq , then the feature will be considered unimportant and set to zero; S36, performing classification prediction on the data processed in step S35; S37, performing loss calculation and back propagation on the data processed in step S36; In step S31, image preprocessing and segmentation include: First, use the dynamic normalization method to perform mean normalization: Where I(x,y) is the pixel value of the input image at position (x,y), μ A and σ A are the mean and standard deviation of the local area A, σ A is a constant; The information entropy is used to evaluate the complexity of each image block and determine the priority of subsequent feature extraction. The information entropy expression of region A is: Where p(i) is the probability distribution of pixel values ​​in region A, and n is the number of possible pixel values; By evaluating the information entropy value, the size and shape of the image block can be dynamically selected. The image set is B = {B1, B2, ..., B k }, in this process, the block strategy follows the following principles: for information entropy H(B i ) is higher than the set value, select the image block that is smaller than the contrast threshold; for the information entropy H(B i ) is lower than the set value, select the image block with a contrast greater than the threshold value; In step S32, the image block is embedded, including: For each preprocessed and segmented image block B i , set up the embedding conversion module, which converts the original image block into a feature vector with semantic information through the attention mechanism: First, the input image block B is transformed into i Translated into query Q, key K and value V: Q=W q ·B i ,K=W k ·B,V=W v ·B; Calculate the similarity score by proposing the query Q with all keys K: Among them, d k is the dimension of the key; The scores are normalized by soft maximization to obtain the attention weights: α = softmax(score); Use the attention weights to perform a weighted summation on the value V to obtain the final feature representation: E(B i )=α·V.

2. The method for intelligent processing of ultrasonic images of hepatic echinococcosis based on deep learning according to claim 1, characterized in that: In step S21, the signal-to-noise ratio evaluation and noise suppression include: using the standard deviation method to evaluate the signal-to-noise ratio of the image, measuring the speckle noise and signal strength in the image; wherein, the signal strength calculation: calculating the average gray value of the image, the average gray value represents the strength of the signal, and the expression of the average gray value is: Where N is the total number of pixels in the area, I(x i ,y i ) is the gray value of the i-th pixel; Calculation of speckle noise intensity: Select an area of ​​1 / 10 the size of the current image in the upper right corner of the image as the background, and calculate the standard deviation of this area: Calculation of signal-to-noise ratio: Among them, NM is the average value of the area; In step S22, adaptive wavelet denoising includes: First, adaptive wavelet denoising is used to remove noise from the image with a signal-to-noise ratio lower than 20 in step S21; in this process, the image needs to be converted into the wavelet domain, and the portion corresponding to the noise in the image wavelet coefficients is removed by threshold processing; If the signal-to-noise ratio after denoising is higher than the set value, the current threshold is maintained; If the signal-to-noise ratio does not improve or decreases, the threshold value is dynamically adjusted based on the feedback; Finally, the denoised wavelet coefficients are converted back to the spatial domain using an inverse wavelet transform to generate an ultrasound image; in step S23, edge preserving filters and non-local mean denoising are performed, including: Apply non-local mean denoising to the image after wavelet denoising. For each pixel in the image to be processed, calculate the similarity by comparing it with other pixels in its neighborhood: Among them, I(i), I(j) are pixel values, h1 is the similarity control parameter, ||I(i)-I(j)|| 2 represents the difference of corresponding pixels, Z(i) is the normalization factor, and the pixels are weighted averaged according to the weights calculated by similarity to obtain the denoised pixel value: Among them, N(i) is the neighborhood set of pixel i; After the non-local mean denoising is completed, the denoised image is processed using an edge-preserving filter to clarify edge and texture features: Among them, I b (x) is the new grayscale value of the pixel at position x in the image after edge filtering, x, y are the coordinates of the processed pixel, G d ,G r are Gaussian functions in the spatial domain and the grayscale domain, W P is the normalization factor.

3. The method for intelligent processing of ultrasonic images of hepatic echinococcosis based on deep learning according to claim 2, characterized in that: In step S24, the grayscale is dynamically expanded, including: Dynamically expand the grayscale of the denoised image, use the histogram equalization method to redistribute the grayscale value of the image, and use the linear mapping rule to gradually expand the pixels concentrated in the set grayscale range to the entire grayscale range; calculate the grayscale histogram of the original image to determine the grayscale value distribution, find the minimum and maximum grayscale values ​​of the image, and convert them into a standard grayscale range through linear mapping, and apply the mapped values ​​to all pixels in the image; In step S25, dynamic block processing includes: Dynamically adjust the size and shape of image blocks according to the texture of the image, and use local gradient and gray-level co-occurrence matrix to calculate the texture features of each area of ​​the image; In step S26, setting an inter-block overlap module includes: An overlapping area is set between adjacent blocks, and the overlapping ratio ranges from 15% to 25%.

4. The method for intelligent processing of hepatic echinococcosis ultrasound images based on deep learning according to claim 3, characterized in that: In step S27, contrast-limited histogram equalization includes: firstly identifying features in the image using an edge detection algorithm by analyzing an edge map of the image; Apply clustering algorithms to edge maps to divide the image into different regions; The adaptive denoising and enhancement model calculates the grayscale histogram of each region and limits the uniform distribution of the histogram according to the set dynamic truncation threshold. The calculation formula of the histogram is: Where H(r) represents the number of pixels with gray level r, I(x,y) represents the gray value of the image at position (x,y), W represents the width of the image, H represents the height of the image, δ represents the Kronecker delta function, and r is the gray level, which is regarded as a constant; The histogram is then processed according to a dynamic cutoff threshold, and the histogram of each region is limited in the following way: Among them, N area represents the total number of pixels in the specific image area currently being processed, and T represents the frequency threshold; After completing the histogram equalization process, the adaptive denoising and enhancement model implements the regional merging operation, and the merging process is represented by the weighted average method: I merged (x,y)=α·I left (x,y)+(1-α)·I right (x,y); Among them, I left (x,y) and I right They are the equalized images of the left and right regions respectively, and the weight factor α ranges from 0 to 1.

5. The method for intelligent processing of ultrasonic images of hepatic echinococcosis based on deep learning according to claim 4, characterized in that: In step S28, image reconstruction includes: A gradient-guided interpolation method is used to eliminate artifacts produced by block processing; First, the gradient changes between adjacent blocks are analyzed, and then the edge-preserving interpolation method is used to combine the pixel values ​​of adjacent blocks through weights to preserve the edge features of the image during the reconstruction process; The processed blocks are merged into a complete ultrasound image; In step S29, image post-processing and optimization include: Apply non-local means denoising to find similar pixels in the same area based on similarity and remove the remaining noise: in, is the pixel value after denoising, (Z(i,j)=∑ k∈Ω w(i,j,k)) is the normalization coefficient, Ω represents all pixel indices in the search window, and w(i,j,k) represents the weight, which is calculated by the similarity metric h is a smoothing parameter that controls the influence range of the weight; Apply an adaptive sharpening filter to sharpen only the edge areas of the image. The adaptive sharpening of the image is expressed as: S(i,j)=I(i,j)+λ·G(i,j); Among them, S(i,j) represents the pixel value after sharpening, G(i,j) represents the gradient value calculated at point I(i,j), and λ represents the parameter controlling the sharpening intensity; Edge detection is achieved by calculating the gradient of pixels. The gradient expression of pixels is:

6. The method for intelligent processing of hepatic echinococcosis ultrasound images based on deep learning according to claim 1, characterized in that: In step S34, the Transformer encoder processes, including: The Transformer encoder consists of multiple stacked encoder layers, each of which consists of two key modules: a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism allows each image block to establish an interactive relationship with all other image blocks, adding an interactive bidirectional mode to the multi-head self-attention mechanism: Attention forward (X)=Attention(Q,K,V); Attention backward (X)=Attention(K,Q,V); The reverse attention mechanism swaps the roles of key and query; Introducing the attention gating mechanism: GatedX)=sigmoid(W g X+b g )⊙X; Among them, W g and b g is a learnable parameter, sigmoid activation function, ⊙ represents element-by-element multiplication; through the gating mechanism, the multi-head self-attention mechanism can adaptively choose to retain or suppress information; The feedforward neural network performs a nonlinear transformation on the output of the multi-head self-attention mechanism, and the output of the multi-head self-attention mechanism is fed into the feedforward neural network: Y = FFN(Attention(X)); The feedforward neural network consists of multiple feedforward layers and activation functions. The operation of each feedforward layer is expressed as: <h2 style=";text-align:left;direction:ltr">Y = ReLU(W)<h2 style=";text-align:left;direction:ltr"> 1Attention <h2 style=";text-align:left;direction:ltr"> (X)+b1)W2+b2; Among them, W1, W2 are weight matrices, b1, b2 are bias terms; Each encoder layer also includes layer normalization and residual connections. Layer normalization ensures that the input of each layer maintains statistical properties: Among them, μ and σ are the mean and standard deviation respectively, and γ and β are learnable parameters; Residual connection is to skip the input to the output.

7. The method for intelligent processing of ultrasonic images of hepatic echinococcosis based on deep learning according to claim 6, characterized in that: In step S36, classification prediction includes: After being processed by the Transformer encoder, the added category tag CLS contains the global semantic information of the entire image. The added category tag CLS is passed as input to the multi-head attention mechanism. The multi-head attention mechanism includes two fully connected layers. A nonlinear activation function is used between the two fully connected layers. The final output layer uses Softmax activation to generate the probability distribution of each category. In step S37, loss calculation and back propagation include: The difference between the model prediction and the true label is calculated using the cross entropy loss function, the backpropagation algorithm calculates the gradient of the loss with respect to the model parameters, and the optimizer is used to update the model parameters.

8. An electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, characterized in that: The memory stores instructions that can be executed by the processor, and the instructions are executed by the processor. The processor is used to execute the deep learning-based intelligent processing method for hepatic echinococcosis ultrasound images as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for intelligent processing of hepatic echinococcosis ultrasound images based on deep learning as described in any one of claims 1 to 7 is implemented.