A Feature Classification Image Processing Method and System Based on the ShrinkNet3D Network
Through the adaptive calibration and attention mechanism of ShrinkNet3D network, the impact of patch file size on feature classification is solved, and the accuracy and confidence of feature classification are improved.
Patent Information
- Application Number
- CN202111476632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In the prior art, patch files of larger sizes are prone to contain multiple features and are difficult to process, and patch files of smaller sizes have low feature extraction accuracy and confidence, resulting in a decrease in feature classification accuracy and confidence.
ShrinkNet3D network is adopted to generate effective features through adaptive calibration convolution, use attention mechanism to zero unimportant features, reduce the weight of easy-to-classify samples, and control models focus on difficult-to-classify samples.
The confidence in feature classification of small-size patch files is improved, and the ability of deep neural networks to extract useful features from noise-containing signals is enhanced, and the accuracy and confidence problems of feature classification are solved.
Smart Images

Figure CN114299276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the field name, and particularly relates to a feature classification image processing method and system based on the ShrinkNet3D network. Background Art
[0002] The ShrinkNet3D network, namely the 3D shrinkage network, and the deep residual shrinkage network is an artificial intelligence algorithm, which is actually a new improvement of the deep residual network. The soft thresholding is introduced as a non-linear layer into the network structure of ResNet, and its purpose is to improve the feature learning effect of deep learning methods on noisy data or complex data.
[0003] In the process of feature classification image processing in the prior art, the input image is segmented to obtain image patches (patch files) of different sizes, and then feature extraction and feature classification are respectively performed on each patch file.
[0004] However, in the process of implementing the technical solution of the present invention in the embodiments of the present application, the inventors of the present application found that the above technology has at least the following technical problems:
[0005] Patch files with larger sizes are prone to contain multiple features and are difficult to process; the smaller the size of the patch file, the fewer features can be extracted, and thus the accuracy of feature extraction and the confidence of feature classification are also higher. However, the features contained in too small patch files are incomplete or the feature distance from adjacent patch files is too close, resulting in a decrease in accuracy and confidence and affecting subsequent feature classification operations.
[0006] Based on this, the present invention designs a feature classification image processing method and system based on the ShrinkNet3D network to solve the above problems. Summary of the Invention
[0007] In order to solve the technical problems mentioned in the current background art, the purpose of the present invention is to provide a feature classification image processing method and system based on the ShrinkNet3D network.
[0008] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0009] A feature classification image processing method based on the ShrinkNet3D network includes the following steps:
[0010] Step 1: Establish an adaptive calibration of long-distance spatial and inter-channel dependencies for each spatial position to enable the convolutional neural network to generate effective features;
[0011] Step 2: Use the attention mechanism to set unimportant features to zero through a soft threshold function, and enhance the neural network to extract useful features from the remaining important features through a signal with noise.
[0012] Step 3: Reduce the weights of easily classifiable samples and control the model training to focus on more difficult-to-classify samples.
[0013] Preferably, the adaptive calibration includes:
[0014] Split the input features of the original dimension H×W×M×C into features X1 and X2 of the new dimension H×W×M×C / 2, and divide the convolution kernel of the original dimension H×W×M×C into K1, K2, K3, and K4 during the convolution operation, so that their dimensions are all the new dimension H×W×M×C / 2 after splitting, and perform feature transformation in two different scale spaces.
[0015] Preferably, the feature transformation in two different scale spaces includes:
[0016] Perform average pooling downsampling on the feature X1 in the new dimension and then perform upsampling. Apply the upsampling result to the calibrated weight. After passing through the Sigmoid function, multiply the weight element-wise with the feature extracted by K3 convolution. Finally, obtain the output feature Y1 through K4 convolution, and then concatenate the output feature Y1 with the feature Y2 obtained by K1 convolution in the original feature space to obtain the final output feature Y.
[0017] Preferably, the upsampling is performed by the bilinear interpolation method.
[0018] Preferably, the processing of the soft threshold function includes:
[0019] The deep neural network automatically determines the selection of the threshold, sets the features close to zero to zero, and retains useful negative features.
[0020] Preferably, the automatic determination of the selection of the threshold includes:
[0021] Establish a sub-network for deep residual processing, and set the threshold of the sub-network to the product of the average value of the absolute value of the feature map and the coefficient α, where, under the action of the Sigmoid function, α ∈ (0,1).
[0022] Preferably, the optimal determination of the threshold includes:
[0023] Perform several convolution operations to reduce the size of the image patch to generate more channels, and repeatedly use the deep residual module multiple times to obtain the optimal threshold.
[0024] Preferably, the reduction of the weights of easily classifiable samples includes:
[0025] Introduce a regulation factor -(1 - p i ) γ to the cross - entropy loss function. If the value of this regulation factor is less than 1, it will play a role in attenuation. Among them, when the value of p i is close to 1, the confidence of the model is higher, and at the same time, the degree of attenuation is greater.
[0026] An image processing system for feature classification based on the ShrinkNet3D network, which is characterized by including:
[0027] A self - calibration convolution module, which is used to control the adaptive calibration for establishing long - distance spatial and channel dependencies at each spatial position, so that the convolutional neural network can obtain effective features;
[0028] A residual shrinkage module, which is used to set unimportant features to zero through a soft - threshold function, and control and strengthen the neural network to extract useful features from the signals containing noise for the retained important features;
[0029] A model training module, which is used to reduce the weights of easily classified samples and control the model training to focus on more difficult - to - classify samples.
[0030] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:
[0031] 1. By using self - calibration convolution operation, the present invention not only ensures the architecture of the original model without introducing additional parameters, but also enables the convolutional neural network to generate more effective features;
[0032] 2. By using the attention mechanism to process features during the residual value shrinkage process, the present invention strengthens the ability of the deep neural network to extract useful features from signals containing noise;
[0033] 3. By reducing the weights of easily classified samples, the present invention realizes that the training model focuses more on processing difficult - to - classify samples, ensuring the data balance of multiple labels in the dataset;
[0034] In summary, the present invention is particularly suitable for feature classification of small - size patch files and can effectively improve the confidence of feature classification. Brief Description of the Drawings
[0035] The following further details the present invention in conjunction with the drawings and specific embodiments:
[0036] Figure 1 It is the 3D neural network structure diagram of the present invention;
[0037] Figure 2 It is the self - calibration convolution flow chart of the present invention;
[0038] Figure 3 This is the flowchart of the deep residual processing of the present invention. Specific implementation manners
[0039] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0040] Embodiment 1
[0041] As Figure 1 shown, wherein, taking the features extracted from the patch file of the image as the input, the SoftMax layer finally outputs the result of classifying the features of the patch file. Among them, SC Convolution represents the self-calibration convolution module (convolution), Shrinkage block represents the residual shrinkage module, and FC represents the linear transformation layer.
[0042] The present invention provides a technical solution: a feature classification image processing method based on the ShrinkNet3D network, including the following steps:
[0043] Step 1: Establish an adaptive calibration for long-distance spatial and inter-channel dependencies at each spatial position, so that the convolutional neural network generates effective features;
[0044] Step 2: Use the attention mechanism to set unimportant features to zero through the soft threshold function, and strengthen the neural network to extract useful features from the signals containing noise for the remaining important features;
[0045] Step 3: Reduce the weights of easily classifiable samples and control the model training to focus on more difficult-to-classify samples.
[0046] It is not difficult to find through the above steps that during the process of feature classification image processing, since the accuracy and confidence of features are crucial, the features contained in too small patch files have problems such as incompleteness or too close distance to the features of adjacent patch files, resulting in a decrease in accuracy and confidence, which affects the subsequent
[0047] feature classification operations. Therefore, the present application provides a classification processing for the features of small-sized patch files to achieve a method for effectively improving the confidence of feature classification.
[0048] It should be noted that in Step 1, it can be seen that this step is an improvement of self-calibrated convolution for the feature transformation process of basic convolution. Without introducing additional parameters, it ensures the architecture of the original model and effectively enhances the standard convolution. In the conventional convolution function, since the size of the convolution kernel is fixed, it directly determines that a larger receptive field cannot be obtained during the feature transformation process. Therefore, the network composed of stacked convolution layers cannot capture higher-level semantics.
[0049] It is also worth noting that in Step 2, it is an improvement of ResNet (Residual Neural Network). In essence, through the attention mechanism, unimportant features are set to zero by the soft threshold function, and the noticed important features are retained, ultimately enhancing the ability of the deep neural network to extract useful features from signals containing noise.
[0050] Furthermore, in Step 3 above, Focal Loss is obtained by improving the standard cross-entropy loss function. By reducing the weights of easy-to-classify samples, the model focuses more on difficult-to-classify samples during training, and it has a good effect on the data imbalance problem of multiple labels in the dataset.
[0051] As Figure 2 shown, where (element-wise summation, the scalar sum (dot addition) of vectors, that is, the elements of each dimension of the vector are added separately as the corresponding dimension elements of the new vector; element-wise multiplication, the scalar product (dot multiplication) of vectors, that is, the elements of each dimension of the vector are multiplied separately as the corresponding dimension elements of the new vector; down-sampling, downsampling; up-sampling, upsampling).
[0052] The adaptive calibration includes:
[0053] The input feature with the original dimension H×W×M×C is split into features X1 and X2 with the new dimension H×W×M×C / 2, and the convolution kernel with the original dimension H×W×M×C is divided into K1, K2, K3, and K4 during the convolution operation, so that their dimensions are all the new dimension H×W×M×C / 2 after splitting, and feature transformation is performed in two different scale spaces.
[0054] Furthermore, the feature transformation in two different scale spaces includes:
[0055] After performing average pooling downsampling on the feature X1 in the new dimension, upsampling is carried out, and the upsampling result is applied to the calibrated weights. After passing through the Sigmoid function, the weights are multiplied element-wise with the features extracted by the K3 convolution. Finally, the output feature Y1 is obtained through the K4 convolution. Then, the output feature Y1 is concatenated with the feature Y2 obtained by the K1 convolution in the original feature space to obtain the final output feature Y.
[0056] Preferably, the upsampling is performed by the bilinear interpolation method.
[0057] In the above embodiment, the input feature in the form of H×W×M×C is split into features X1 and X2 with a size of H×W×M×C / 2 through self-calibrated convolution. During the convolution operation, the convolution kernel with a dimension of H×W×M×C is also correspondingly divided into four parts, each with a dimension of H×W×M×C / 2. Thereafter, feature transformation is performed in two different scale spaces.
[0058] And when outputting features, the formula Y1 = ((K3×X1)·σ(Up(AvgPool(X1)×K2)+X1))×K4 is adopted; where AvgPool and Up respectively correspond to the downsampling and upsampling operations, σ is the Sigmoid function, and K1, K2, K3, and K4 are four sub-convolution kernels with the same dimension divided from the convolution kernel. In the scale space where the self-calibrated convolution is located, first, average pooling downsampling is performed on the feature X1, then upsampling is carried out by the bilinear interpolation method, and the upsampling result is used to form the weights for calibration. After passing through the Sigmoid function, these weights are multiplied element-wise with the features extracted by the K3 convolution, and finally, the output feature Y1 is obtained through the K4 convolution. Thereafter, it is concatenated with the feature Y2 obtained by the K1 convolution in the original feature space to obtain the final output feature Y.
[0059] More specifically, the processing of the soft threshold function includes:
[0060] The deep neural network automatically determines the selection of the threshold, sets the features close to zero to zero, and retains the useful negative features.
[0061] In this embodiment, it can be seen that in many signal denoising methods, the soft threshold is often used as a key step. In this embodiment, the soft threshold function is embedded into the deep residual shrinkage network as a non-linear transformation layer, and its function is as follows: where x is the input feature and T is the threshold.
[0062] It should be noted that what the soft thresholding achieves is not to set all negative features to zero, but to set the features close to zero to zero, so that useful negative features can be retained, thereby achieving the denoising effect. And the selection of the threshold is automatically determined by the deep neural network.
[0063] As shown in Figure 3 Figure 3 , where (thresholding), the selection of the automatically determined threshold includes:
[0064] A sub-network for depth residual processing is established, and the threshold of the sub-network is the product of the average value of the absolute value of the feature map and the coefficient α. Under the action of the Sigmoid function, α ∈ (0, 1).
[0065] Furthermore, the optimal determination of the threshold includes:
[0066] Perform a number of convolution operations to reduce the size of the image patch to generate more channels, and repeatedly use the depth residual module multiple times to obtain the optimal threshold.
[0067] It is not difficult to find through the above embodiments that there is a small sub-network in this module, and its function is to adaptively set the threshold. The threshold set by this sub-network is actually the product of the average value of the absolute value of the feature map and the coefficient α. Under the action of the Sigmoid function, α is a number between 0 and 1. This method ensures that the threshold will be a relatively small positive number, so that the output will not all be zero. This module corresponds to a shrinkage threshold for each channel of the input feature, so the final threshold is obtained in the form of a vector. Using this adaptive threshold acquisition method, the depth residual module can effectively denoise the input features, thereby improving the feature extraction ability of the network.
[0068] Moreover, when constructing the ShrinkNet3D neural network based on the depth shrinkage module, first perform a number of convolution operations to reduce the size of the image patch, thereby generating more channels, and then repeatedly use the depth residual module multiple times to obtain the optimal threshold.
[0069] It should also be noted that the reduction of the weight of easily classified samples includes:
[0070] Introduce a regulation factor -(1 - p i ) γ to the cross-entropy loss function. If the value of this regulation factor is less than 1, it will play a decaying role. Among them, when the value of p i is close to 1, the confidence of the model is higher, and at the same time, the degree of attenuation is greater.
[0071] In this embodiment, the traditional cross-entropy loss function is as shown in the formula: And the Focal Loss is as shown in the formula where N represents the number of classification categories, y is the label. When the current classification is i, yi is 1, otherwise yi is 0. pi represents the probability that the neural network outputs category i.
[0072] It can be obtained from the formula that, compared with the cross-entropy loss function, Focal Loss adds a pre-term in front of it, -(1 - p i ) γ is the introduced adjustment factor, and this value is less than 1, which can play a role in attenuation. When the value of p i is close to 1, it means that the confidence of the model is higher. At this time, the degree of attenuation will be greater, so as to achieve the purpose of reducing the weight of easy-to-classify samples.
[0073] Embodiment 2
[0074] The present invention also provides a feature classification image processing system based on the ShrinkNet3D network, including:
[0075] A self-calibrating convolution module, which is used to control the adaptive calibration for establishing long-distance spatial and channel dependencies at each spatial position, so that the convolutional neural network can obtain effective features;
[0076] A residual shrinkage module, which is used to set unimportant features to zero through a soft threshold function, and control the neural network to extract useful features from the retained important features through signals containing noise; and
[0077] A model training module, which is used to reduce the weight of easy-to-classify samples and control the model training to focus on more difficult-to-classify samples.
[0078] The above embodiments only illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A feature classification image processing method based on the ShrinkNet3D network, characterized in that It includes the following steps: Step 1: Establish adaptive calibration of long-distance spatial and inter-channel dependencies for each spatial position to enable the convolutional neural network to generate effective features. Among them, the adaptive calibration includes: splitting the input features of the original dimension H×W×M×C into features X1 and X2 of the new dimension H×W×M×C / 2, and dividing the convolutional kernel of the original dimension H×W×M×C into K1, K2, K3, and K4 during the convolutional operation, so that their dimensions are all the new dimension H×W×M×C / 2 after splitting, and performing feature transformation in two different-scale spaces. The feature transformation in two different-scale spaces includes: performing average pooling downsampling on the feature X1 in the new dimension and then performing upsampling, applying the upsampling result to the calibrated weight, and after passing through the Sigmoid function, multiplying the weight element-wise with the feature extracted by the convolution of K3, and finally obtaining the output feature Y1 through the convolution of K4. Then, splicing the output feature Y1 with the feature Y2 obtained by the convolution of K1 in the original feature space to obtain the final output feature Y. The upsampling is performed by the bilinear interpolation method. Step 2: Use the attention mechanism to set unimportant features to zero through the soft threshold function, and strengthen the neural network to extract useful features from the remaining important features with noisy signals. Among them, the processing of the soft threshold function includes: the deep neural network automatically determines the selection of the threshold, sets the features close to zero to zero, and retains useful negative features. The automatic determination of the selection of the threshold includes: establishing a sub-network for deep residual processing, and setting the threshold of the sub-network to the product of the average value of the absolute value of the feature map and the coefficient α, where, under the action of the Sigmoid function, α∈(0,1). The optimal determination of the threshold includes: performing several convolutional operations to reduce the size of the image patch to generate more channels, and repeatedly using the deep residual module multiple times to obtain the optimal threshold. Step 3: Reduce the weight of easily classifiable samples to control the model training to focus on more difficult-to-classify samples. Among them, reducing the weight of easily classifiable samples includes: introducing a regulation factor -(1 - p i ) γ to the cross-entropy loss function. If the value of this regulation factor is less than 1, it plays a decaying role. Among them, i represents the current category, and p i represents the probability that the neural network outputs category i. When the value of p i is close to 1, the confidence of the model is higher, and the degree of decay is greater.
2. A feature classification image processing system applied to the feature classification image processing method based on the ShrinkNet3D network described in claim 1, characterized in that, It includes: Self-calibration convolution module, which is used to control the adaptive calibration of establishing long-distance spatial and channel dependencies for each spatial position, enabling the convolutional neural network to obtain effective features; Residual shrinkage module, which is used to set unimportant features to zero through the soft threshold function, and control and strengthen the neural network to extract useful features from the remaining important features with noisy signals; Model training module, which is used to reduce the weight of easily classified samples and control the model training to focus on more difficult-to-classify samples.
Citation Information
Patent Citations
Rock slice image recognition method based on residual shrinkage module and attention mechanism
CN113486929A
Voiceprint recognition method based on short voice
CN113488058A