Construction Method of Bidirectional Adaptive Reweighted Neural Network for Image Processing
By building a bidirectional adaptive reweighting neural network, the cross-regional and multi-scale reweighting module combined with attention mechanism is used to solve the dependence of image classification tasks on the data set construction time, and efficient image classification and feature extraction are achieved.
Patent Information
- Application Number
- CN202510378986.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the prior art, image classification tasks rely too much on a large amount of standard data, resulting in a large amount of time spent on data set construction, thereby reducing the efficiency of image classification tasks.
A two-way adaptive reweighting neural network is built, and image features are reweighted and fused through cross-region reweighting modules and adaptive multi-scale global to local reweighting modules. Combined with the attention mechanism, image features are reweighted and fused, reducing dependence on labeled data, and image processing results are obtained through the fully connected layer.
This method can effectively balance global and local information without relying on region of interest (ROI) annotation or hyperparameter calibration, improving the accuracy and efficiency of image classification, reducing information loss and high-value deviation.
Smart Images

Figure QLYQS_14 
Figure QLYQS_26 
Figure QLYQS_33
Abstract
Description
Technical Field
[0001] The present invention discloses a construction method of a bidirectional adaptive reweighted neural network for image processing, belonging to the technical field of neural networks. Background Art
[0002] In the field of computer vision, traditional image processing methods may not be able to capture the subtle but crucial details required accurately. The emergence of deep learning technology has opened up new ways to improve the accuracy of image classification. Deep neural networks (DNNs) have demonstrated excellent capabilities in directly learning hierarchical features from data, reducing the dependence on manual feature engineering. However, these methods require a large amount of labeled data, including bounding boxes or segmentation information, and the creation of such datasets usually requires a large amount of manpower. The progress of imaging technology has led to the generation of a large number of datasets based on film and digital images, and there are significant differences between these datasets on different devices. These differences pose challenges to training high-quality and generalizable models. In addition, most existing models adopt rigid segmentation in multi-instance learning, which may cause damage and fragmentation of information in image classification. These models also heavily rely on hyperparameters, and the optimal hyperparameters vary for different datasets, which further hinders the generalization process of the models. Summary of the Invention
[0003] The purpose of the present invention is to provide a construction method of a bidirectional adaptive reweighted neural network for image processing, so as to solve the problem in the prior art that the image classification task overly relies on a large amount of standard data, resulting in a large consumption of time for dataset construction and thus reducing the efficiency of the image classification task.
[0004] In the construction method of the bidirectional adaptive reweighted neural network for image processing, the image to be processed is input into a convolutional neural network feature extractor to obtain image features , and is input into the bidirectional adaptive reweighted neural network. In the bidirectional adaptive reweighted neural network is divided into two branches. One branch is input into the cross-region reweighting module to obtain an output , and one branch is input into the adaptive multi-scale global-to-local reweighting module to obtain an output . Combining the attention mechanism, and are reduced to a vector of a fixed size, and the image processing result is obtained through a fully connected layer.
[0005] The cross-region reweighting module includes a region attention module and a cross-region attention module.
[0006] The region attention module compresses the dimension of to 512 to obtain image features , and It is divided into 64 blocks, and each block is locally reweighted using the two-dimensional selective state space model SS2D to obtain image features , and take as the output of the regional attention module .
[0007] SS2D includes:
[0008] ;
[0009] In the formula, represents the fusion operation, represents the module operation in the two-dimensional selective state space model, , , , respectively represent flattening along four directions.
[0010] The module includes that the input of the module generates a matrix and the time step parameter , and the hidden state is updated as:
[0011] ;
[0012] In the formula, is the natural constant, is the input of the module, is the hidden state at time
[0013] The output is:
[0014] ;
[0015] Arrange the of all time steps to obtain the output of the module.
[0016] The cross-region attention module divides the input into two branches. One branch combines the adjustable parameter to perform matrix multiplication, and then splits it into three inputs , and performs weight calculation:
[0017] ;
[0018] ;
[0019] ;
[0020] In the formula, is the number of single-dimensional cuts, is the total side length of the block, is the th side length of the block, , , are the weights of the three input corresponding outputs respectively, represents the th key, represents the total number of keys, , are activation functions.
[0021] The second branch of the cross-region attention module and obtain the output after matrix multiplication and then input into the SS2D module. Multiply , , to get .
[0022] The adaptive multi-scale global-to-local reweighting module performs max pooling on , and then inputs into the global-scale module, quarter-scale module, and sixteenth-scale module simultaneously, and combines different scale modules for weighted processing:
[0023] ;
[0024] In the formula, , is the scale module, including the global-scale module , quarter-scale module , and sixteenth-scale module . includes the weighted final result of includes the weighted final result of includes the weighted final result of represents the weight, represents the weighted final result, represents element-wise multiplication, including unifying the size of the output information of with weight information, then performing embedding operations separately, and finally performing feature fusion to obtain ;
[0025] Process by combining with 1×1 convolution, and , , Perform 1×1 convolution processing on each of the three branches, add two of the three convolution results through matrix addition, and then add the addition result and the third convolution result through matrix addition to obtain :
[0026] ;
[0027] In the formula, is an adjustable parameter.
[0028] Combined with the attention mechanism, reducing and to a vector of a fixed size includes:
[0029] ;
[0030] In the formula, represents or , is a fully connected layer, is the parameter of the fully connected layer, represents the product of the length and width, represents the number of channels, represents when it represents and when it represents represents matrix multiplication, represents the output of the attention mechanism;
[0031] Input into the fully connected layer:
[0032] ;
[0033] In the formula, is or 's predicted probability, is the activation function, and are the weight and bias of the linear regression layer respectively;
[0034] Obtain the image processing result :
[0035] ;
[0036] In the formula, is 's predicted probability, is 's predicted probability.
[0037] Loss Function of Bidirectional Adaptive Reweighted Neural Network is as follows:
[0038] ;
[0039] In the formula, is the binary result of image processing, , The two values of correspond to two binary results, = 1, is the architecture parameter of the bidirectional adaptive reweighted neural network, , represents the regularization term, is the weight used to alleviate sample imbalance.
[0040] Compared with the prior art, the present invention has the following beneficial effects: The present invention can well balance global and local information, without relying on Region of Interest (ROI) annotation, nor requiring calibration of hyperparameters; it integrates large-scale information to evaluate small-scale regions and effectively highlights detailed information, thereby improving the classification effect; it adopts a cross-region module to integrate information from different regions, reducing the loss of image information and the dependence on hyperparameters that may be caused by rigid segmentation, and at the same time avoiding the high-value deviation that may be brought by global attention. Specific Embodiments
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0042] A construction method of a bidirectional adaptive reweighted neural network for image processing, which inputs the image to be processed into a convolutional neural network feature extractor to obtain image features , and inputs into the bidirectional adaptive reweighted neural network. In the bidirectional adaptive reweighted neural network is divided into two branches. One branch inputs into the cross-region reweighting module to obtain the output , and the other branch inputs into the adaptive multi-scale global-to-local reweighting module to obtain the output . Combining the attention mechanism, and are reduced to a vector of a fixed size, and the image processing result is obtained through a fully connected layer.
[0043] The cross-region reweighting module includes a region attention module and a cross-region attention module.
[0044] The region attention module compresses the dimension of to 512 to obtain image features , divides into 64 blocks, and each block uses a two-dimensional selective state space model SS2D for local reweighting to obtain image features , and takes as the output of the region attention module .
[0045] SS2D includes:
[0046] ;
[0047] In the formula, represents the fusion operation, represents the module operation in the two-dimensional selective state space model, , , , respectively represent flattening along four directions.
[0048] The module's input undergoes a linear transformation to generate a matrix and a time step parameter , and the hidden state is updated as:
[0049] ;
[0050] In the formula, is the natural constant, is the module's input, is the hidden state at time
[0051] The output is:
[0052] ;
[0053] Arrange the of all time steps to obtain the module output.
[0054] The cross-region attention module divides the input into two branches. One branch combines the adjustable parameter for matrix multiplication, and then splits it into three inputs , and performs weight calculation:
[0055] ;
[0056] ;
[0057] ;
[0058] In the formula, is the number of single-dimensional cuts, is the total side length of the block, is the th side length of the block, , , are the weights of the three input corresponding outputs respectively, represents the th key, represents the total number of keys, , are activation functions.
[0059] The second branch of the cross-region attention module and are input into the SS2D module after matrix multiplication to obtain the output . , , are multiplied to obtain .
[0060] The adaptive multi-scale global-to-local reweighting module performs max pooling on , and then inputs into the global-scale module, quarter-scale module, and sixteenth-scale module simultaneously, and combines different scale modules for weighted processing:
[0061] ;
[0062] In the formula, , is the scale module, including the global-scale module , quarter-scale module , sixteenth-scale module , includes the weighted final result of, includes the weighted final result of, includes the weighted final result of, represents the weight, represents the weighted final result, represents element-wise multiplication, including The output information is unified in size, then embedded operations are performed separately, and finally feature fusion is carried out to obtain ;
[0063] For Combined with 1×1 convolution for processing, and , , Are used as three branches to perform 1×1 convolution processing respectively. Matrix addition is performed on two of the three convolution processing results, and then matrix addition is performed on the addition result and the third convolution processing result to obtain :
[0064] ;
[0065] In the formula, Is an adjustable parameter.
[0066] Combined with the attention mechanism to reduce And To a vector of fixed size, including:
[0067] ;
[0068] In the formula, Represents Or , Is a fully connected layer, Is the parameter of the fully connected layer, Represents the product of length and width, Represents the number of channels, When it represents , When it represents , Represents matrix multiplication, Represents the output of the attention mechanism;
[0069] Input Into the fully connected layer:
[0070] ;
[0071] In the formula, Is Or 's predicted probability, Is the activation function, And Are the weight and bias of the linear regression layer respectively;
[0072] Obtain the image processing result :
[0073] ;
[0074] In the formula, is the predicted probability of is the predicted probability of
[0075] The loss function of the bidirectional adaptive reweighted neural network is:
[0076] ;
[0077] In the formula, is the binary result of image processing, , The two values of correspond to two binary results, = 1, is the architecture parameter of the bidirectional adaptive reweighted neural network, , represents the regularization term, is the weight used to alleviate sample imbalance.
[0078] The present invention uses two widely recognized datasets, CBIS-DDSM and INbreast, and an in-house dataset to evaluate the performance of the BARNet (the model of the present invention) method. The CBIS-DDSM dataset contains 3,071 mammograms, which have been carefully annotated for classification and segmentation, covering various pathological types such as masses and calcifications. The INbreast dataset contains 410 digital mammograms, which are from multiple angles of 115 different cases and have been detailedly classified and segmented. It covers not only masses and calcifications, but also various lesion types such as asymmetry and distortion, and is accompanied by comprehensive annotation information. The medical records from February 2021 to October 2022 were carefully evaluated to identify patients with histologically confirmed breast cancer and benign tumors. This dataset contains full-field mammography images of 214 patients, totaling 406, including 113 benign tumors and 293 malignant tumors.
[0079] Table 1 presents the data distributions of the CBIS-DDSM mammogram dataset and the internal dataset. Regarding the CBIS-DDSM dataset, referring to its original data source, it was found that some images contain both tumors and calcifications, resulting in their simultaneous appearance in the tumor training set and the calcification dataset. Therefore, all duplicate images were removed from the database, and the dataset was randomly divided into five equal parts, with a training set to test set ratio of 4:1, and the cross-validation technique was used. The positive to negative sample ratio of the INbreast dataset is 3.1:1, showing a significant imbalance. To address this issue, the INbreast dataset was also randomly divided into five equal parts, with a training set to test set ratio of 4:1, and the cross-validation method was used. For the internal dataset, the same processing strategy as the INbreast dataset was adopted.
[0080] Table 1. Data Distributions of the CBIS-DDSM Mammogram Dataset and the Internal Dataset
[0081] ;
[0082] The present invention uses deep learning techniques and the Py-Torch framework to process medical image data. For the INbreast dataset, the Otsu segmentation algorithm was used to identify and crop the breast region, removing unnecessary backgrounds, and the images were resized to 800×800. The internal dataset was also processed similarly and resized to 448×448. For the CBIS-DDSM dataset containing artifacts, the images were resized to 448×448, binarized, and masked with the largest connected component to retain the largest connected region. All images were normalized by dividing the pixel values by 255 to map them to the [0,1] interval. To prevent overfitting, in each training epoch, the images were randomly adjusted and rotated, with the rotation angle ranging from -25 degrees to +25 degrees.
[0083] The present invention is implemented in PyTorch, and the experiments are performed on Tesla V100 GPUs. The DenseNet-169 architecture based on the code provided by the torchvision library is adopted. Some minor adjustments are made, replacing the final classification layer with a multi-scale module that can extract features from global to local perspectives and a position embedding module from local to global, and finally forming a fixed-length linear regression layer for output. The initialization of the backbone network weights uses the pre-trained weights from ImageNet to ensure a fair comparison with existing research. In the experiments using the INbreast and CBIS-DDSM datasets, the CNN is also initialized with the pre-trained weights from ImageNet and updated with a smaller learning rate. These pre-trained weights can be obtained from the public packages provided by the torchvision library. The model training is carried out using the Adam optimization algorithm. The learning rate of the logistic regression layer is initialized to , while the learning rate of the convolutional neural network (CNN) layer is set to . A strategy of learning rate decay of 0.98 every 10 epochs is also implemented, and the loss function hyperparameter is set to . The model training stops after 150 epochs, and the weights are retained for testing. Given that the number of training epochs required for different datasets varies, to prevent overfitting caused by overtraining, the model with the highest accuracy is saved during the training process.
[0084] Table 2 shows the details of the experimental data obtained by integrating the adaptive multi-scale global-to-local reweighting module AMG2L and the cross-region reweighting module L2GCR into the baseline convolutional neural network (CNN) for standard mammogram classification. This table evaluates the impact of these two modules on the model performance from a quantitative perspective. Specifically, the accuracy and AUC values listed in Table 2 are the average results of five independent tests, providing statistical robustness for the experimental results. These values not only reflect the overall performance of the model in the classification task but also reveal the specific contributions of AMG2L and L2GCR to the model performance by comparing the performance before and after the introduction of these two modules.
[0085] Table 2. Accuracy and AUC values of different baseline models combined with AMG2L and L2GCR on the INBREAST dataset
[0086] ;
[0087] In Table 2, Alexnet, Resnet-34, Resnet-50, VGG19, and DenseNet-169 are deep learning models of the prior art. In these experiments, the average improvement value measures the performance difference between the baseline model with AMG2L and L2GCR introduced and the original model. This metric reflects the improvement in model performance macroscopically and indicates the positive impact of integrating AMG2L and L2GCR on model performance. At the same time, the maximum improvement value represents the largest performance improvement recorded in these five tests, revealing the potential performance gain that AMG2L and L2GCR can bring to the model under optimal conditions. The experimental results show that integrating AMG2L and L2GCR into the baseline CNN model significantly improves the AUC and accuracy. This finding is significant because it not only confirms the effectiveness of AMG2L and L2GCR in enhancing model performance but also shows that their addition helps the model identify lesion areas more accurately. This is crucial for improving the accuracy and reliability of mammogram classification, especially in the detection and diagnosis of early lesions. In addition, the increase in parameters and computational cost brought about by these improvements can be ignored. This is particularly noteworthy because computational resources and model complexity are often important considerations in practical applications. The efficiency of AMG2L and L2GCR means that they can bring significant performance improvements to the model without significantly increasing the computational burden.
[0088] In the present invention, the performance of BARNet is compared with existing benchmark methods on three datasets. Traditional models usually rely on a multi-stage training process or require manually annotated data, and need to annotate edge information or regions of interest (ROIs) during the training and testing processes, which results in a high training cost. Recent studies have explored the application of convolutional neural networks (CNNs) in the field of mammogram classification, and these networks have shown effectiveness in other fields. However, these methods often ignore the specific features of mammogram images or rely on hyperparameters. The experimental results show that the accuracy and AUC of BARNet are close to the optimal results, and it has extremely high specificity, which means that BARNet performs excellently in reducing false positives (wrongly identifying non-lesion areas as lesion areas). BARNet achieves full automation without manual annotation or hyperparameter tuning, and can directly extract classification labels from case results for training, thus realizing an end-to-end automated training process. To ensure the fairness of the experiment, all conditions are strictly kept consistent except for the model itself, and the model is the only variable that is modified. This experimental design with single variable change is applied in both the CBIS-DDSM and internal dataset experiments. The results show that BARNet performs excellently on multiple evaluation metrics and does not require the use of manual annotation and dual-view information. Especially in key metrics such as accuracy and precision, there is a significant improvement compared to the baseline network. By integrating global and local information and fusing multi-scale features, BARNet can more accurately identify lesion areas and classify them. The model not only considers the small size and irregular shape features of breast cancer lesions, but also excludes irrelevant areas and comprehensively judges by integrating information at different scales.
[0089] In the experiments conducted on the internal dataset, the same preprocessing method as the INbreast dataset was adopted, and the results showed that BARNet performed extremely well in various evaluation metrics including AUC, accuracy, specificity, and precision. This indicates that BARNet is very effective in reducing false positives, identifying lesions, and distinguishing positive and negative samples. The comprehensive experimental results show that BAR-Net performs well in the breast cancer classification task, which is mainly attributed to the following points: First, the feature extractor makes full use of the advantages of the convolutional neural network (CNN), and the feature reuse characteristic of the DenseNet feature extractor helps to reduce overfitting to a certain extent. Second, the AMG2L module uses the attention mechanism to reweight the intra-region and inter-region features, evaluating the lesions from local to global. The L2GCR module uses the spatial state through multi-scale input and the forward attention characteristic of the SSM to judge the local lesion situation from global to local. Finally, the two results are fused through a fully connected layer to obtain the final classification decision, effectively balancing the information at different scales. It is worth noting that the attention mechanism of the SSM has a linear complexity, which significantly reduces the training cost compared with the quadratic complexity of the traditional Transformer and does not require additional manual annotation, which is an obvious advantage.
[0090] The influence of different partition numbers in the L2GCR module on the model accuracy (Acc) and area under the curve (AUC) was studied in detail. The research results show that when the partition number exceeds 20, the AUC value of the model tends to be stable and is no longer significantly affected by the partition number. At the same time, the change in the accuracy value is very small, and with the increase of the partition number, its standard deviation is controlled within 4%. This result strongly proves that the L2GCR module can effectively reduce the adverse impact of rigid segmentation on the model performance. In addition, the influence of different segmentation methods on the model is relatively small. However, when the partition number decreases, the L2GCR module becomes a simple superposition of two global attention modules. At this time, the AUC value decreases significantly and is more vulnerable to high-value deviation.
[0091] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a bidirectional adaptive reweighted neural network for image processing, characterized in that: Input the image to be processed into the convolutional neural network feature extractor to obtain image features ,Will Input bidirectional adaptive reweighted neural network, bidirectional adaptive reweighted neural network It is divided into two branches, one of which inputs the cross-region reweighting module to obtain the output , an input adaptive multi-scale global to local reweighting module to get the output , combined with the attention mechanism and Reduce it to a vector of fixed size and get the image processing result through the fully connected layer; The cross-region reweighting module includes a regional attention module and a cross-region attention module; The regional attention module will The dimension is compressed to 512 to obtain image features ,Will Divided into 64 blocks, each block uses the two-dimensional selective state space model SS2D to perform local reweighting to obtain image features ,Will As the output of the region attention module ; SS2D includes: ; In the formula, represents the fusion operation, represents the two-dimensional selectivity state space model Module operation, , , , Respectively represent flattening along four directions; Modules include, The input of the module is linearly transformed to generate the matrix and the time step parameter , the hidden state is updated to: ; In the formula, is a natural constant, yes The module input, yes Always hide the status; Output for: ; All time steps Sequence sorting Module output; The cross-region attention module divides the input into two branches, one of which combines the adjustable parameters Perform matrix multiplication and then split into three inputs , and calculate the weight: ; ; ; In the formula, is the number of unidimensional cuts, is the total side length of the block, It is The side length of a block, , , There are three inputs The corresponding output weight, Indicates keys, Represents the total number of keys, , is the activation function; The second branch of the cross-region attention module After matrix multiplication, the output is input into the SS2D module. ,Will , , Perform matrix multiplication to get ; The adaptive multi-scale global to local reweighting module Perform maximum pooling, and then The global scale module, the quarter scale module, and the sixteenth scale module are input simultaneously, and weighted processing is performed in combination with different scale modules: ; In the formula, , It is a scale module, including the global scale module , quarter-scale module , 1 / 16 scale module , include The final result of the weighting is include The final result of the weighting is include The final result of the weighting is Right of expression, represents the final result of weighting, Represents element-by-element multiplication, including the weight information The output information of is unified in size, and then embedded respectively, and finally feature fusion is performed to obtain ; right Combined with 1×1 convolution, , , As three branches, 1×1 convolution processing is performed respectively, two of the three convolution processing results are matrix added, and then the addition result and the third convolution processing result are matrix added to obtain : ; In the formula, It is an adjustable parameter.
2. The method for constructing a bidirectional adaptive reweighted neural network for image processing according to claim 1, characterized in that: Combined with the attention mechanism and Reduction to a fixed-size vector includes: ; In the formula, express or , is a fully connected layer, are the fully connected layer parameters, represents the product of length and width, Indicates the number of channels, Time Representative , Time Representative , represents matrix multiplication, represents the output of the attention mechanism; Will Input fully connected layer: ; In the formula, yes or The predicted probability of is the activation function, and are the weights and biases of the linear regression layer respectively; Get the image processing results : ; In the formula, yes The predicted probability of yes The predicted probability of .
3. The method for constructing a bidirectional adaptive reweighted neural network for image processing according to claim 2, characterized in that: Loss function of bidirectional adaptive reweighted neural network for: ; In the formula, is the binary result of image processing, , The two values of correspond to two binary results, is the predicted probability corresponding to the binary outcome, =1, is the architectural parameter of the bidirectional adaptive reweighted neural network, , represents the regularization term, is the weight used to alleviate sample imbalance.
Citation Information
Patent Citations
Remote sensing image complex scene semantic segmentation method and system
CN117372686A
Forest fire early warning method and system based on super-resolution neural operator
CN119672544A