Quality control classification method and system suitable for laryngeal knot DR images

By adopting a deep learning algorithm with multi-scale and multi-dimensional attention mechanism in the Adam's apple DR image classification, combined with image cropping and pixel value standardization technology, the problem of insufficient ability to identify subtle structures and complex features in the existing technology is solved, and higher classification accuracy and stability are achieved.

CN120088524APending Publication Date: 2025-06-03WUHAN JULEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411937883.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing Adam's apple DR image classification method based on CNN model has shortcomings in identifying the subtle structures in the image and understanding complex features, resulting in the improvement of classification accuracy and stability.

Method used

The deep learning algorithm with multi-scale and multi-dimensional attention mechanism is adopted, and pre-processed through image cropping and pixel value standardization techniques, and combined with multiple cascading residual blocks and attention and multi-scale aggregation modules (CBAM spatial channel attention module and MSAA multi-scale feature aggregation module) for model training.

Benefits of technology

It improves the refinement analysis ability of Adam's apple DR images, enhances the recognition ability of Adam's apple characteristics at different scales, and significantly improves the accuracy and stability of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088524A_ABST
    Figure CN120088524A_ABST
Patent Text Reader

Abstract

The invention relates to a quality control classification method and system suitable for laryngeal knot DR images, and the method comprises the steps: carrying out the preprocessing of each obtained historical laryngeal knot DR image, and obtaining a preprocessed image; dividing a training set, a verification set and a test set according to a preset distribution proportion based on the preprocessed historical laryngeal knot DR image set; the training set and the verification set are input into a classification network model for model training, the classification network model is composed of a plurality of cascaded residual blocks, and after output of each residual block, an attention and multi-scale aggregation module is integrated, so that an attention and multi-scale aggregation model is formed; the attention and multi-scale aggregation module is composed of a CBAM space channel attention module and an MSAA multi-scale feature aggregation module which are arranged in sequence; inputting the test set into the trained classification network model, and optimizing model parameters based on a model performance evaluation result; and acquiring a real-time laryngeal knot DR image, inputting the real-time laryngeal knot DR image into the trained and optimized classification network model, and processing the real-time laryngeal knot DR image to obtain a corresponding classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a quality control classification method and system applicable to laryngeal prominence DR images. Background Art

[0002] With the development of computer vision and deep learning technologies, and the increasing application of DR images in the medical field, especially in the analysis of laryngeal prominence DR images, the automated quality control classification technology based on deep learning models can effectively improve the diagnostic accuracy and efficiency. Therefore, quality control and classification of DR images to ensure that the image quality is suitable for diagnosis is an important issue in medical imaging.

[0003] Currently, in order to improve the analysis ability of laryngeal prominence DR images, the medical imaging community widely adopts deep learning methods. Especially considering the application of convolutional neural networks (CNNs) in image classification tasks, applying CNN models to the quality control classification of laryngeal prominence DR images has become a mainstream trend and has achieved remarkable results. However, the existing laryngeal prominence DR image classification methods based on CNN models still have some limitations, such as insufficient recognition ability for fine structures in images and inability to effectively understand complex features in images. Therefore, it is necessary to study deep learning algorithms based on multi-scale and multi-dimensional attention mechanisms to enhance the refined analysis ability of laryngeal prominence DR images and improve the accuracy and stability of classification by capturing key features at different scales and dimensions in the images. Summary of the Invention

[0004] In order to achieve the accuracy and stability of DR image classification, the purpose of the present invention is to provide a quality control classification method and system applicable to laryngeal prominence DR images, and the specific technical solutions adopted are as follows:

[0005] In a first aspect, the present application discloses a quality control classification method applicable to laryngeal prominence DR images, and the method includes:

[0006] S1. Preprocess each obtained historical laryngeal prominence DR image based on image cropping and pixel value normalization techniques to obtain a corresponding preprocessed image;

[0007] S2. Divide the preprocessed historical laryngeal prominence DR image set into a training set, a validation set, and a test set according to a preset allocation ratio;

[0008] S3, inputting the training set and the validation set into a classification network model for model training, wherein the classification network model is composed of a plurality of cascaded residual blocks, and an attention and multi-scale aggregation module is integrated after the output of each residual block, and the attention and multi-scale aggregation module is composed of a sequentially arranged CBAM spatial channel attention module and an MSAA multi-scale feature aggregation module;

[0009] S4, inputting the test set into the trained classification network model, and tuning the model parameters based on the model performance evaluation results;

[0010] S5. Obtain a real-time DR image of the Adam's apple and input it into a trained and optimized classification network model, and obtain the corresponding classification result through forward propagation processing.

[0011] Further, in step S1, for each historical DR image of the Adam's apple, the image cropping and pixel value normalization technology is used to preprocess each acquired historical DR image of the Adam's apple to obtain a corresponding preprocessed image, including:

[0012] S11, according to the principle of minimizing the background area and retaining the integrity of the Adam's apple, the acquired historical Adam's apple DR image is cropped to obtain a corresponding cropped image;

[0013] S12, normalizing the value of each pixel in the cropped image to the interval [0, 1] according to a linear normalization method to obtain a corresponding zoomed image;

[0014] S13, performing Z-score normalization processing on the zoomed image to obtain a preprocessed image whose image grayscale values ​​present a standard normal distribution.

[0015] Furthermore, during the model training process, in step S3, after the residual feature map is obtained by processing the residual block and input into the CBAM spatial channel attention module, the method includes:

[0016] Performing a global average pooling operation and a global maximum pooling operation on the residual feature map to obtain a corresponding global pooling feature vector;

[0017] The obtained global pooling feature vectors are superimposed and merged to obtain a fused global feature vector;

[0018] The global feature vector is input into the fully connected layer for linear transformation processing, and the result of the linear transformation is activated using the sigmoid activation function to obtain a channel attention feature map.

[0019] Further, during the model training process, in step S3, after the residual feature map obtained through processing by the residual block is input into the CBAM spatial channel attention module, the method further includes:

[0020] Performing average pooling operation on the channel dimension of the residual feature map and max pooling operation on the channel dimension to obtain the corresponding channel average pooling feature map and channel max pooling feature map;

[0021] Superposing and merging the obtained channel average pooling feature map and channel max pooling feature map to obtain a fused global feature map;

[0022] Performing a 7×7 convolution operation on the global feature map and activating the obtained convolution feature map using the sigmoid activation function to obtain a spatial attention feature map.

[0023] Further, during the model training process, in step S3, after the residual feature map obtained through processing by the residual block is input into the CBAM spatial channel attention module, the method further includes:

[0024] Performing element-wise multiplication superposition fusion based on the residual feature map, channel attention feature map, and spatial attention feature map to obtain a fused feature map.

[0025] Further, during the model training process, in step S3, after the fused feature map is obtained through processing by the CBAM spatial channel attention module and the fused feature map is input into the MSAA multi-scale feature aggregation module, the method includes:

[0026] Performing a 1×1 convolution operation on the fused feature map and performing average pooling followed by two 1×1 convolution operations to obtain a first convolution feature map and a channel attention feature map, where a relu activation function for increasing non-linearity is inserted between the two 1×1 convolution operations;

[0027] Performing convolution operations on the first convolution feature map based on multiple convolution kernels of different sizes and performing multi-scale fusion based on the obtained second convolution feature maps to obtain a multi-scale fusion feature map;

[0028] Based on the multi-scale fusion feature map, performing a pooling operation and then passing through a 7×7 convolution followed by a sigmoid activation function, and after superposing the multi-scale fusion feature map for spatial aggregation, obtaining a multi-scale spatial fusion feature map;

[0029] Performing element-wise multiplication superposition fusion on the channel attention feature map and the multi-scale spatial fusion feature map to obtain a comprehensive feature map that fuses channel attention and spatial attention;

[0030] Superimpose and merge the comprehensive feature map with the fused feature map obtained by processing through the CBAM spatial channel attention module to obtain an enhanced attention feature map of the corresponding dimension.

[0031] Further, during the model training process, in step S3, after obtaining the enhanced attention feature maps of different dimensions, the method further includes:

[0032] Perform superimposed aggregation based on the enhanced attention feature maps of each dimension to obtain an aggregated feature map;

[0033] Input the aggregated feature map into a linear layer and activate it through a sigmoid activation function to obtain a predicted classification result.

[0034] Further, before model training, it is determined to use an ADAM optimizer for parameter update. In the ADAM optimizer, the initial learning rate is set to 1×10 -2 and the learning rate decay coefficient is 3×10 -5 , and at the same time, the batch size is set to 16, the number of training epochs is 100, and the cross-entropy loss function is used as the loss metric of the model.

[0035] In a second aspect, the present application also discloses a quality control classification system applicable to laryngeal prominence DR images. The system includes an image preprocessing module, a dataset allocation module, a model training module, a parameter tuning module, and a real-time classification prediction module, where:

[0036] The image preprocessing module is used to preprocess each acquired historical laryngeal prominence DR image based on image cropping and pixel value normalization techniques to obtain a corresponding preprocessed image;

[0037] The dataset allocation module is used to divide the preprocessed historical laryngeal prominence DR image set into a training set, a validation set, and a test set according to a preset allocation ratio;

[0038] The model training module is used to input the training set and the validation set into a classification network model for model training. Among them, the classification network model is composed of multiple cascaded residual blocks, and after the output of each residual block, an attention and multi-scale aggregation module is incorporated. The attention and multi-scale aggregation module is composed of a sequentially arranged CBAM spatial channel attention module and an MSAA multi-scale feature aggregation module;

[0039] The parameter tuning module is used to input the test set into the trained classification network model and tune the model parameters based on the model performance evaluation results;

[0040] The real-time classification prediction module is used to obtain real-time DR images of the laryngeal prominence and input them into the trained and optimized classification network model, and obtain corresponding classification results through forward propagation processing.

[0041] In a third aspect, the present application also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the quality control classification method applicable to DR images of the laryngeal prominence.

[0042] The present invention has the following beneficial effects:

[0043] 1) Through image cropping and pixel value standardization techniques, irrelevant information and noise in DR images of the laryngeal prominence can be effectively removed, and at the same time, the pixel value distribution becomes more uniform, thereby improving the image quality.

[0044] 2) In the classification network model, a structure of multiple cascaded residual blocks is adopted, which can effectively extract deep features of DR images of the laryngeal prominence. At the same time, by integrating an attention and multi-scale aggregation module (including a CBAM spatial-channel attention module and an MSAA multi-scale feature aggregation module), key features can be further highlighted and irrelevant features can be suppressed, improving the model's recognition ability for laryngeal prominence features at different scales. These designs enable the model to show higher accuracy and robustness in classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 It is a flowchart of a quality control classification method applicable to DR images of the laryngeal prominence provided by an embodiment of the present invention.

[0047] Figure 2 It is an overall structural block diagram of the classification network model.

[0048] Figure 3 It is a structural block diagram of the MSAA multi-scale feature aggregation module.

[0049] Figure 4 It is a system structure diagram of a quality control classification system applicable to DR images of the laryngeal prominence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a quality control classification method and system applicable to laryngeal DR images according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0052] The following specifically describes the specific solution of a quality control classification method and system applicable to laryngeal DR images provided by the present invention in conjunction with the accompanying drawings.

[0053] Please refer to Figure 1 , which shows a method flowchart of a quality control classification method applicable to laryngeal DR images provided by an embodiment of the present invention. The method includes:

[0054] Step S1: Based on image cropping and pixel value normalization techniques, preprocess each obtained historical laryngeal DR image to obtain a corresponding preprocessed image.

[0055] Specifically, the present application first removes irrelevant background regions using image cropping techniques according to the approximate position of the larynx in the DR image, following the principle of minimizing the background region and preserving the integrity of the larynx, and retains the effective region containing the larynx. Then, the obtained cropped image is normalized, that is, the pixel values in the image are converted to a standard interval range to eliminate pixel value differences caused by different parameter settings between different DR scanning devices.

[0056] Step S2: Based on the preprocessed historical laryngeal DR image set, divide it into a training set, a validation set, and a test set according to a preset allocation ratio.

[0057] Specifically, the present application divides the preprocessed historical laryngeal DR image set into a training set, a validation set, and a test set according to a ratio of 8:1:1. Among them, the training set is used for model training, the validation set is used to adjust model parameters during training to prevent overfitting, and the test set is used to evaluate model performance.

[0058] Step S3: Input the training set and the validation set into the classification network model for model training. The classification network model is composed of multiple cascaded residual blocks. After the output of each residual block, an attention and multi-scale aggregation module is incorporated. The attention and multi-scale aggregation module consists of a sequentially arranged CBAM (Convolutional Block Attention Module) spatial-channel attention module and an MSAA (Multi-Scale Aggregation of Attributes) multi-scale feature aggregation module.

[0059] Specifically, please refer to Figure 2 . During the model training process, first, feature extraction and deep representation learning are performed through 4 cascaded ResBlock (Residual Block) residual blocks. Among them, after each ResBlock residual block is processed, a residual feature map of the corresponding dimension will be obtained. These residual feature maps will be respectively input into the classification network model. The CBAM spatial-channel attention module first selectively focuses on the key features in the space and channels to highlight the important features and suppress the irrelevant features. Then, the MSAA multi-scale feature aggregation module performs multi-scale feature fusion to capture the local and global features in the feature map, improving the model's recognition ability for laryngeal prominence features of different scales.

[0060] Step S4: Input the test set into the trained classification network model, and optimize the model parameters based on the model performance evaluation results.

[0061] Specifically, this application will input the test set into the trained classification network model and evaluate the model's performance according to the accuracy (ACC), sensitivity, and specificity. Among them, the calculation formula of the accuracy (ACC) includes:

[0062]

[0063] where TP represents the number of samples correctly predicted as the positive class, TN represents the number of samples correctly predicted as the negative class, FP represents the number of negative-class samples mispredicted as the positive class, and FN represents the number of positive-class samples mispredicted as the negative class.

[0064] The calculation formula of the sensitivity includes:

[0065] The calculation formula of the specificity includes:

[0066] According to the results of the three indicators of accuracy (ACC), sensitivity, and specificity, when it is determined that at least one of the accuracy, sensitivity, and specificity is lower than a preset threshold, this application will adjust the hyperparameters of the model (including learning rate, batch size, regularization coefficient, etc.) in combination with an automated tuning tool.

[0067] Step S5: Obtain a real-time DR image of the laryngeal prominence and input it into the trained and optimized classification network model, and obtain the corresponding classification result through forward propagation processing.

[0068] Specifically, after obtaining the real-time DR image of the laryngeal prominence, this application will first perform image cropping and pixel value normalization processing on it. Then, the preprocessed image is input into the trained and optimized classification network model, and the classification result is obtained through forward propagation processing, including the situation with a laryngeal prominence and the situation without a laryngeal prominence. It should be noted that this result will be further used to assist doctors in diagnosing laryngeal prominence diseases or formulating treatment plans.

[0069] As can be seen from the above, a quality control classification method applicable to DR images of the laryngeal prominence disclosed in this application can effectively remove irrelevant information and noise in the DR images of the laryngeal prominence through image cropping and pixel value normalization techniques, and at the same time make the pixel value distribution more uniform, thereby improving the image quality; a structure of multiple cascaded residual blocks is adopted in the classification network model, which can effectively extract the deep features of the DR images of the laryngeal prominence. At the same time, by integrating an attention and multi-scale aggregation module (including a CBAM spatial channel attention module and an MSAA multi-scale feature aggregation module), key features can be further highlighted and irrelevant features can be suppressed, improving the model's recognition ability for laryngeal prominence features of different scales. These designs enable the model to show higher accuracy and robustness in the classification task.

[0070] In one embodiment, in step S1, for each historical DR image of the laryngeal prominence, the preprocessing of each obtained historical DR image of the laryngeal prominence is performed based on image cropping and pixel value normalization techniques to obtain the corresponding preprocessed image, including:

[0071] Step S11: Crop the obtained historical DR image of the laryngeal prominence according to the principle of minimizing the background area and preserving the integrity of the laryngeal prominence to obtain the corresponding cropped image.

[0072] Step S12: Normalize the value of each pixel point in the cropped image to the interval [0, 1] in a linear normalization manner to obtain the corresponding scaled image.

[0073] Specifically, this application will determine the maximum and minimum values of the pixels in the cropped image, and normalize the value of each pixel in the image to the interval [0, 1] according to the linear normalization formula, that is, (original pixel value - maximum value) / (maximum value - minimum value).

[0074] Step S13, perform Z-score normalization processing on the scaled image to obtain a preprocessed image with the gray values of the image following a standard normal distribution.

[0075] Specifically, this application will calculate the mean and standard deviation of all pixels in the scaled image, and convert the value of each pixel in the scaled image to a value following a standard normal distribution according to the Z-score normalization processing formula, that is, Z-score = (image pixel value - mean) / standard deviation, so as to reduce the convergence time in model training and improve the accuracy of the model.

[0076] In the above embodiments, on the one hand, by minimizing the background area and preserving the integrity of the Adam's apple, redundant information unrelated to the Adam's apple can be removed, reducing noise interference, which helps improve the accuracy and efficiency of subsequent image processing and analysis. On the other hand, normalizing the pixel values to the interval [0, 1] can eliminate the differences in pixel values between different images, making the image data have a unified scale, which helps the model to better learn and generalize. Finally, based on the Z-score normalization processing, the gray value distribution of the image conforms to the standard normal distribution, that is, the mean is 0 and the standard deviation is 1, which helps eliminate the skewness and kurtosis of the image data and improve the stability and performance of the model.

[0077] In one of the embodiments, during the model training process, in step S3, after the residual feature map obtained through the processing of the residual block is input into the CBAM spatial channel attention module, the method includes: performing global average pooling operation and global maximum pooling operation on the residual feature map respectively to obtain corresponding global pooling feature vectors; superimposing and merging the obtained global pooling feature vectors to obtain a fused global feature vector; inputting the global feature vector into a fully connected layer for linear transformation processing, and using the sigmoid activation function to activate the result of the linear transformation to obtain a channel attention feature map.

[0078] Specifically, the CBAM spatial channel attention module will perform global average pooling operation and global maximum pooling operation on the residual feature map respectively through the following formula, and perform superimposing and merging processing on the obtained global pooling feature vectors: AvgPool(X) + MaxPool(X), where X represents the input residual feature map, AvgPool(X) represents performing global average pooling operation on X, and MaxPool(X) represents performing global maximum pooling operation on X.

[0079] Further, AvgPool(X) + MaxPool(X) is used as the global feature vector input to the fully connected layer FC, which is linearly transformed by the fully connected layer FC and activated by the sigmoid activation function σ. The above processing process can be understood with reference to the formula: F ca = σ(FC(AvgPool(X) + MaxPool(X))), where F ca represents the obtained channel attention feature map.

[0080] In one embodiment, during the model training process, in step S3, after the residual feature map obtained through the residual block processing is input to the CBAM spatial channel attention module, the method further includes: performing an average pooling operation on the channel dimension of the residual feature map and a max pooling operation on the channel dimension to obtain the corresponding channel average pooling feature map and channel max pooling feature map; superimposing and merging the obtained channel average pooling feature map and channel max pooling feature map to obtain a fused global feature map; performing a 7×7 convolution operation on the global feature map, and using the sigmoid activation function to activate the obtained convolution feature map to obtain a spatial attention feature map.

[0081] Specifically, the CBAM spatial channel attention module will perform an average pooling operation on the channel dimension of the input residual feature map (i.e., calculating the average value of all pixel values within a specified window for each channel to obtain a channel average pooling feature map with the same number of channels) and a max pooling operation on the channel dimension (i.e., calculating the maximum value of all pixel values within a specified window for each channel to obtain a channel max pooling feature map with the same number of channels). And the obtained channel average pooling feature map and channel max pooling feature map are concatenated along the channel dimension, i.e., AvgPool C (X) + MaxPool C (X), to obtain a fused global feature map. Finally, a 7×7 convolution operation is performed on the obtained global feature map, and the sigmoid activation function σ is used to activate the obtained convolution feature map, i.e., F sa = σ(C 7×7 (AvgPool C (X) + MaxPool C (X))), to obtain a spatial attention feature map F sa . Wherein, AvgPool C (X) represents performing an average pooling operation on the channel dimension of the input residual feature map X, MaxPool C (X) represents performing a max pooling operation on the channel dimension of the input residual feature map X, and C 7×7Indicates a convolution operation based on a 7×7 size convolution kernel.

[0082] In one embodiment, during the model training process, in step S3, after the residual feature map obtained through processing by the residual block is input into the CBAM spatial channel attention module, the method further includes: performing element-wise multiplication superposition fusion based on the residual feature map, the channel attention feature map, and the spatial attention feature map to obtain a fused feature map.

[0083] Specifically, this application will use the formula to perform element-wise multiplication superposition fusion on the obtained residual feature map X, channel attention feature map F ca , and spatial attention feature map F sa , that is, perform multiplication calculations for each channel and spatial position. During the process, first, the channel attention feature map F ca is applied to the residual feature map X for per-channel multiplication operations. Among them, each channel value in the channel attention feature map is multiplied by all pixel values in the corresponding channel of the residual feature map X. Then, the feature map weighted by channel attention is multiplied by the spatial attention feature map for per-spatial position multiplication operations. Among them, the weight of each spatial position in the spatial attention feature map is multiplied by all channel values in the corresponding spatial position to emphasize or suppress specific regions in the feature map. Thus, through such element-wise multiplication superposition fusion, each pixel value in the residual feature map is affected by both channel attention and spatial attention, thereby obtaining a more refined feature representation.

[0084] In one of the embodiments, during the model training process, in step S3, after obtaining the fused feature map through the processing of the CBAM spatial channel attention module and inputting the fused feature map into the MSAA multi-scale feature aggregation module, the method includes: performing a 1×1 convolution operation on the fused feature map respectively, and performing average pooling followed by two 1×1 convolution operations to obtain a first convolutional feature map and a channel attention feature map, wherein a relu activation function for increasing non-linearity is inserted between the two 1×1 convolution operations; performing convolution operations on the first convolutional feature map respectively based on multiple convolutional kernels of different sizes, and performing multi-scale fusion on the obtained second convolutional feature maps respectively to obtain a multi-scale fused feature map; based on the multi-scale fused feature map, performing a pooling operation and then passing through a 7×7 convolution followed by a sigmoid activation function, and after superimposing the multi-scale fused feature map for spatial aggregation, obtaining a multi-scale spatial fused feature map; performing element-wise multiplication and superposition fusion on the channel attention feature map and the multi-scale spatial fused feature map to obtain a comprehensive feature map integrating channel attention and spatial attention; superimposing and merging the comprehensive feature map with the fused feature map obtained through the processing of the CBAM spatial channel attention module to obtain an enhanced attention feature map of the corresponding dimension.

[0085] Specifically, please refer to Figure 3 , this application will perform convolution operations on the input first convolutional feature map respectively with convolutional kernels of sizes 7×7, 5×5, and 3×3 to obtain second convolutional feature maps of corresponding sizes. Then, the obtained 3 second convolutional feature maps are fused through an addition operation to obtain a multi-scale fused feature map M. Specifically, it can be understood with reference to the formula: M = ∑ k∈{3,5,7} C k×k (X 1 ), where C k×k (X 1 ) represents performing a convolution operation on the input first convolutional feature map X 1 based on a convolutional kernel of size k×k.

[0086] Furthermore, please refer to Figure 3 , this application will perform a pooling operation based on the multi-scale fused feature map M to reduce the size of the feature map and extract key information therefrom. Subsequently, a 7×7 convolutional kernel is used to perform a convolution operation on the pooled feature map to capture more extensive context information. Finally, a sigmoid activation function is applied for normalization of the eigenvalue, where the normalized eigenvalue will be superimposed and fused onto the multi-scale fused feature map M in a point-by-point multiplication manner to achieve spatial aggregation. The above process can be understood with reference to the following formula: where M sa represents the multi-scale spatial fused feature map.

[0087] Further, the present application will fuse the channel attention feature map and the multi-scale spatial fusion feature map in a point-by-point multiplication manner for each element. This fusion method can ensure that the feature values at each position are jointly affected by channel attention and spatial attention, thereby obtaining a comprehensive feature map that fuses the two attention mechanisms.

[0088] In one embodiment, during the model training process, in step S3, after obtaining the enhanced attention feature maps of different dimensions, the method further includes: performing stacking aggregation based on the enhanced attention feature maps of each dimension to obtain an aggregated feature map; inputting the aggregated feature map into a linear layer and then activating it through a sigmoid activation function to obtain a predicted classification result.

[0089] Specifically, please refer to Figure 2 , on different feature dimensions, after being jointly processed by the CBAM spatial-channel attention module and the MSAA multi-scale feature aggregation module, the enhanced attention feature maps y 1 , y 2 , y 3 , y 4 are obtained. The present application will perform stacking aggregation on the enhanced attention feature maps of these 4 dimensions through the following formula to obtain the corresponding aggregated feature map Y, where:

[0090] Y = y 1 ⊕ y 2 ⊕ y 3 ⊕ y 4 .

[0091] In one embodiment, before model training, it is determined to use the ADAM optimizer for parameter update. In the ADAM optimizer, the initial learning rate is set to 1×10 -2 , and the learning rate decay coefficient is 3×10 -5 . At the same time, the batch size is set to 16, the number of training epochs is 100, and the cross-entropy loss function is used as the loss metric of the model.

[0092] Specifically, before the start of model training, the present application has carried out detailed configuration and initialization work. First, it is determined to use the ADAM optimizer to update the network parameters. It should be noted that the ADAM optimizer is widely adopted for its adaptive learning rate adjustment and good convergence performance. In this embodiment, the present application sets the initial learning rate of the ADAM optimizer to 1×10 -2 , and the learning rate decay coefficient is 3×10 -5, which also means that as the number of training rounds increases, the learning rate will gradually decrease according to the preset exponential decay coefficient. Specifically, the learning rate is adjusted after each round of training based on the current learning rate and the learning rate decay coefficient, usually by multiplying by (1 - learning rate decay coefficient), so as to ensure that the model can converge quickly in the initial stage of training, and can finely adjust the parameters in the later stage of training to avoid overfitting.

[0093] Furthermore, this application also sets the batch size to 16, which means that in each iteration, the model will process the data of 16 samples simultaneously. And, the number of training rounds is set to 100, that is, the model will traverse the entire training dataset 100 times. During each traversal, the model will perform forward propagation on the data of each batch, calculate the loss, and then perform backpropagation to update the network parameters. Finally, this application also sets to use the cross-entropy loss function as the loss metric of the model. Among them, the cross-entropy loss function is a commonly used loss function in classification tasks, which can measure the difference between the predicted distribution of the model and the actual distribution. During training, by minimizing the cross-entropy loss, the prediction of the model can be made closer to the true label.

[0094] In the above embodiments, detailed configuration work was carried out before model training, including selecting an optimizer, setting the learning rate and its decay coefficient, determining the batch size and the number of training rounds, and selecting a loss function, etc. These configurations provide a solid foundation for the training of the model and help the model achieve good performance in the subsequent training process.

[0095] Please refer to Figure 4 , a quality control classification system applicable to laryngeal DR images disclosed in this application, the system includes an image preprocessing module, a dataset allocation module, a model training module, a parameter tuning module, and a real-time classification prediction module, where:

[0096] The image preprocessing module is used to preprocess each acquired historical laryngeal DR image based on image cropping and pixel value normalization techniques to obtain the corresponding preprocessed image.

[0097] The dataset allocation module is used to divide the preprocessed historical laryngeal DR image set into a training set, a validation set, and a test set according to a preset allocation ratio.

[0098] The model training module is used to input the training set and the validation set into the classification network model for model training. Among them, the classification network model is composed of multiple cascaded residual blocks, and after the output of each residual block, an attention and multi-scale aggregation module is incorporated. The attention and multi-scale aggregation module is composed of a sequentially arranged CBAM spatial-channel attention module and an MSAA multi-scale feature aggregation module.

[0099] The parameter tuning module is used to input the test set into the trained classification network model and tune the model parameters based on the model performance evaluation results.

[0100] The real-time classification and prediction module is used to obtain real-time laryngeal DR images and input them into the trained and tuned classification network model, and obtain corresponding classification results through forward propagation processing.

[0101] In one embodiment, the above modules are also used to implement the enterprise industrial chain identification method based on the large language model described in any one of the foregoing method embodiments, and this application does not make any limitations in this regard.

[0102] As can be seen from the above, a quality control classification system applicable to laryngeal DR images disclosed in this application can effectively remove irrelevant information and noise in laryngeal DR images through image cropping and pixel value standardization techniques, and at the same time make the pixel value distribution more uniform, thereby improving the image quality; a multi-stage cascaded residual block structure is adopted in the classification network model, which can effectively extract the deep features of laryngeal DR images. At the same time, by integrating the attention and multi-scale aggregation module (including the CBAM spatial-channel attention module and the MSAA multi-scale feature aggregation module), key features can be further highlighted and irrelevant features can be suppressed, improving the model's recognition ability for laryngeal features of different scales. These designs enable the model to show higher accuracy and robustness in the classification task.

[0103] A computer-readable storage medium disclosed in this application stores a computer program, and when the computer program is executed by a processor, it implements the quality control classification method applicable to laryngeal DR images.

[0104] As can be seen from the above, a computer-readable storage medium disclosed in this application can effectively remove irrelevant information and noise in laryngeal DR images through image cropping and pixel value standardization techniques, and at the same time make the pixel value distribution more uniform, thereby improving the image quality; a multi-stage cascaded residual block structure is adopted in the classification network model, which can effectively extract the deep features of laryngeal DR images. At the same time, by integrating the attention and multi-scale aggregation module (including the CBAM spatial-channel attention module and the MSAA multi-scale feature aggregation module), key features can be further highlighted and irrelevant features can be suppressed, improving the model's recognition ability for laryngeal features of different scales. These designs enable the model to show higher accuracy and robustness in the classification task.

[0105] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

Claims

1. A quality control classification method for DR images of the Adam's apple, characterized in that: The method comprises: S1. Based on image cropping and pixel value standardization technology, each acquired historical Adam's apple DR image is preprocessed to obtain a corresponding preprocessed image; S2, based on the preprocessed historical laryngeal apple DR image set, dividing the training set, validation set, and test set according to a preset allocation ratio; S3, inputting the training set and the validation set into a classification network model for model training, wherein the classification network model is composed of a plurality of cascaded residual blocks, and an attention and multi-scale aggregation module is integrated after the output of each residual block, and the attention and multi-scale aggregation module is composed of a sequentially arranged CBAM spatial channel attention module and an MSAA multi-scale feature aggregation module; S4, inputting the test set into the trained classification network model, and tuning the model parameters based on the model performance evaluation results; S5. Obtain a real-time DR image of the Adam's apple and input it into a trained and optimized classification network model, and obtain the corresponding classification result through forward propagation processing.

2. The method according to claim 1, characterized in that In step S1, for each historical DR image of the Adam's apple, the image cropping and pixel value normalization technology is used to preprocess each acquired historical DR image of the Adam's apple to obtain a corresponding preprocessed image, including: S11, according to the principle of minimizing the background area and retaining the integrity of the Adam's apple, the acquired historical Adam's apple DR image is cropped to obtain a corresponding cropped image; S12, normalizing the value of each pixel in the cropped image to the interval [0, 1] according to a linear normalization method to obtain a corresponding zoomed image; S13, performing Z-score normalization processing on the zoomed image to obtain a preprocessed image whose image grayscale values ​​present a standard normal distribution.

3. The method according to claim 1, characterized in that: During the model training process, in step S3, after the residual feature map is obtained by processing the residual block and input into the CBAM spatial channel attention module, the method includes: Performing a global average pooling operation and a global maximum pooling operation on the residual feature map to obtain a corresponding global pooling feature vector; The obtained global pooling feature vectors are superimposed and merged to obtain a fused global feature vector; The global feature vector is input into the fully connected layer for linear transformation processing, and the result of the linear transformation is activated using the sigmoid activation function to obtain a channel attention feature map.

4. The method according to claim 3, characterized in that During the model training process, in step S3, after the residual feature map is obtained by processing the residual block and input into the CBAM spatial channel attention module, the method further includes: Performing channel-dimensional average pooling operations and channel-dimensional maximum pooling operations on the residual feature map to obtain corresponding channel-dimensional average pooling feature maps and channel-dimensional maximum pooling feature maps; The obtained channel average pooling feature map and channel maximum pooling feature map are superimposed and merged to obtain a fused global feature map; A 7×7 convolution operation is performed on the global feature map, and the obtained convolution feature map is activated using a sigmoid activation function to obtain a spatial attention feature map.

5. The method according to claim 3, characterized in that: During the model training process, in step S3, after the residual feature map is obtained by processing the residual block and input into the CBAM spatial channel attention module, the method further includes: Based on the residual feature map, the channel attention feature map, and the spatial attention feature map, element-level multiplication and superposition fusion are performed to obtain a fused feature map.

6. The method according to claim 1, characterized in that During the model training process, in step S3, after the fused feature map is obtained through the CBAM spatial channel attention module and the fused feature map is input into the MSAA multi-scale feature aggregation module, the method includes: Performing a 1×1 convolution operation and average pooling followed by two 1×1 convolution operations on the fused feature map, respectively, to obtain a first convolution feature map and a channel attention feature map, wherein a relu activation function for increasing nonlinearity is inserted between the two 1×1 convolution operations; Based on a plurality of convolution kernels of different sizes, the first convolution feature maps are respectively convolved, and multi-scale fusion is performed based on the obtained second convolution feature maps to obtain a multi-scale fusion feature map; Based on the multi-scale fusion feature map, a pooling operation is performed and a 7×7 convolution followed by a sigmoid activation function is performed, and after superimposing the multi-scale fusion feature map for spatial aggregation, a multi-scale spatial fusion feature map is obtained; The channel attention feature map is element-wise multiplied and superimposed with the multi-scale spatial fusion feature map to obtain a comprehensive feature map that combines channel attention and spatial attention; The comprehensive feature map is superimposed and merged with the fused feature map obtained by processing the CBAM spatial channel attention module to obtain an enhanced attention feature map of the corresponding dimension.

7. The method according to claim 6, characterized in that During the model training process, in step S3, after obtaining enhanced attention feature maps of different dimensions, the method further includes: Based on the enhanced attention feature maps of each dimension, superposition and aggregation are performed to obtain an aggregated feature map; The aggregated feature map is input into the linear layer and activated by the sigmoid activation function to obtain the predicted classification result.

8. The method according to claim 1, characterized in that Before model training, determine to use ADAM optimizer for parameter update. The initial learning rate in ADAM optimizer is set to 1×10 -2 , and the learning rate decay coefficient is 3×10 -5 , and set the batch size to 16, the number of training rounds to 100, and the cross entropy loss function as the loss metric of the model.

9. A quality control classification system for DR images of the Adam's apple, characterized in that: The system includes an image preprocessing module, a data set allocation module, a model training module, a parameter tuning module, and a real-time classification prediction module, wherein: The image preprocessing module is used to preprocess each acquired historical Adam's apple DR image based on image cropping and pixel value standardization technology to obtain a corresponding preprocessed image; The data set allocation module is used to divide the training set, the validation set, and the test set according to a preset allocation ratio based on the preprocessed historical laryngeal apple DR image set; The model training module is used to input the training set and the validation set into the classification network model for model training, wherein the classification network model is composed of a plurality of cascaded residual blocks, and an attention and multi-scale aggregation module is integrated after the output of each residual block, and the attention and multi-scale aggregation module is composed of a sequentially arranged CBAM spatial channel attention module and an MSAA multi-scale feature aggregation module; The parameter tuning module is used to input the test set into the trained classification network model and tune the model parameters based on the model performance evaluation results; The real-time classification prediction module is used to obtain real-time Adam's apple DR images and input them into the trained and optimized classification network model, and obtain the corresponding classification results through forward propagation processing.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the quality control classification method applicable to Larynx DR images is implemented.

Citation Information

Cited By

  • Method and system for predicting postoperative recurrence risk of embedded attention-enhancing PELD

    CN121709272A