Distorted image classification method and device, equipment and storage medium
By using convolutional neural networks and improved IRVFL classifiers, distorted images taken by different types of camera devices are classified, which solves the problem of poor results when using the same distortion correction model, and improves correction accuracy and model robustness.
Patent Information
- Application Number
- CN202510488811.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Due to the different degree of distortion, when the same distortion correction model is used for correction, the images captured by different types of imaging equipment have poor results and even serious deviations, resulting in the need to classify the distorted images to select a suitable correction model.
Distorted images are classified using convolutional neural networks and improved IRVFL classifiers. By introducing regularization coefficients and error vectors, the output weight of the model is optimized, and the robustness and generalization ability of the model are improved.
The accuracy and correction accuracy of distorted image classification are improved, the algorithm is avoided from falling into local optimal solutions, and the stability and reliability of the model are enhanced.
Smart Images

Figure CN120032186A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method, device, equipment and storage medium for classifying distorted images. Background Art
[0002] In the context of digital twin technology, video fusion technology is widely used to empower the virtual world and lead twin applications into the era of intelligent IoT and virtual-real symbiosis. Digital twin technology achieves real-time monitoring and prediction of physical systems by mapping data from the physical world to the virtual world. As an important part of this process, video fusion technology provides rich visual information for the virtual world by integrating and processing video data from different sources.
[0003] However, in practical applications, video images are usually taken by various types of cameras, including fisheye cameras, ordinary ball cameras, and gun cameras. The images taken by these devices have different degrees of distortion. Specifically, the images taken by fisheye cameras usually have large radial distortion, resulting in serious deformation of the edge of the image. The images taken by ordinary ball cameras and gun cameras have a smaller degree of distortion, but there is still a certain degree of distortion in some cases. Since the images taken by different types of cameras have different degrees of distortion, if the same distortion correction model is used for correction, the correction effect will be poor or even serious deviations will occur. Therefore, it is necessary to classify the distorted images so as to select different correction models to improve the accuracy and reliability of distortion correction. Therefore, a classification method for distorted images is urgently needed. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device and storage medium for classifying distorted images, which are used to classify distorted images and improve the accuracy of distorted image classification.
[0005] In a first aspect, the present application provides a method for classifying a distorted image, the method comprising: Inputting the distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type of the distorted image; wherein the distorted image classification model includes a convolutional neural network and an IRVFL classifier, and the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into the RVFL classifier; The distorted image classification model is trained in the following manner: Inputting the training samples into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training samples include a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; Obtaining an error vector according to the marked distortion type and the predicted distortion type; Obtaining a total model risk through the error vector, the regularization coefficient, and the output weight vector in the IRVFL classifier; If the total model risk does not meet the specified conditions, the model parameters and output weight vectors in the IRVFL classifier are adjusted, and then the step of inputting each distortion feature vector into the IRVFL classifier for classification is returned until the total model risk meets the specified conditions, and the training of the distorted image classification model is terminated.
[0006] In the embodiment of the present application, the distorted image is first classified by introducing an IRVFL classifier obtained by introducing a regularization coefficient into the RVFL classifier. Since the improved IRVFL algorithm introduces a regularization coefficient and determines the total model risk through an error vector, the regularization coefficient and the output weight in the IRVFL classifier, the global optimization capability of the algorithm is improved, and the algorithm is prevented from falling into a local optimal solution during the classification process, thereby improving the robustness and generalization ability of the model and enhancing the accuracy of image classification.
[0007] In a possible embodiment, the convolutional neural network is obtained by replacing the ReLU activation function in the residual block in the Resnet50 convolutional neural network with a SELU activation function.
[0008] In the embodiment of the present application, the SELU activation function can avoid the common "dead" neuron problem in the ReLU function, that is, when the input is always negative, the neuron output is always zero. SELU also has non-zero output in the negative interval, thus avoiding the problem of neuron "death". In addition, the SELU activation function has a self-normalization characteristic, which can automatically adjust the output distribution of neurons during training to keep it within a stable range, and the self-normalization characteristic reduces the problems of gradient disappearance and gradient explosion during training, thereby improving the training efficiency and stability of the model; SELU activation function, especially when processing complex and diverse distorted image data, SELU can better capture the characteristics of distorted images and improve classification accuracy.
[0009] In one embodiment, the convolutional neural network includes an initial convolutional layer, a maximum pooling layer, a first residual block group, a second residual block group, a third residual block group, a fourth residual block group, an average pooling layer, and a fully connected layer; The training samples are input into the convolutional neural network for feature extraction to obtain the distortion feature vectors, including: Using the initial convolution layer to perform a convolution operation on a first training distorted image to obtain a first feature vector, wherein the first training distorted image is any distorted image in the training samples; Performing maximum pooling on the first feature vector using the maximum pooling layer to obtain a second feature vector; Performing continuous feature extraction on the second feature vector using a plurality of residual blocks to obtain a third feature vector, wherein the plurality of residual blocks include the first residual block group, the second residual block group, the third residual block group, and the fourth residual block group; Using the average pooling layer to perform an average pooling operation on the third eigenvector to obtain a fourth eigenvector; The fully connected layer is used to perform a fully connected operation on the fourth feature vector to obtain a distorted feature vector corresponding to the first training distorted image.
[0010] In one embodiment, the training sample is obtained by: Acquire multiple distorted images with different distortion types; Correcting the color of each distorted image to obtain each corrected distorted image; Performing distortion type balancing on the corrected distorted images to obtain filtered distorted images; The filtered distorted images are divided into test samples, the training samples and verification samples.
[0011] In the embodiment of the present application, the color of each distorted image is first corrected, and then the distortion type of each distorted image is balanced before determining the training sample, thereby ensuring the accuracy of the obtained training samples to further improve the classification effect of the model.
[0012] In one embodiment, the color of each distorted image is corrected to obtain each corrected distorted image, including: Preliminarily adjusting the colors of the distorted images by using a color correction technology to obtain intermediate distorted images; The color of each intermediate distorted image is respectively corrected by using a color constancy algorithm to obtain each corrected distorted image.
[0013] In the embodiment of the present application, the distorted images are adjusted respectively by combining the color correction technology with the color constancy algorithm, thereby ensuring a more comprehensive optimization of the quality of the distorted images, improving the brightness, contrast and color consistency of the distorted images under different lighting conditions, and reducing image distortion, thereby improving the accuracy and reliability of subsequent processing.
[0014] In one embodiment, before inputting each distortion feature vector into an IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector, the method further includes: The PCA algorithm is used to reduce the dimension of each distorted feature vector to obtain each distorted feature vector after the dimension reduction, and each distorted feature vector after the dimension reduction is determined as the distorted feature vector.
[0015] In the embodiment of the present application, by reducing the dimension of each distorted feature vector, the dimension of big data is reduced, thereby reducing the computational complexity and improving the model training efficiency.
[0016] In one embodiment, the total model risk is obtained by using the error vector, the regularization coefficient and the output weight vector in the IRVFL classifier, including: The total model risk is obtained by the following formula: ; Wherein, E is the total model risk, C is a preset constant value, is the error vector, is the regularization coefficient, is the output weight vector.
[0017] In one embodiment, after completing the training of the distorted image classification model, the method further includes: Inputting the verification sample into the convolutional neural network for feature extraction to obtain a distortion feature vector of each distorted image in the verification sample; wherein the verification sample includes a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting the distortion feature vector of each distorted image into the trained distorted image classification model to obtain the predicted distortion type of each distorted image in the verification sample; Obtaining the accuracy of the trained distorted image classification model according to the annotated distortion type of each distorted image in the verification sample and the predicted distortion type of each distorted image; If the accuracy is not greater than a specified threshold, the distorted image classification model is retrained.
[0018] In the embodiment of the present application, the accuracy of the trained distorted image classification model is verified by using verification samples, and if it is determined that the accuracy is not greater than a specified threshold, the distorted image classification model needs to be retrained, thereby improving the accuracy of distorted image classification.
[0019] In a second aspect, the present application provides a distorted image classification device, the device comprising: A distortion type determination module, used for inputting the distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type; the distorted image classification model includes a convolutional neural network and an IRVFL classifier, and the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into the RVFL classifier; The training module is used to train the distorted image classification model in the following manner: Inputting the training samples into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training samples include a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; Obtaining an error vector according to the marked distortion type and the predicted distortion type; Obtaining a total model risk through the error vector, the regularization coefficient, and the output weight vector in the IRVFL classifier; If the total model risk does not meet the specified conditions, the model parameters and output weight vectors in the IRVFL classifier are adjusted, and then the step of inputting each distortion feature vector into the IRVFL classifier for classification is returned until the total model risk meets the specified conditions, and the training of the distorted image classification model is terminated.
[0020] In a possible embodiment, the convolutional neural network is obtained by replacing the ReLU activation function in the residual block in the Resnet50 convolutional neural network with a SELU activation function.
[0021] In a possible embodiment, the convolutional neural network includes an initial convolutional layer, a maximum pooling layer, a first residual block group, a second residual block group, a third residual block group, a fourth residual block group, an average pooling layer and a fully connected layer; The training module is also used for: Using the initial convolution layer to perform a convolution operation on a first training distorted image to obtain a first feature vector, wherein the first training distorted image is any distorted image in the training samples; Performing maximum pooling on the first feature vector using the maximum pooling layer to obtain a second feature vector; Performing continuous feature extraction on the second feature vector using a plurality of residual blocks to obtain a third feature vector, wherein the plurality of residual blocks include the first residual block group, the second residual block group, the third residual block group, and the fourth residual block group; Using the average pooling layer to perform an average pooling operation on the third eigenvector to obtain a fourth eigenvector; The fully connected layer is used to perform a fully connected operation on the fourth feature vector to obtain a distorted feature vector corresponding to the first training distorted image.
[0022] In a possible embodiment, the device further includes: The training sample determination module is used to obtain the training sample in the following manner: Acquire multiple distorted images with different distortion types; Correcting the color of each distorted image to obtain each corrected distorted image; Performing distortion type balancing on the corrected distorted images to obtain filtered distorted images; The training samples are obtained according to the filtered distorted images.
[0023] In a possible embodiment, the training sample determination module is further used to: Preliminarily adjusting the colors of the distorted images by using a color correction technology to obtain intermediate distorted images; The color of each intermediate distorted image is respectively corrected by using a color constancy algorithm to obtain each corrected distorted image.
[0024] In a possible implementation, the device further includes: A dimensionality reduction module is used for inputting each distortion feature vector into an IRVFL classifier for classification, and before obtaining the predicted distortion type corresponding to each distortion feature vector, reducing the dimension of each distortion feature vector using a PCA algorithm to obtain each distortion feature vector after dimensionality reduction, and determining each distortion feature vector after dimensionality reduction as each distortion feature vector.
[0025] In a possible implementation manner, the training module is further used to: The total model risk is obtained by the following formula: ; Wherein, E is the total model risk, C is a preset constant value, is the error vector, is the regularization coefficient, is the output weight vector.
[0026] In a possible embodiment, the device further includes: A feature extraction module, which is used to input the verification sample into the convolution neural network for feature extraction after the training of the distorted image classification model is completed, so as to obtain the distortion feature vector of each distorted image in the verification sample; wherein the verification sample includes multiple distorted images and annotated distortion types corresponding to the multiple distorted images; An input module, used for inputting the distortion feature vector of each distorted image into the trained distorted image classification model to obtain the predicted distortion type of each distorted image in the verification sample; A verification module, used to obtain the accuracy of the trained distorted image classification model according to the annotated distortion type of each distorted image in the verification sample and the predicted distortion type of each distorted image; The judgment module is used to retrain the distorted image classification model if the accuracy is not greater than a specified threshold.
[0027] In a third aspect, the present application provides an electronic device, including: A memory for storing program instructions; The processor is used to call the program instructions stored in the memory, and execute the steps included in any one of the methods in the first aspect according to the obtained program instructions.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes any one of the methods described in the first aspect.
[0029] In a fifth aspect, the present application provides a computer program product, comprising: a computer program code, when the computer program code is run on a computer, the computer executes any one of the methods described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 One of the flowcharts of a method for classifying distorted images provided in an embodiment of the present application; Figure 2 A schematic diagram of a process for determining training samples provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a convolutional neural network provided in an embodiment of the present application; Figure 4 A schematic diagram of a process for extracting features for a convolutional neural network provided in an embodiment of the present application; Figure 5 A second flow chart of a method for classifying distorted images provided in an embodiment of the present application; Figure 6 A schematic diagram of a distorted image classification device provided in an embodiment of the present application; Figure 7 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0032] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application can mean at least two, for example, two, three or more, and the embodiments of the present application are not limited.
[0033] In the technical solution of this application, the collection, dissemination, and use of data are in compliance with the requirements of relevant national laws and regulations.
[0034] Before introducing the distorted image classification method provided by the embodiment of the present application, in order to facilitate understanding, the technical background of the embodiment of the present application is first introduced in detail.
[0035] At present, video images are usually taken by various types of camera equipment, including fisheye cameras, ordinary ball cameras and gun cameras. The images taken by these devices have different degrees of distortion. Specifically, the images taken by fisheye cameras usually have large radial distortion, resulting in serious deformation of the edge of the image. The images taken by ordinary ball cameras and gun cameras have a smaller degree of distortion, but there is still a certain degree of distortion in some cases. Since the images taken by different types of camera equipment have different degrees of distortion, if the same distortion correction model is used for correction, it will lead to poor correction effect or even serious deviation. Therefore, it is necessary to classify the distorted images so as to select different correction models to improve the accuracy and reliability of distortion correction. Therefore, a classification method for distorted images is urgently needed.
[0036] Therefore, the present application provides a method for classifying distorted images. The distorted images are classified by using an IRVFL classifier obtained by introducing a regularization coefficient into the RVFL classifier. Since the improved IRVFL algorithm introduces a regularization coefficient and determines the total model risk through the error vector, the regularization coefficient and the output weight in the IRVFL classifier, the global optimization capability of the algorithm is improved, and the algorithm is prevented from falling into a local optimal solution during the classification process, thereby improving the robustness and generalization capability of the model and enhancing the accuracy of image classification. Below, the scheme of the present disclosure is described in detail in conjunction with the accompanying drawings.
[0037] like Figure 1 FIG. 1 is a flow chart of a method for classifying distorted images, which specifically includes the following steps: Step 101: inputting the training sample into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training sample includes a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; First, the method of determining the training samples in the embodiment of the present application is described. Figure 2 As shown, it is a flowchart of determining training samples, which may specifically include the following steps: Step 201: Acquire multiple distorted images of different distortion types; Step 202: Correct the color of each distorted image to obtain each corrected distorted image; In a possible implementation, step 202 may be specifically implemented as follows: using a color correction technique to preliminarily adjust the colors of the distorted images to obtain intermediate distorted images; and using a color constancy algorithm to perform color correction on the intermediate distorted images to obtain the corrected distorted images.
[0038] The color correction technology in the embodiment of the present application may be a gamma correction algorithm, etc. The specific algorithm may be set according to the specific actual situation, and the embodiment of the present application does not limit the color correction technology.
[0039] Step 203: performing distortion type balance on the corrected distorted images to obtain filtered distorted images; In the embodiment of the present application, oversampling technology and / or undersampling technology are used to balance the distortion types of the corrected distorted images. In addition, in the embodiment of the present application, the pixels of each distorted image after filtering are adjusted, and the pixels of any distorted image are adjusted to Pixel.
[0040] Step 204: Obtain the training samples according to the filtered distorted images.
[0041] In the embodiment of the present application, 20% of each distorted image after filtering is used as a test sample. The remaining 80% of each distorted image is divided into training samples and verification samples in a ratio of 8:2. However, the embodiment of the present application does not limit the division method of each sample, and the specific method can be set according to the actual situation.
[0042] The convolutional neural network in the embodiment of the present application is obtained by replacing the ReLU activation function in the residual block of the Resnet50 convolutional neural network with the SELU activation function. Wherein, formula (1) is the mathematical model of the SELU activation function: ……(1); Among them, x is the input value of SELU activation function, is the output value of the SELU activation function, a is a preset first weight value, and b is a preset second weight value.
[0043] In the embodiment of the present application, b is set to 1.0507 and a is set to 1.67326. However, the specific values of b and a in the embodiment of the present application are not limited, and the specific values of b and a can be set according to specific actual conditions.
[0044] like Figure 3 As shown in Figure 1, it is a schematic diagram of the structure of a convolutional neural network. Figure 3 It can be seen that the convolutional neural network 300 includes an initial convolutional layer 301, a maximum pooling layer 302, a first residual block group 303, a second residual block group 304, a third residual block group 305, a fourth residual block group 306, an average pooling layer 307 and a fully connected layer 308. The following is a detailed description of step 101 in the embodiment of the present application in conjunction with the structure of the convolutional neural network in the embodiment of the present application. Figure 4The above is a schematic diagram of the process of feature extraction by convolutional neural network, which may specifically include the following steps: Step 401: performing a convolution operation on a first training distorted image using the initial convolution layer to obtain a first feature vector, wherein the first training distorted image is any distorted image in the training samples; The initial convolutional layer in the embodiment of the present application includes 64 The convolution kernel of . And the size of the first eigenvector obtained is .
[0045] Step 402: using the maximum pooling layer to perform maximum pooling on the first feature vector to obtain a second feature vector; The size of the maximum pooling layer in the embodiment of the present application is , so the size of the second eigenvector is .
[0046] Step 403: performing continuous feature extraction on the second feature vector using a plurality of residual blocks to obtain a third feature vector, wherein the plurality of residual blocks include the first residual block group, the second residual block group, the third residual block group and the fourth residual block group; The first residual block group in the embodiment of the present application includes four convolution residual blocks. The four convolution residual blocks each include a convolution layer + a batch normalization layer (BN) + a scaled exponential linear unit (SELU), and each convolution layer contains 64 Convolution kernel, 64 Convolution kernel and 256 The size of the feature vector obtained by the first residual block group is .
[0047] The second residual block group in the embodiment of the present application includes five convolution residual blocks, each of which includes a convolution layer + a batch normalization layer (BN) + a scaled exponential linear unit (SELU). Each convolution layer contains 128 Convolution kernel, 128 Convolution kernel and 512 The size of the feature vector obtained by the second residual block group becomes .
[0048] The third residual block group in the embodiment of the present application includes seven residual blocks. Each residual block includes a convolution layer + a batch normalization layer (BN) + a scaled exponential linear unit (SELU). The convolution layer of each residual block contains 256 Convolution kernel, 256 Convolution kernel and 1024 The size of the feature vector output by the third residual block group becomes .
[0049] The fourth residual block group in the embodiment of the present application includes four residual blocks. Each residual block includes a convolution layer + a batch normalization layer (BN) + a scaled exponential linear unit (SELU), and the convolution layer in each residual block contains 512 Convolution kernel, 512 Convolution kernel and 2048 Convolution kernel. The size of the feature vector output by the fourth residual block group becomes .
[0050] Step 404: using the average pooling layer to perform an average pooling operation on the third feature vector to obtain a fourth feature vector; The size of the average pooling layer in the embodiment of the present application is .
[0051] Step 405: Use the fully connected layer to perform a fully connected operation on the fourth eigenvector to obtain a distorted eigenvector corresponding to the first training distorted image.
[0052] The distorted feature vector obtained in the embodiment of the present application has a feature attribute of 2048 dimensions.
[0053] It should be noted that the structure of each layer in the Resnet50 convolutional neural network in the embodiment of the present application is only used for illustration and does not limit the structure of each layer in the embodiment of the present application.
[0054] In order to improve the training efficiency of the model, in a possible implementation, before executing step 102, the PCA algorithm is used to reduce the dimension of each distorted feature vector to obtain each distorted feature vector after dimensionality reduction, and the each distorted feature vector after dimensionality reduction is determined as the distorted feature vector.
[0055] In the embodiment of the present application, the parameters of the PCA algorithm are set to 0.99, retaining at least 99% of the principal component variance. The PCA algorithm automatically calculates the principal component variance based on the input data, and selects the least principal component so that the sum of the variances of these principal components is greater than or equal to the set variance share. The low-dimensional features after dimensionality reduction are used as the number of input layer nodes of the improved RVFL, so that the data becomes more compact and efficient, the number of input layer nodes of the improved RVFL is optimized, and the efficiency and accuracy of model training are improved.
[0056] Step 102: inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into the RVFL classifier; The regularization coefficient in the embodiment of the present application is used to train the IRVFL classifier, so that the model of the trained IRVFL classifier is more accurate.
[0057] Step 103: obtaining an error vector according to the marked distortion type and the predicted distortion type; In one embodiment, step 103 may be specifically implemented as follows: for any distortion feature vector, subtract the predicted distortion type corresponding to the distortion feature vector from the marked distortion type to obtain an error vector of the distortion feature vector; and determine the average value of the error vectors of the distortion feature vectors as the error vector.
[0058] In the embodiment of the present application, each distortion type can be represented by a different numerical value. It can be set according to the actual situation, and the embodiment of the present application is not limited here. In addition, the method of determining the error vector in the embodiment of the present application is only used for illustration, and the method of determining the error vector in the embodiment of the present application is not limited. The specific method of determining the error vector in the embodiment of the present application can be set according to the specific actual situation.
[0059] Step 104: Obtain the total model risk through the error vector, the regularization coefficient and the output weight vector in the IRVFL classifier; wherein the total model risk can be obtained through formula (2): ……(2); Wherein, E is the total model risk, C is a preset constant value, is the error vector, is the regularization coefficient, is the output weight vector.
[0060] Step 105: Determine whether the total model risk meets the specified conditions, if not, execute step 106, if yes, execute step 107; The specified condition in the embodiment of the present application is to minimize the value of the total model risk.
[0061] Step 106: After adjusting the model parameters and output weight vector in the IRVFL classifier, return to step 102; In the embodiment of the present application, the output weight vector needs to be adjusted to the optimal value, wherein the optimal output weight vector can be obtained by formula (3): ……(3); in, is the optimal output weight vector, H is the preset target output matrix, Y is the actual output matrix in the IRVFL classifier, and C is a preset constant value.
[0062] Step 107: End the training of the distorted image classification model; In the embodiment of the present application, after the training of the classification model of the distorted image is completed, the trained distorted image classification model is verified using the verification sample. If the verification fails, the distorted image classification model needs to be retrained. If the verification passes, the trained distorted image classification model is tested using the test sample to determine the performance of the distorted image classification model. The following is an introduction to the method of using the verification sample to verify the trained distorted image classification model in the embodiment of the present application. In a possible embodiment, a verification sample is input into the convolutional neural network for feature extraction to obtain a distortion feature vector of each distorted image in the verification sample; wherein the verification sample includes multiple distorted images and annotated distortion types corresponding to the multiple distorted images; the distortion feature vector of each distorted image is input into a trained distorted image classification model to obtain a predicted distortion type of each distorted image in the verification sample; based on the annotated distortion type of each distorted image in the verification sample and the predicted distortion type of each distorted image, the accuracy of the trained distorted image classification model is obtained; if the accuracy is not greater than a specified threshold, the distorted image classification model is retrained; if the accuracy is greater than a specified threshold, the training of the distorted image classification model is terminated.
[0063] It should be noted that the specified threshold in the embodiment of the present application can be set according to the specific actual situation, and the embodiment of the present application does not limit the specific value of the specified threshold.
[0064] Step 108: inputting the distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type of the distorted image; wherein the distorted image classification model includes a convolutional neural network and an IRVFL classifier.
[0065] The distortion types in the embodiments of the present application include radial distortion, barrel distortion, pincushion distortion, and mixed distortion.
[0066] In order to further understand the classification method of distorted images in this application, Figure 5 FIG. 1 is a flow chart of a method for classifying distorted images, which specifically includes the following steps: Step 501: Acquire multiple distorted images of different distortion types; Step 502: Correct the color of each distorted image to obtain each corrected distorted image; Step 503: performing distortion type balance on the corrected distorted images to obtain filtered distorted images; Step 504: obtaining the training samples according to the filtered distorted images; Step 505: input the training sample into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training sample includes multiple distorted images and the labeled distortion types corresponding to the multiple distorted images, and the convolution neural network is obtained by replacing the ReLU activation function in the residual block in the Resnet50 convolution neural network with the SELU activation function; Step 506: using a PCA algorithm to reduce the dimension of each distortion feature vector to obtain each distortion feature vector after dimension reduction, and determining each distortion feature vector after dimension reduction as each distortion feature vector; Step 507: inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; Step 508: Obtain an error vector according to the marked distortion type and the predicted distortion type; Step 509: Obtaining the total model risk through the error vector, the regularization coefficient and the output weight vector in the IRVFL classifier; Step 510: Determine whether the total model risk meets the specified conditions, if yes, execute step 512, if no, execute step 511; Step 511: after adjusting the model parameters and output weight vector in the IRVFL classifier, return to step 507; Step 512: End the training of the distorted image classification model; Step 513: input the distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type of the distorted image.
[0067] Based on the same inventive concept, an embodiment of the present application provides a distorted image classification device. The effect of the distorted image classification device is similar to that of the aforementioned method, and will not be described in detail herein.
[0068] Figure 6 The figure is a schematic diagram of the structure of a distorted image classification device according to an embodiment of the present disclosure.
[0069] like Figure 6 As shown, the distorted image classification device 600 of the present disclosure may include a distortion type determination module 610 and a training module 620 .
[0070] The distortion type determination module 610 is configured to input a distortion image to be detected into a pre-trained distortion image classification model to obtain the distortion type; the distortion image classification model includes a convolutional neural network and an IRVFL classifier, and the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into the RVFL classifier; The training module 620 is configured to train the distortion image classification model in the following manner: Input the training samples into the convolutional neural network for feature extraction to obtain respective distortion feature vectors, where the training samples include multiple distortion images and the labeled distortion types respectively corresponding to the multiple distortion images; Input the respective distortion feature vectors into the IRVFL classifier for classification to obtain the predicted distortion types corresponding to the respective distortion feature vectors; According to the labeled distortion type and the predicted distortion type, obtain an error vector; Through the error vector, the regularization coefficient, and the output weight vector in the IRVFL classifier, obtain the total model risk; If the total model risk does not meet the specified condition, then after adjusting the model parameters and the output weight vector in the IRVFL classifier, return to the step of inputting the respective distortion feature vectors into the IRVFL classifier for classification until the total model risk meets the specified condition, and then end the training of the distortion image classification model.
[0071] In a possible embodiment, the convolutional neural network is obtained by replacing the ReLU activation function in the residual block of the Resnet50 convolutional neural network with the SELU activation function.
[0072] In a possible embodiment, the convolutional neural network includes an initial convolutional layer, a max pooling layer, a first residual block group, a second residual block group, a third residual block group, a fourth residual block group, an average pooling layer, and a fully connected layer; The training module 620 is further configured to: Perform a convolution operation on a first training distortion image by using the initial convolutional layer to obtain a first feature vector, where the first training distortion image is any one of the distortion images in the training samples; Perform max pooling on the first feature vector by using the max pooling layer to obtain a second feature vector; Perform continuous feature extraction on the second feature vector by using multiple residual blocks to obtain a third feature vector, where the multiple residual blocks include the first residual block group, the second residual block group, the third residual block group, and the fourth residual block group; Using the average pooling layer to perform an average pooling operation on the third eigenvector to obtain a fourth eigenvector; The fully connected layer is used to perform a fully connected operation on the fourth feature vector to obtain a distorted feature vector corresponding to the first training distorted image.
[0073] In a possible embodiment, the device further includes: The training sample determination module 630 is used to obtain the training sample by: Acquire multiple distorted images with different distortion types; Correcting the color of each distorted image to obtain each corrected distorted image; Performing distortion type balancing on the corrected distorted images to obtain filtered distorted images; The training samples are obtained according to the filtered distorted images.
[0074] In a possible embodiment, the training sample determination module 630 is further configured to: Preliminarily adjusting the colors of the distorted images by using a color correction technology to obtain intermediate distorted images; The color of each intermediate distorted image is respectively corrected by using a color constancy algorithm to obtain each corrected distorted image.
[0075] In a possible implementation, the device further includes: The dimension reduction module 640 is used to input the distortion feature vectors into the IRVFL classifier for classification, and before obtaining the predicted distortion type corresponding to each distortion feature vector, use the PCA algorithm to reduce the dimension of each distortion feature vector to obtain each distortion feature vector after dimension reduction, and determine each distortion feature vector after dimension reduction as each distortion feature vector.
[0076] In a possible implementation, the training module 620 is further configured to: The total model risk is obtained by the following formula: ; Wherein, E is the total model risk, C is a preset constant value, is the error vector, is the regularization coefficient, is the output weight vector.
[0077] In a possible embodiment, the device further includes: A feature extraction module 650 is used to input the verification sample into the convolution neural network for feature extraction after the training of the distorted image classification model is completed, so as to obtain the distortion feature vector of each distorted image in the verification sample; wherein the verification sample includes multiple distorted images and annotated distortion types corresponding to the multiple distorted images; An input module 660, configured to input the distortion feature vector of each distorted image into the trained distorted image classification model to obtain a predicted distortion type of each distorted image in the verification sample; Verification 670, for obtaining the accuracy of the trained distorted image classification model according to the annotated distortion type of each distorted image in the verification sample and the predicted distortion type of each distorted image; The judgment module 680 is used to retrain the distorted image classification model if the accuracy is not greater than a specified threshold.
[0078] Based on the same technical concept, the embodiment of the present application also provides an electronic device 700, such as Figure 7 As shown, it includes at least one processor 701 and a memory 702 connected to the at least one processor. The specific connection medium between the processor 701 and the memory 702 is not limited in the embodiment of the present application. Figure 7 For example, the processor 701 and the memory 702 are connected via a bus 703. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0079] Among them, the processor 701 is the control center of the electronic device, which can use various interfaces and lines to connect various parts of the electronic device, and realize data processing by running or executing instructions stored in the memory 702 and calling data stored in the memory 702. Optionally, the processor 701 may include one or more processing units, and the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user pages, and applications, etc., and the modem processor mainly processes the issuance of instructions. It is understandable that the above-mentioned modem processor may not be integrated into the processor 701. In some embodiments, the processor 701 and the memory 702 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on independent chips.
[0080] Processor 701 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the problem location method in combination with the system can be directly embodied as a hardware processor for execution, or can be executed by a combination of hardware and software modules in the processor.
[0081] The memory 702 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 702 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, and the like. The memory 702 is any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 702 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0082] In the embodiment of the present application, the memory 702 stores a computer program. When the program is executed by the processor 701, the processor 701 executes the distorted image classification method described above.
[0083] Since the electronic device is the electronic device in the method in the embodiment of the present application, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0084] Based on the same inventive concept, the embodiment of the present application provides a computer-readable storage medium, and the computer program product includes: computer program code, when the computer program code is executed on a computer, the computer executes any of the distorted image classification methods discussed above. Therefore, the implementation of the above-mentioned computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be repeated.
[0085] Based on the same inventive concept, the embodiment of the present application further provides a computer program product, which includes: computer program code, when the computer program code is run on a computer, the computer executes any of the distorted image classification methods discussed above. Since the principle of solving the problem by the above computer program product is similar to that of the distorted image classification method, the implementation of the above computer program product can refer to the implementation of the method, and the repeated parts will not be repeated.
[0086] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0087] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0088] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0089] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of user operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.
[0090] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for classifying distorted images, characterized in that: The method comprises: Inputting the distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type of the distorted image; wherein the distorted image classification model includes a convolutional neural network and an IRVFL classifier, and the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into the RVFL classifier; The distorted image classification model is trained in the following manner: Inputting the training samples into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training samples include a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; Obtaining an error vector according to the marked distortion type and the predicted distortion type; Obtaining a total model risk through the error vector, the regularization coefficient, and the output weight vector in the IRVFL classifier; If the total model risk does not meet the specified conditions, the model parameters and output weight vectors in the IRVFL classifier are adjusted, and then the step of inputting each distortion feature vector into the IRVFL classifier for classification is returned until the total model risk meets the specified conditions, and the training of the distorted image classification model is terminated.
2. The method according to claim 1, characterized in that The convolutional neural network is obtained by replacing the ReLU activation function in the residual block in the Resnet50 convolutional neural network with the SELU activation function.
3. The method according to claim 2, characterized in that The convolution neural network includes an initial convolution layer, a maximum pooling layer, a first residual block group, a second residual block group, a third residual block group, a fourth residual block group, an average pooling layer and a fully connected layer; The training samples are input into the convolutional neural network for feature extraction to obtain the distortion feature vectors, including: Using the initial convolution layer to perform a convolution operation on a first training distorted image to obtain a first feature vector, wherein the first training distorted image is any distorted image in the training samples; Performing maximum pooling on the first feature vector using the maximum pooling layer to obtain a second feature vector; Performing continuous feature extraction on the second feature vector using a plurality of residual blocks to obtain a third feature vector, wherein the plurality of residual blocks include the first residual block group, the second residual block group, the third residual block group, and the fourth residual block group; Using the average pooling layer to perform an average pooling operation on the third eigenvector to obtain a fourth eigenvector; The fully connected layer is used to perform a fully connected operation on the fourth feature vector to obtain a distorted feature vector corresponding to the first training distorted image.
4. The method according to claim 1, characterized in that The training samples are obtained by: Acquire multiple distorted images with different distortion types; Correcting the color of each distorted image to obtain each corrected distorted image; Performing distortion type balancing on the corrected distorted images to obtain filtered distorted images; The training samples are obtained according to the filtered distorted images.
5. The method according to claim 4, characterized in that Correcting the colors of the distorted images to obtain the corrected distorted images includes: Preliminarily adjusting the colors of the distorted images by using a color correction technology to obtain intermediate distorted images; The color of each intermediate distorted image is respectively corrected by using a color constancy algorithm to obtain each corrected distorted image.
6. The method according to claim 1, characterized in that Before inputting each distortion feature vector into the IRVFL classifier for classification to obtain the predicted distortion type corresponding to each distortion feature vector, the method further includes: The PCA algorithm is used to reduce the dimension of each distorted feature vector to obtain each distorted feature vector after the dimension reduction, and each distorted feature vector after the dimension reduction is determined as the distorted feature vector.
7. The method according to claim 1, characterized in that The total model risk is obtained by using the error vector, the regularization coefficient and the output weight vector in the IRVFL classifier, including: The total model risk is obtained by the following formula: ; Wherein, E is the total model risk, C is a preset constant value, is the error vector, is the regularization coefficient, is the output weight vector.
8. The method according to claim 1, characterized in that After completing the training of the distorted image classification model, the method further includes: Inputting the verification sample into the convolutional neural network for feature extraction to obtain a distortion feature vector of each distorted image in the verification sample; wherein the verification sample includes a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting the distortion feature vector of each distorted image into the trained distorted image classification model to obtain the predicted distortion type of each distorted image in the verification sample; Obtaining the accuracy of the trained distorted image classification model according to the annotated distortion type of each distorted image in the verification sample and the predicted distortion type of each distorted image; If the accuracy is not greater than a specified threshold, the distorted image classification model is retrained.
9. A distorted image classification device, characterized in that: The device comprises: A distortion type determination module, used for inputting a distorted image to be detected into a pre-trained distorted image classification model to obtain the distortion type of the distorted image; the distorted image classification model includes a convolutional neural network and an IRVFL classifier, and the IRVFL classifier is a classifier obtained by introducing a regularization coefficient into an RVFL classifier; The training module is used to train the distorted image classification model in the following manner: Inputting the training samples into the convolution neural network for feature extraction to obtain each distortion feature vector, wherein the training samples include a plurality of distorted images and annotated distortion types corresponding to the plurality of distorted images; Inputting each distortion feature vector into the IRVFL classifier for classification to obtain a predicted distortion type corresponding to each distortion feature vector; Obtaining an error vector according to the marked distortion type and the predicted distortion type; Obtaining a total model risk through the error vector, the regularization coefficient, and the output weight vector in the IRVFL classifier; If the total model risk does not meet the specified conditions, the model parameters and output weight vectors in the IRVFL classifier are adjusted, and then the step of inputting each distortion feature vector into the IRVFL classifier for classification is returned until the total model risk meets the specified conditions, and the training of the distorted image classification model is terminated.
10. An electronic device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory, and execute the steps included in any one of the methods of claims 1-8 according to the obtained program instructions.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Calibration method and device of image acquisition equipment, readable storage medium and system
CN110610522A
Single-phase PWM rectifier power device IGBT open-circuit fault diagnosis method based on RVFL
CN115828094A
Fabric defect classification method for optimizing RVFL based on RIME optimization algorithm
CN118262147A
Contrastive explanations for images with monotonic attribute functions
US20210056355A1