Image classification method and system based on neural network

Through the combination of dimensionality reduction shear processing and lightweight recognition models, combined with the weight adjustment and secondary recognition of differentiated databases, the problems of waste of computing resources and feature omissions in the existing technology that are prone to confusing category recognition are solved, and efficient and accurate image classification is achieved.

CN119992224AActive Publication Date: 2025-05-13CHINA UNICOM (SHANDONG) IND INTERNET CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510457709.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

Smart Images

  • Figure CN119992224A_ABST
    Figure CN119992224A_ABST
Patent Text Reader

Abstract

The invention discloses an image classification method and system based on a neural network, mainly relates to the technical field of image classification, and is used for solving the problems that an existing multi-stage neural network mechanism method easily causes computing resource waste or key feature omission and an existing dimension reduction processing method may lose key discrimination information. Comprising the following steps: cutting a to-be-recognized image corresponding to a wrong first recognition result, and inputting the to-be-recognized image into an image recognition model to obtain a second recognition result; when the second recognition result is consistent with the first recognition result, determining next distinguishing points from the distinguishing target information in sequence according to the relationship between the weight values of the distinguishing target information, and determining the next distinguishing point which can determine that the first recognition result is correct for the first time as a distinguishing point to be adjusted; and adjusting the distinguishing target weight value of the distinguishing point to be adjusted to be greater than the distinguishing target weight value of the first distinguishing point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image classification, and in particular to a neural network-based image classification method and system. Background Art

[0002] Existing neural network-based image classification technologies generally use end-to-end deep learning models for feature extraction and classification decisions, but they have significant limitations when dealing with easily confused categories with similar visual features. Traditional methods usually improve classification accuracy by increasing model depth or introducing attention mechanisms, but such solutions will lead to a surge in computational complexity and make it difficult to perform targeted optimization for specific error-prone categories.

[0003] Some improvement schemes attempt to improve reliability through a multi-stage neural network mechanism, such as superimposing a neural network module after preliminary classification. However, there is a defect of rigid verification strategy, which is mainly reflected in the fixed neural network logic and the fixed neural network judgment benchmark. It consumes a lot of computing power during operation and has a long running time, which easily leads to waste of computing resources or omission of key features.

[0004] In addition, although existing dimensionality reduction processing methods (such as principal component analysis or convolutional downsampling) can reduce the amount of input data, they are not deeply integrated with the feature enhancement requirements of error-prone categories and may lose key discriminant information. For example, in the garbage sorting scenario, the combination of local reflective pixels of plastic bottles and glass bottles plays a decisive role in classification, while traditional dimensionality reduction operations will weaken such details. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present application provides a neural network-based image classification method and system to solve the problem that the existing multi-stage neural network mechanism method easily leads to waste of computing resources or omission of key features, and the existing dimensionality reduction processing method may lose key discriminant information.

[0006] In a first aspect, the present application provides an image classification method based on a neural network, the method comprising: The image to be identified is subjected to dimensionality reduction and shearing processing, and then input into a preset lightweight recognition model to obtain a first recognition result; when the first recognition result belongs to a preset error-prone type, a preset distinction database is called to obtain distinction target information and distinction target weight values; wherein the distinction target information at least includes: target pixel combination, target contour, texture feature, color distribution feature, edge gradient distribution, and the distinction target weight values ​​at least include: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, edge gradient distribution weight value; according to the relationship between the distinction target weight values, a first distinction point is determined from the distinction target information, and according to whether the first distinction point exists in the image to be identified, whether the first recognition result is correct is determined; the image to be identified corresponding to the erroneous first recognition result is sheared, and then input into the image recognition model to obtain a second recognition result; when the second recognition result is consistent with the first recognition result, according to the relationship between the distinction target information weight values, the next distinction point is determined from the distinction target information in turn, and the next distinction point that can determine the first recognition result to be correct for the first time is determined as the distinction point to be adjusted; the distinction target weight value of the distinction point to be adjusted is adjusted to be greater than the distinction target weight value of the first distinction point.

[0007] In one implementation of the present application, before performing dimensionality reduction and shearing processing on the image to be recognized and then inputting the image into a preset lightweight recognition model to obtain a first recognition result, the method further includes: Construct a preset lightweight recognition model with 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the stride is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; Add a global average pooling layer instead of a fully connected layer at the end of the network of the preset lightweight recognition model.

[0008] In one implementation of the present application, the image to be recognized is subjected to dimensionality reduction and shearing processing, and then input into a preset lightweight recognition model to obtain a first recognition result, which specifically includes: The resolution of the image to be identified is compressed to the preset resolution through the convolution downsampling layer and the high-frequency features are retained. At the same time, the Grad-CAM algorithm is used to locate the salient area and crop it to obtain the identification area. The principal component analysis technique is used to map the RGB three channels of the recognition area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

[0009] In one implementation of the present application, when the first recognition result belongs to a preset error-prone type, before calling a preset distinction database to obtain distinction target information, the method further includes: Through the preset information adjustment interface, operations such as adding, deleting, modifying and checking the distinguishing target information can be performed.

[0010] In one implementation of the present application, determining the first distinguishing point from the distinguishing target information according to the relationship between the distinguishing target weight values ​​specifically includes: The distinguishing target information with the largest distinguishing target weight value is determined as the first distinguishing point.

[0011] In a second aspect, the present application provides an image classification system based on a neural network, the system comprising: The acquisition module is used to perform dimensionality reduction and shearing processing on the image to be identified, and then input the preset lightweight recognition model to obtain the first recognition result; the acquisition module is used to call the preset distinction database when the first recognition result belongs to the preset error-prone type, and obtain the distinction target information and the distinction target weight value; wherein the distinction target information at least includes: target pixel combination, target contour, texture feature, color distribution feature, edge gradient distribution, and the distinction target weight value at least includes: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, edge gradient distribution weight value; the result module is used to distinguish the target weight value according to the relationship between the distinction target weight values. A system is provided for determining a first distinguishing point from distinguishing target information, and determining whether a first recognition result is correct according to whether the first distinguishing point exists in the image to be recognized; the image to be recognized corresponding to the erroneous first recognition result is cut and then input into the image recognition model to obtain a second recognition result; an adjustment module is provided for determining the next distinguishing point from the distinguishing target information in turn according to the relationship between the weight values ​​of the distinguishing target information when the second recognition result is consistent with the first recognition result, and determining the next distinguishing point that can determine the first recognition result to be correct for the first time as the distinguishing point to be adjusted; and adjusting the distinguishing target weight value of the distinguishing point to be adjusted to be greater than the distinguishing target weight value of the first distinguishing point.

[0012] In one implementation of the present application, the acquisition module includes a model building unit, Used to build a preset lightweight recognition model with 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the stride is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; Add a global average pooling layer instead of a fully connected layer at the end of the network of the preset lightweight recognition model.

[0013] In one implementation of the present application, the acquisition module includes a shearing unit, It is used to compress the resolution of the image to be identified to the preset resolution through the convolution downsampling layer and retain the high-frequency features. At the same time, it combines the Grad-CAM algorithm to locate the salient area and crop it to obtain the identification area; The principal component analysis technique is used to map the RGB three channels of the recognition area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

[0014] In one implementation of the present application, the acquisition module includes an acquisition unit, Used to add, delete, modify and check the target information through the preset information adjustment interface.

[0015] In one implementation of the present application, the adjustment module includes an adjustment unit. The distinction target information for determining the maximum distinction target weight value is the first distinction point.

[0016] Those skilled in the art can understand that the present application has at least the following beneficial effects: This application provides an image classification method and system based on neural network: 1. Improvements to the problem of wasted computing resources: This application adopts a combination of "dimensionality reduction and shearing processing + lightweight recognition model". By first reducing the dimension to reduce the amount of data, and then using a lightweight model for preliminary recognition, the computational complexity can be reduced. Compared with the full-scale calculation mode of traditional multi-stage neural networks, this application only calls complex models for error-prone types, effectively avoiding resource waste. This technical feature achieves a balance between computational efficiency and recognition accuracy through a dynamic resource allocation strategy.

[0017] When the initial recognition result is an error-prone type, the high-energy recognition model is called only when necessary through the progressive verification process of "differential database comparison → weight adjustment → secondary recognition". This hierarchical processing avoids the defects of blindly using complex models in the existing technology.

[0018] 2. Improvements to the problem of missing key features: This application includes five types of target information, including target pixel combination, texture features, color distribution features, etc., and dynamically determines key discrimination points through weight value relationships (for example, the first discrimination point that can be verified to be correct is used as the point to be adjusted). Compared with the traditional fixed weight dimensionality reduction method, this application retains high-frequency key features through a weight feedback mechanism, solving the problem of loss of discrimination information caused by static feature selection in the prior art.

[0019] The preset error-prone feature weights in the distinguishing database (such as the edge gradient distribution weight value) essentially form a feature enhancement mechanism for specific scenarios. The adaptive optimization of feature discrimination capability is achieved through the gradient adjustment of the weight value (for example, the weight value of the point to be adjusted is greater than the previous distinguishing point).

[0020] 3. Systematic improvement of technical effects: This application uses the coordinated cooperation of lightweight models and complex models to assign most of the routine recognition tasks to low-energy models while ensuring the recognition accuracy (secondary recognition verification). The overall efficiency improvement is in line with the judgment basis of the three-step method on the actual solution effect of technical problems.

[0021] The weight adjustment mechanism enables the system to dynamically optimize the feature discrimination priority according to the actual recognition results (for example, the adjustment of color distribution weight value > texture weight value). This closed-loop feedback mechanism breaks through the limitations of the traditional neural network fixed parameter mode. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0023] Figure 1 This is a flow chart of a neural network-based image classification method provided in an embodiment of the present application.

[0024] Figure 2 It is a schematic diagram of the internal structure of a neural network-based image classification system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] It should be understood by those skilled in the art that the embodiments described below are only preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work should still fall within the protection scope of the present disclosure.

[0026] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0027] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0028] The embodiment provides an image classification method based on a neural network, such as Figure 1 As shown, the method provided in the embodiment of the present application mainly includes the following steps: Step 110: Perform dimensionality reduction and shearing processing on the image to be recognized, and then input it into a preset lightweight recognition model to obtain a first recognition result.

[0029] In some embodiments, before performing dimensionality reduction and shearing processing on the image to be recognized and then inputting the image into a preset lightweight recognition model to obtain a first recognition result, the method further includes: Construct a preset lightweight recognition model with 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the stride is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; Add a global average pooling layer instead of a fully connected layer at the end of the network of the preset lightweight recognition model.

[0030] It is understandable that this step uses a small 3×3 convolution kernel and a step size of 2×2 to reduce the number of model parameters and the amount of calculation, thereby improving the recognition speed. Depthwise separable convolution decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution, which can reduce the number of model parameters and the storage requirements of the model. The global average pooling layer can avoid the overfitting problem caused by the fully connected layer, while retaining important features in the image, which helps to improve recognition accuracy.

[0031] The image to be recognized is subjected to dimensionality reduction and shearing processing, and then input into a preset lightweight recognition model to obtain a first recognition result, which specifically includes: The resolution of the image to be identified is compressed to the preset resolution through the convolution downsampling layer and the high-frequency features are retained. At the same time, the Grad-CAM algorithm is used to locate the salient area and crop it to obtain the identification area. The principal component analysis technique is used to map the RGB three channels of the recognition area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

[0032] It should be noted that due to the use of lightweight recognition models and dimensionality reduction and shearing processing, the entire recognition process is more efficient and can complete the image recognition task in a shorter time. By focusing on key areas and removing redundant information, the model can more accurately identify target objects or scenes in the image.

[0033] For example, Assume that the application scenario of license plate recognition needs to identify the license plate in the vehicle image. In this scenario, you can follow the above steps: Build a lightweight recognition model with 5 layers of convolution, each with a kernel size of 3×3 and a stride of 2×2, and use depthwise separable convolution. Add a global average pooling layer at the end of the model. Perform dimensionality reduction on the vehicle image to remove noise and redundant information in the image. Then, perform shearing according to the position of the license plate in the image, retaining only the license plate area. Input the license plate image after dimensionality reduction and shearing into the lightweight recognition model, the model performs feature extraction and classification operations on the license plate image, and outputs the license plate number as the first recognition result. In this way, while ensuring recognition accuracy, the efficiency of license plate recognition can be improved to meet the needs of actual application scenarios.

[0034] Step 120: When the first recognition result belongs to a preset error-prone type, call a preset distinction database to obtain distinction target information and a distinction target weight value.

[0035] It should be noted that the target information for distinguishing includes at least: target pixel combination, target contour, texture features, color distribution features, and edge gradient distribution; the target weight value for distinguishing includes at least: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, and edge gradient distribution weight value.

[0036] It is understandable that this step can quickly identify possible recognition errors by presetting the error-prone type judgment, providing a prerequisite for the subsequent call to the distinction database. The distinction database stores specific information and weight values ​​for error-prone types. Calling this database can provide strong support for subsequent feature comparison and weight adjustment. By obtaining detailed target information and corresponding weight values, image features can be more accurately compared and analyzed, thereby improving recognition accuracy. Special processing for error-prone types enhances the system's robustness to complex images and noise.

[0037] When the first recognition result belongs to the preset error-prone type, before calling the preset distinction database to obtain the distinction target information, the method may further include: Through the preset information adjustment interface, operations such as adding, deleting, modifying and checking the distinguishing target information can be performed.

[0038] It is understandable that the present application allows users to flexibly manage the distinguishing target information through the interface, thereby improving the maintainability and adaptability of the system. Adjusting the distinguishing target information according to actual needs helps to optimize the recognition performance and improve the recognition efficiency and accuracy.

[0039] Specific examples: Assume an application scenario of license plate recognition, in which some special license plates are easily misrecognized due to their special format and complex characters. In this case, you can follow the steps below: First, the license plate image is preliminarily recognized to obtain a first recognition result.

[0040] Determine whether the first recognition result belongs to a preset error-prone type.

[0041] If it is an error-prone type, the preset distinction database is called to obtain the distinction target information and weight values ​​for these special license plates. For example, for police license plates, it may be necessary to focus on the contour features, color distribution features, and edge gradient distribution of the license plate.

[0042] Through the preset information adjustment interface, necessary operations of adding, deleting, modifying and checking are performed on the obtained distinguishing target information to ensure the accuracy and completeness of the information.

[0043] The obtained distinguishing target information and weight values ​​are used to identify and analyze the license plate image again, and finally an accurate recognition result is obtained.

[0044] In this way, the accuracy and robustness of license plate recognition can be improved, especially when dealing with special license plates, better recognition results can be achieved.

[0045] Step 130: determine the first distinguishing point from the distinguishing target information according to the relationship between the distinguishing target weight values, and determine whether the first recognition result is correct according to whether the first distinguishing point exists in the image to be recognized; cut the image to be recognized corresponding to the erroneous first recognition result, and then input it into the image recognition model to obtain the second recognition result.

[0046] It should be noted that in the image recognition process, the first distinction point is determined from the distinction target information according to the relationship between the distinction target weight values, and the correctness of the first recognition result is judged accordingly. When the first recognition result is wrong, the corresponding image to be recognized is cut and re-recognized, which enhances the robustness of the recognition.

[0047] By focusing on the distinguishing target information with the largest weight value, the key features in the image can be more accurately identified, thereby improving the accuracy of recognition. Focusing the recognition focus on the distinguishing target information with the largest weight value helps reduce unnecessary calculations and optimize the recognition process.

[0048] ‌Example‌: Assume that in the license plate recognition scenario, the distinguishing target information includes the outline, color, characters and other features of the license plate, and the corresponding weight values ​​are 0.3, 0.2, and 0.5 respectively. At this time, the weight value of the character feature is the largest, so it is determined as the first distinguishing point. During the recognition process, the system will focus on the features of the license plate characters to improve the accuracy of recognition.

[0049] By checking whether the first distinguishing point exists, the correctness of the first recognition result can be quickly verified without complicated comparison and analysis, which helps to reduce unnecessary repeated recognition and improve overall recognition efficiency.

[0050] ‌Example‌: Continuing with the example of license plate recognition, if the first recognition result incorrectly recognizes the license plate characters, and the first distinguishing point (character feature) does exist in the image to be recognized but is not correctly recognized, then the system will determine that the first recognition result is wrong.

[0051] When the first recognition result is judged to be wrong, the corresponding image to be recognized is cut to remove redundant information, and then re-input into the image recognition model for recognition to obtain a second recognition result.

[0052] Removing redundant information through clipping helps the model to more accurately identify key features in the image and enhances the robustness of recognition. Re-recognition helps correct errors in the first recognition result and improves the overall recognition accuracy.

[0053] ‌Example‌: In the license plate recognition scenario, if the first recognition result mistakenly identifies the license plate characters as other characters, the system will crop the license plate image to remove redundant information such as the background, and then re-enter the recognition model for recognition. Assuming that the second recognition correctly identifies the license plate characters, the second recognition result is the correct recognition result.

[0054] According to the relationship between the weight values ​​of the distinguishing targets, the first distinguishing point is determined from the distinguishing target information, which may be specifically: The distinguishing target information with the largest distinguishing target weight value is determined as the first distinguishing point.

[0055] Step 140, when the second recognition result is consistent with the first recognition result, determine the next distinction point from the distinction target information in turn according to the relationship between the weight values ​​of the distinction target information, and determine the next distinction point that can first determine that the first recognition result is correct as the distinction point to be adjusted; adjust the distinction target weight value of the distinction point to be adjusted to be greater than the distinction target weight value of the first distinction point.

[0056] It should be noted that by determining the next distinguishing point in sequence, the accuracy of recognition can be verified step by step to ensure that the recognition result is correct, which helps to optimize the recognition process and improve recognition efficiency.

[0057] ‌Example‌: Assume that in the license plate recognition scenario, the distinguishing target information includes the features of the license plate's contour, color, characters, and position, and the corresponding weight values ​​are 0.3, 0.2, 0.4, and 0.1, respectively. In the first round of recognition, the character feature is determined as the first distinguishing point, and the recognition result is judged based on this. In the second round of recognition, when the second recognition result is consistent with the first recognition result, the system will consider the features such as contour, color, and position as the next distinguishing point for verification.

[0058] By determining the difference points to be adjusted, the reliability of recognition can be further improved, the accuracy of recognition results can be ensured, and it is helpful to optimize the weight distribution of the recognition model and improve the recognition performance of the model.

[0059] ‌Example‌: Continuing with the example of license plate recognition, suppose that in the second round of recognition, the contour feature is verified as the next distinguishing point and is the feature that can determine the correctness of the first recognition result for the first time. At this time, the contour feature will be determined as the distinguishing point to be adjusted.

[0060] By adjusting the weight value, the weight distribution of the recognition model can be optimized, so that the model pays more attention to the features that have an important impact on the recognition results. This helps to improve the adaptability of the model and enable it to better cope with complex and changing recognition scenarios.

[0061] ‌Example‌: In the license plate recognition scenario, if the contour feature is determined as a distinguishing point to be adjusted, then its corresponding weight value will be adjusted to be greater than the weight value of the character feature. In this way, in the subsequent recognition process, the model will pay more attention to the contour features of the license plate, thereby improving the accuracy and reliability of recognition.

[0062] Based on the previous description, this embodiment adopts a combination of "dimensionality reduction and shearing processing + lightweight recognition model". By first reducing the dimension to reduce the amount of data, and then using a lightweight model for preliminary recognition, the computational complexity can be reduced. Compared with the full-scale calculation mode of traditional multi-stage neural networks, this application only calls complex models for error-prone types, effectively avoiding waste of resources. This technical feature achieves a balance between computational efficiency and recognition accuracy through a dynamic resource allocation strategy. When the preliminary recognition result is an error-prone type, the high-energy recognition model is called only when necessary through the progressive verification process of "distinguishing database comparison → weight adjustment → secondary recognition". This layered processing avoids the defects of blindly using complex models in the prior art.

[0063] This embodiment includes five types of target information such as "target pixel combination, texture features, color distribution features", and dynamically determines key discrimination points through weight value relationships (such as "the first discrimination point that can be verified to be correct is used as the point to be adjusted"). Compared with the traditional fixed weight dimensionality reduction method, this solution retains high-frequency key features through a weight feedback mechanism, solving the problem of loss of discrimination information caused by static feature selection in the prior art. The preset error-prone type feature weights in the discrimination database (such as edge gradient distribution weight values) essentially form a feature enhancement mechanism for specific scenarios. Through the gradient adjustment of weight values ​​(such as "the weight value of the point to be adjusted is greater than the previous discrimination point"), adaptive optimization of feature discrimination capabilities is achieved.

[0064] This embodiment uses the collaboration of lightweight models and complex models to assign most of the routine recognition tasks to low-energy models while ensuring recognition accuracy (secondary recognition verification). The overall efficiency improvement is in line with the judgment basis of the actual solution effect of technical problems in the three-step method. The weight adjustment mechanism enables the system to dynamically optimize the feature discrimination priority (such as the adjustment of "color distribution weight value > texture weight value") according to the actual recognition results. This closed-loop feedback mechanism breaks through the limitations of the traditional neural network fixed parameter mode.

[0065] In addition, this application Figure 2 The present invention provides an image classification system based on a neural network. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes: The acquisition module 210 is used to perform dimensionality reduction and shearing processing on the image to be recognized, and then input it into a preset lightweight recognition model to obtain a first recognition result.

[0066] The acquisition module 210 includes a model building unit, which is used to construct a preset lightweight recognition model including 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the step size is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; a global average pooling layer is added at the end of the network of the preset lightweight recognition model instead of the fully connected layer.

[0067] The acquisition module 210 includes a shearing unit, which is used to compress the resolution of the image to be identified to a preset resolution and retain high-frequency features through a convolution downsampling layer, and at the same time use the Grad-CAM algorithm to locate the significant area and crop to obtain the identification area; the principal component analysis technology is used to map the RGB three channels of the identification area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

[0068] The acquisition module 220 is used to call the preset distinction database to obtain the distinction target information and the distinction target weight value when the first recognition result belongs to the preset error-prone type; wherein the distinction target information at least includes: target pixel combination, target contour, texture feature, color distribution feature, edge gradient distribution, and the distinction target weight value at least includes: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, edge gradient distribution weight value.

[0069] The acquisition module 220 includes an acquisition unit, which is used to perform addition, deletion, modification and query operations on the distinguishing target information through a preset information adjustment interface.

[0070] The result module 230 is used to determine the first distinguishing point from the distinguishing target information according to the relationship between the distinguishing target weight values, and determine whether the first recognition result is correct according to whether the first distinguishing point exists in the image to be recognized; the image to be recognized corresponding to the erroneous first recognition result is cut, and then input into the image recognition model to obtain the second recognition result.

[0071] The adjustment module 240 is used to determine the next distinction point from the distinction target information in turn according to the relationship between the distinction target information weight values ​​when the second recognition result is consistent with the first recognition result, and determine the next distinction point that can first determine that the first recognition result is correct as the distinction point to be adjusted; adjust the distinction target weight value of the distinction point to be adjusted to be greater than the distinction target weight value of the first distinction point.

[0072] The adjustment module 240 includes an adjustment unit, which is used to determine the distinguishing target information with the largest distinguishing target weight value as the first distinguishing point.

[0073] So far, the technical solutions of the present disclosure have been described in combination with the above multiple embodiments, but it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principles of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.

Claims

1. A neural network-based image classification method, characterized in that: The method comprises: Performing dimensionality reduction and shearing processing on the image to be recognized, and then inputting it into a preset lightweight recognition model to obtain a first recognition result; When the first recognition result belongs to the preset error-prone type, calling the preset distinction database to obtain the distinction target information and the distinction target weight value; wherein the distinction target information at least includes: target pixel combination, target contour, texture feature, color distribution feature, edge gradient distribution, and the distinction target weight value at least includes: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, edge gradient distribution weight value; According to the relationship between the weight values ​​of the distinguishing targets, a first distinguishing point is determined from the distinguishing target information, and whether the first distinguishing point exists in the image to be recognized is determined whether the first recognition result is correct; the image to be recognized corresponding to the wrong first recognition result is cut and then input into the image recognition model to obtain a second recognition result; When the second recognition result is consistent with the first recognition result, the next distinction point is determined from the distinction target information in turn according to the relationship between the weight values ​​of the distinction target information, and the next distinction point that can first determine that the first recognition result is correct is determined as the distinction point to be adjusted; the distinction target weight value of the distinction point to be adjusted is adjusted to be greater than the distinction target weight value of the first distinction point.

2. The neural network-based image classification method according to claim 1, characterized in that: Before performing dimensionality reduction and shearing processing on the image to be recognized and then inputting it into a preset lightweight recognition model to obtain a first recognition result, the method further includes: Construct a preset lightweight recognition model with 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the stride is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; Add a global average pooling layer instead of a fully connected layer at the end of the network of the preset lightweight recognition model.

3. The neural network-based image classification method according to claim 1, characterized in that: The image to be recognized is subjected to dimensionality reduction and shearing processing, and then input into a preset lightweight recognition model to obtain a first recognition result, which specifically includes: The resolution of the image to be identified is compressed to the preset resolution through the convolution downsampling layer and the high-frequency features are retained. At the same time, the Grad-CAM algorithm is used to locate the salient area and crop it to obtain the identification area. The principal component analysis technique is used to map the RGB three channels of the recognition area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

4. The neural network-based image classification method according to claim 1, characterized in that: When the first recognition result belongs to the preset error-prone type, before calling the preset distinction database to obtain the distinction target information, the method further includes: Through the preset information adjustment interface, operations such as adding, deleting, modifying and checking the distinguishing target information can be performed.

5. The neural network-based image classification method according to claim 1, characterized in that: According to the relationship between the weight values ​​of the distinguishing targets, determining the first distinguishing point from the distinguishing target information specifically includes: The distinguishing target information with the largest distinguishing target weight value is determined as the first distinguishing point.

6. A neural network-based image classification system, characterized in that: The system comprises: An acquisition module, used for performing dimensionality reduction and shearing processing on the image to be recognized, and then inputting the image into a preset lightweight recognition model to obtain a first recognition result; An acquisition module, used for calling a preset distinction database to acquire distinction target information and a distinction target weight value when the first recognition result belongs to a preset error-prone type; wherein the distinction target information at least includes: target pixel combination, target contour, texture feature, color distribution feature, edge gradient distribution, and the distinction target weight value at least includes: target pixel weight value, target contour weight value, texture feature weight value, color distribution feature weight value, edge gradient distribution weight value; A result module is used to determine the first distinguishing point from the distinguishing target information according to the relationship between the distinguishing target weight values, and determine whether the first recognition result is correct according to whether the first distinguishing point exists in the image to be recognized; the image to be recognized corresponding to the wrong first recognition result is cut and then input into the image recognition model to obtain a second recognition result; The adjustment module is used to determine the next distinction point from the distinction target information in turn according to the relationship between the weight values ​​of the distinction target information when the second recognition result is consistent with the first recognition result, and determine the next distinction point that can first determine that the first recognition result is correct as the distinction point to be adjusted; adjust the distinction target weight value of the distinction point to be adjusted to be greater than the distinction target weight value of the first distinction point.

7. The neural network-based image classification system according to claim 6, characterized in that: The acquisition module includes a model building unit, Used to build a preset lightweight recognition model with 5 layers of convolution, and the convolution kernel size of each layer is set to 3×3, the stride is 2×2, and the depthwise separable convolution is used instead of the standard convolution operation; Add a global average pooling layer instead of a fully connected layer at the end of the network of the preset lightweight recognition model.

8. The neural network-based image classification system according to claim 6, characterized in that: The acquisition module includes a shear unit, It is used to compress the resolution of the image to be identified to the preset resolution through the convolution downsampling layer and retain the high-frequency features. At the same time, it combines the Grad-CAM algorithm to locate the salient area and crop it to obtain the identification area; The principal component analysis technique is used to map the RGB three channels of the recognition area into a single-channel grayscale image to obtain the image after dimensionality reduction and shearing processing.

9. The neural network-based image classification system according to claim 6, characterized in that: The acquisition module includes an acquisition unit, Used to add, delete, modify and check the target information through the preset information adjustment interface.

10. The neural network-based image classification system according to claim 6, characterized in that: The adjustment module includes an adjustment unit, The distinction target information for determining the maximum distinction target weight value is the first distinction point.

Citation Information

Patent Citations

  • License plate recognition method and device

    CN111832337A

  • License plate recognition system based on deep neural network

    CN112232351A

  • Artificial intelligence-based image recognition method and related device

    CN113569889A

  • License plate recognition method and device, terminal and computer readable storage medium

    CN114882492A

  • Character recognition method, device, equipment, system and medium

    CN117809307A