High-precision image recognition algorithm and system based on deep learning

By using the ResNet model in deep learning-based image recognition algorithm and integrating adversarial sample defense mechanism, the problem of algorithms identifying errors when facing new images is solved, improving generalization capabilities and avoiding the risk of adversarial samples.

CN120014355APending Publication Date: 2025-05-16SOFT CLOUD IND NETWORK DATA (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510113403.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing deep learning-based image recognition algorithms are prone to recognition errors when facing new and unseen images, and there are security risks of identification errors caused by adversarial samples.

Method used

ResNet is used as the basic model and different adversarial sample defense mechanisms are integrated to generate multiple alternative models. Improve the model's defense capabilities through adversarial training, gradient masking and adversarial sample detection. Data preprocessing and enhancement techniques improve data quality, train and evaluate models to select the optimal model.

Benefits of technology

It significantly improves the generalization ability of image recognition algorithms, reduces the occurrence of new and unseen image recognition errors, and effectively avoids recognition errors caused by adversarial samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014355A_ABST
    Figure CN120014355A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision image recognition algorithm and system based on deep learning, and relates to the technical field of image recognition, and the algorithm comprises the steps: S1, selecting ResNet as a basic model, sequentially integrating different antagonistic sample defense mechanisms, and generating a plurality of alternative models; s2, collecting and marking image data, and preprocessing and enhancing the data; s3, training the alternative model by using the training data set; s4, weighing the precision and calculation complexity of the model, selecting an optimal alternative model as a final model, and performing optimization and adjustment; and S5, deploying the final model to an actual application scene, continuously monitoring the performance of the model, and performing updating and optimization according to actual requirements. According to the high-precision image recognition algorithm and system based on deep learning, the generalization ability is high, recognition errors are not likely to occur in the face of new and unseen images, and recognition errors caused by adversarial samples can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and specifically to a high-precision image recognition algorithm and system based on deep learning. Background Art

[0002] Image recognition is a technology that uses computers to process, analyze and understand images to identify objects, scenes or information in the image. This technology has been widely used in various fields, such as medical image analysis, autonomous driving, security monitoring, face recognition, etc. Image recognition algorithms are the core of image recognition. These algorithms are usually based on knowledge in fields such as mathematics, statistics and computer science, and process and analyze images through specific calculation methods and processes. There are many types of image recognition algorithms, including traditional image processing algorithms and algorithms based on deep learning.

[0003] Traditional image processing algorithms mainly rely on manually designed features and classifiers. They recognize objects in images by preprocessing, extracting features, and classifying them. These algorithms usually require more complex image processing, and the recognition effect is limited by feature design and classifier performance. With the development of deep learning technology, image recognition algorithms based on deep learning have gradually become mainstream. These algorithms automatically learn feature representations in images by building deep neural network models, and use these features for classification and recognition. Deep learning algorithms have powerful feature learning and generalization capabilities, so they have achieved remarkable results in the field of image recognition.

[0004] However, in actual use, although the existing deep learning-based image recognition algorithms show high accuracy in specific fields and scenarios, their generalization ability is still limited, and they are prone to recognition errors when faced with some new and unseen images. In addition, existing image recognition algorithms have certain security risks. For example, through technologies such as generative adversarial networks (GANs), some adversarial samples can be generated. These adversarial samples may be very similar to real images visually, but they can deceive image recognition algorithms and cause the algorithms to produce incorrect recognition results.

[0005] Therefore, there is an urgent need to improve this shortcoming. The present invention studies and improves the existing technology and its shortcomings, and provides a high-precision image recognition algorithm and system based on deep learning. Summary of the invention

[0006] The purpose of the present invention is to provide a high-precision image recognition algorithm and system based on deep learning to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] In the first aspect, a high-precision image recognition algorithm based on deep learning comprises the steps of:

[0009] S1. Select ResNet as the basic model and integrate different adversarial sample defense mechanisms in turn to generate multiple candidate models;

[0010] S2, collect and annotate image data, preprocess and enhance the data;

[0011] S3, dividing the processed random data into a training data set and a test data set, and using the training data set to train the candidate model;

[0012] S4. Use the test data set to evaluate each candidate model, weigh the accuracy and computational complexity of the model, select the best candidate model as the final model, and optimize and adjust it, including adjusting hyperparameters, adding regularization terms, using dropout and other methods to prevent overfitting, and adjusting the model structure, such as adding convolutional layers, changing the size of the convolution kernel, etc., to further improve the performance of the model;

[0013] S5. Deploy the final model to actual application scenarios, such as image classification systems, object detection systems, etc., to ensure that the model can correctly receive input images and output prediction results, continuously monitor the performance of the model, and update and optimize it according to actual needs.

[0014] Furthermore, in step S1, each candidate model adopts a different defense mechanism or a combination of defense mechanisms, and the adversarial sample defense mechanism includes but is not limited to adversarial training, gradient masking, and adversarial sample detection;

[0015] The adversarial training: by introducing adversarial samples during the training process, the model learns how to resist these adversarial attacks. The implementation method is: using the adversarial attack algorithm to generate adversarial samples, and sending the generated adversarial samples together with the original samples to the ResNet model for training, and then using the test data set to evaluate the performance of the model, including indicators such as accuracy and recall rate;

[0016] The gradient masking: by modifying the activation function or other parameters of the model, it is difficult for the attacker to use the gradient information of the model to construct adversarial samples. This method reduces the success rate of adversarial attacks by increasing the sensitivity or uncertainty of the model to gradient information. The implementation method is: based on the ResNet model, the activation function or other parameters are modified to achieve gradient masking, and the modified model is used for training, and the performance of the model is evaluated using a test data set, including indicators such as the success rate of adversarial attacks;

[0017] The adversarial sample detection: by introducing a detection algorithm to identify whether the input sample is an adversarial sample and taking corresponding defense measures. These detection algorithms can make judgments based on model output, gradient information, abnormal behavior in the feature space, etc. The implementation method is: use adversarial samples and normal samples to train a detector to identify whether the input sample is an adversarial sample, and integrate the trained detector into the ResNet model to detect the input sample. According to the output results of the detector, take measures such as rejection, warning or correction for the input identified as an adversarial sample.

[0018] Furthermore, the anti-attack algorithm selects FGSM or PGD, as follows:

[0019] FGSM: Generate adversarial samples by adding a small perturbation with the same sign as the gradient to the original input sample; the specific steps of FGSM are:

[0020] Compute gradients: The gradient of the model with respect to the input sample needs to be calculated, which usually involves backpropagating the derivative of the model's loss function with respect to the input sample;

[0021] Generate perturbations: Generate a small perturbation based on the sign of the gradient. The size of this perturbation is controlled by a hyperparameter ε, which determines the strength of the perturbation.

[0022] Add perturbation: Add this perturbation to the original input sample to obtain an adversarial sample.

[0023] FGSM formula: Adversarial sample = original sample + ε*sign(gradient).

[0024] PGD: Generate stronger adversarial examples by applying the FGSM attack iteratively multiple times and projecting the adversarial examples back to the feasible set of the original data (usually the L∞ or L2 norm ball) after each iteration; the specific steps of PGD are:

[0025] Initialization: Start with an original sample, which can be considered as an adversarial sample for the first iteration;

[0026] Calculate gradient: In each iteration, calculate the gradient of the model for the current adversarial sample;

[0027] Update the adversarial example: Update the adversarial example according to the gradient, which usually involves adding a small perturbation in the direction of the gradient and may include a projection step to ensure that the adversarial example remains in the feasible set;

[0028] Iteration: Repeat the above steps until the predetermined number of iterations is reached or other stopping conditions are met.

[0029] PGD ​​formula (simplified version, without projection step): For each iteration k: adversarial_k+1 = adversarial_k + α*sign(gradient_k), where α is the learning rate, which controls the step size of the perturbation in each iteration. In practical applications, PGD usually also includes a projection step to ensure that the generated adversarial examples meet specific constraints (such as L∞ or L2 norm constraints).

[0030] Furthermore, the step S2 specifically includes the following steps:

[0031] S21. Data collection: collected image data, and the collected image data must be able to cover various situations in actual application scenarios;

[0032] S22, Data Annotation: Assign a correct label to each image for use in the training process;

[0033] S23, data preprocessing: including image normalization, resizing, image denoising and other operations to improve the model's ability to extract image features, as follows:

[0034] Normalization: Use min-max normalization (linearly scale pixel values ​​to a specified minimum and maximum value) or Z-score normalization (normalize according to the mean and standard deviation of the image so that the pixel values ​​conform to the standard normal distribution) to scale the pixel values ​​of the image to a fixed range, such as [0, 1] or [-1, 1], which helps speed up the convergence of the model and improve performance;

[0035] Resizing: resize the image to the same size for input into the model by scaling (scaling the image proportionally to the target size), cropping (cropping out a region of a specified size from the image, which can be a center crop or a random crop), and padding (adding a margin around the image to resize it, the padding value can be a constant, mirrored, or reflected, etc.);

[0036] Image denoising: Use filtering algorithms such as median filtering and Gaussian filtering to remove noise from the image (the noise may come from sensor noise during image acquisition, distortion during transmission, etc.) to improve the quality and clarity of the image;

[0037] S24, Data Augmentation: Generate new training samples by randomly transforming the original image to increase the diversity of the data, and the random transformation specifically includes:

[0038] Rotation: Rotate the image at random angles to enhance the model’s ability to recognize rotated objects;

[0039] Translation: Random translation of the image in the horizontal or vertical direction to simulate the situation where the object appears in different positions;

[0040] Scaling: Randomly adjust the size or scale of the image to help the model learn objects of different sizes;

[0041] Flip: Flip the image horizontally or vertically to enhance the model's ability to recognize different perspectives;

[0042] Cropping: Randomly select an area from the image for cropping to increase the diversity of the image. In the object detection task, cropping can simulate the position changes of objects in different scenes;

[0043] Color transformation: adjust the color attributes of the image, such as brightness, contrast, and saturation, to simulate different lighting conditions or shooting environments to improve the model's robustness to lighting changes;

[0044] Noise addition: Add random noise (such as Gaussian noise, salt and pepper noise, etc.) to the image to simulate the impact of different sensor noises on image quality and increase the robustness of the model.

[0045] Furthermore, in step S24, data enhancement also includes the following method:

[0046] Random erasing: Randomly cover an area on the image so that the pixel value of the area becomes a constant or random value. In this way, the model learns to correctly identify the target even when information is missing, which is used to improve the robustness of the model and prevent overfitting.

[0047] Mixed images: Linearly combine two images to form a new image. This method allows the model to learn more abstract features and enhance generalization capabilities.

[0048] Target box transformation: In the object detection task, the position and size of the target box are updated accordingly according to the transformation of the image (such as scaling, rotation, cropping, etc.).

[0049] Furthermore, in step S3, the test data set and the training data set are independent of each other, and the data distribution is consistent, and the training data set accounts for 70% to 80%, and the test data set accounts for 20% to 30%.

[0050] Furthermore, in step S3, the operation process of model training is:

[0051] Initialize the model: create a model instance and set hyperparameters such as learning rate, batch size, number of training rounds, etc.

[0052] Loss function and optimizer: select a loss function suitable for the task, such as cross entropy loss for classification problems and mean square error loss for regression problems, and select an optimization algorithm, such as Adam, SGD, etc., and set hyperparameters such as learning rate;

[0053] Training process: includes outer loop (epoch) and inner loop (batch);

[0054] The outer loop: traverses each round of training, and specifically performs the following operations: batch loading of training data through the data loader, and resetting indicators such as total loss, number of iterations, etc.;

[0055] The inner loop: traverses each batch of data, and the specific operations are:

[0056] Forward propagation: pass the input data through the model and calculate the output;

[0057] Calculate loss: Use the loss function to calculate the error between the model output and the target value;

[0058] Back propagation: calculate gradients, clear gradients, and perform back propagation;

[0059] Update parameters: Update model parameters through the optimizer;

[0060] Record indicators: accumulated loss, number of update iterations, etc.

[0061] Record logs: After each round, calculate the average loss and training time, and record log information.

[0062] Furthermore, in step S4, the model evaluation process: calculate the accuracy (such as precision, recall rate, F1 score, etc.) and computational complexity (such as model size, inference time, etc.) of the model, record the evaluation results of each model, including accuracy and computational complexity; and based on the evaluation results, weigh the model accuracy and computational complexity, as follows:

[0063] Accuracy trade-off: Determine the threshold of accuracy requirements based on actual needs, and select models with accuracy exceeding the threshold as candidate models;

[0064] Computational complexity trade-off: Considering the resource constraints in actual application scenarios (such as computing resources, storage resources, etc.), among the candidate models, the model with lower computational complexity and meeting the accuracy requirements is selected as the final model;

[0065] Comprehensive evaluation: If there are multiple models with similar performance in terms of accuracy and computational complexity, other factors such as model stability and interpretability are further considered to select the best candidate model as the final model.

[0066] Furthermore, in step S5, continuous monitoring is achieved by collecting data and evaluating it in actual application scenarios, and the evaluation indicators include accuracy, response time, etc., and the means of model updating and optimization include adding new training data, adjusting the model structure (such as the number of layers, the number of neurons, etc.) or parameters (such as learning rate, batch size, etc.), introducing new adversarial defense mechanisms, etc.

[0067] In a second aspect, a high-precision image recognition system based on deep learning is applied to the high-precision image recognition algorithm based on deep learning as described above, comprising:

[0068] Image input module: responsible for obtaining image data from various devices (such as cameras, scanners, etc.) and passing it as raw input data to subsequent modules;

[0069] Data processing module: responsible for data labeling, preprocessing, and data enhancement of image data;

[0070] Alternative model generation module: responsible for selecting ResNet as the basic model, and integrating different adversarial sample defense mechanisms in turn to generate multiple alternative models;

[0071] Data partitioning module: responsible for dividing the processed random data into training data set and test data set;

[0072] Model training module: responsible for training the candidate model using the training data set;

[0073] Final model generation module: responsible for evaluating each candidate model using the test data set, calculating the accuracy and computational complexity of the model, and weighing the accuracy and computational complexity of the model based on the evaluation results, and selecting the best candidate model as the final model;

[0074] Optimization and adjustment module: responsible for performing optimization operations such as hyperparameter tuning and model structure adjustment on the final model to improve the performance of the model;

[0075] Model deployment module: responsible for deploying the optimized final model to actual application scenarios, such as intelligent monitoring, human-computer interaction, autonomous driving and other fields;

[0076] Monitoring and update module: responsible for continuously monitoring the performance of the model, including indicators such as accuracy and response time. If the model performance deteriorates, the model will be updated and optimized according to changes in actual needs to adapt to new application scenarios and data distribution.

[0077] The present invention provides a high-precision image recognition algorithm and system based on deep learning, which has the following beneficial effects:

[0078] The present invention selects ResNet as the basic model, which lays a solid foundation for the generalization ability of the model, and generates multiple candidate models by integrating different adversarial sample defense strategies. It cooperates with data preprocessing and enhancement to improve data quality, and then uses the training data set to train the candidate model, and uses the test data set to evaluate each candidate model. The accuracy and computational complexity of the model are weighed, and the optimal candidate model is selected as the final model. After optimization and adjustment, the obtained image recognition algorithm has strong generalization ability, is not prone to recognition errors when facing new and unseen images, and can avoid recognition errors caused by adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 A schematic diagram of the steps of a high-precision image recognition algorithm based on deep learning in the present invention;

[0080] Figure 2 This is a general architecture diagram of a high-precision image recognition system based on deep learning in the present invention. DETAILED DESCRIPTION

[0081] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0082] like Figure 1 As shown, a high-precision image recognition algorithm based on deep learning includes the following steps:

[0083] S1. Select ResNet as the basic model, and integrate different adversarial sample defense mechanisms in turn to generate multiple candidate models; in this embodiment, each candidate model adopts a different defense mechanism or a combination of defense mechanisms, and the adversarial sample defense mechanism includes but is not limited to adversarial training, gradient masking, and adversarial sample detection;

[0084] Adversarial training: By introducing adversarial samples during the training process, the model learns how to resist these adversarial attacks. The implementation method is: Use the adversarial attack algorithm to generate adversarial samples, and send the generated adversarial samples and the original samples to the ResNet model for training, and then use the test data set to evaluate the performance of the model, including indicators such as accuracy and recall rate; Among them, the adversarial attack algorithm selects FGSM or PGD, as follows:

[0085] FGSM: Generate adversarial samples by adding a small perturbation with the same sign as the gradient to the original input sample; in this embodiment, the specific steps of FGSM are:

[0086] Compute gradients: The gradient of the model with respect to the input sample needs to be calculated, which usually involves backpropagating the derivative of the model's loss function with respect to the input sample;

[0087] Generate perturbations: Generate a small perturbation based on the sign of the gradient. The size of this perturbation is controlled by a hyperparameter ε, which determines the strength of the perturbation.

[0088] Add perturbation: Add this perturbation to the original input sample to obtain an adversarial sample.

[0089] FGSM formula: Adversarial sample = original sample + ε*sign(gradient).

[0090] PGD: Generate stronger adversarial samples by applying the FGSM attack multiple times iteratively and projecting the adversarial samples back to the feasible set of the original data (usually the L∞ or L2 norm sphere) after each iteration; in this embodiment, the specific steps of PGD are:

[0091] Initialization: Start with an original sample, which can be considered as an adversarial sample for the first iteration;

[0092] Calculate gradient: In each iteration, calculate the gradient of the model for the current adversarial sample;

[0093] Update the adversarial example: Update the adversarial example according to the gradient, which usually involves adding a small perturbation in the direction of the gradient and may include a projection step to ensure that the adversarial example remains in the feasible set;

[0094] Iteration: Repeat the above steps until the predetermined number of iterations is reached or other stopping conditions are met.

[0095] PGD ​​formula (simplified version, without projection step): For each iteration k: adversarial_k+1 = adversarial_k + α*sign(gradient_k), where α is the learning rate, which controls the step size of the perturbation in each iteration. In practical applications, PGD usually also includes a projection step to ensure that the generated adversarial examples meet specific constraints (such as L∞ or L2 norm constraints).

[0096] Gradient masking: By modifying the activation function or other parameters of the model, it is difficult for attackers to use the gradient information of the model to construct adversarial samples. This method reduces the success rate of adversarial attacks by increasing the sensitivity or uncertainty of the model to gradient information. Implementation method: Based on the ResNet model, modify the activation function or other parameters to achieve gradient masking, and use the modified model for training. Use the test data set to evaluate the performance of the model, including indicators such as the success rate of adversarial attacks;

[0097] Adversarial sample detection: By introducing detection algorithms to identify whether input samples are adversarial samples and taking corresponding defense measures, these detection algorithms can make judgments based on model output, gradient information, abnormal behavior in feature space, etc. Implementation method: Use adversarial samples and normal samples to train a detector to identify whether the input samples are adversarial samples, and integrate the trained detector into the ResNet model to detect the input samples. According to the output results of the detector, take measures such as rejection, warning or correction for the input identified as adversarial samples;

[0098] S2, collect and annotate image data, preprocess and enhance the data;

[0099] S21. Data collection: collected image data, and the collected image data must be able to cover various situations in actual application scenarios;

[0100] S22, Data Annotation: Assign a correct label to each image for use in the training process;

[0101] S23, data preprocessing: including image normalization, resizing, image denoising and other operations to improve the model's ability to extract image features, as follows:

[0102] Normalization: Use min-max normalization (linearly scale pixel values ​​to a specified minimum and maximum value) or Z-score normalization (normalize according to the mean and standard deviation of the image so that the pixel values ​​conform to the standard normal distribution) to scale the pixel values ​​of the image to a fixed range, such as [0, 1] or [-1, 1], which helps speed up the convergence of the model and improve performance;

[0103] Resizing: resize the image to the same size for input into the model by scaling (scaling the image proportionally to the target size), cropping (cropping out a region of a specified size from the image, which can be a center crop or a random crop), and padding (adding a margin around the image to resize it, the padding value can be a constant, mirrored, or reflected, etc.);

[0104] Image denoising: Use filtering algorithms such as median filtering and Gaussian filtering to remove noise from the image (the noise may come from sensor noise during image acquisition, distortion during transmission, etc.) to improve the quality and clarity of the image;

[0105] S24, Data Augmentation: Generate new training samples by randomly transforming the original image to increase the diversity of the data, and the random transformation specifically includes:

[0106] Rotation: Rotate the image at random angles to enhance the model’s ability to recognize rotated objects;

[0107] Translation: Randomly translate the image horizontally or vertically to simulate objects appearing in different positions.

[0108] Scaling: Randomly adjust the size or scale of the image to help the model learn objects of different sizes;

[0109] Flip: Flip the image horizontally or vertically to enhance the model's ability to recognize different perspectives;

[0110] Cropping: Randomly select an area from the image for cropping to increase the diversity of the image. In the object detection task, cropping can simulate the position changes of objects in different scenes;

[0111] Color transformation: adjust the color attributes of the image, such as brightness, contrast, and saturation, to simulate different lighting conditions or shooting environments to improve the model's robustness to lighting changes;

[0112] Noise addition: Add random noise (such as Gaussian noise, salt and pepper noise, etc.) to the image to simulate the impact of different sensor noises on image quality and increase the robustness of the model;

[0113] In addition, data enhancement also includes the following methods:

[0114] Random erasing: Randomly cover an area on the image so that the pixel value of the area becomes a constant or random value. In this way, the model learns to correctly identify the target even when information is missing, which is used to improve the robustness of the model and prevent overfitting.

[0115] Mixed images: Linearly combine two images to form a new image. This method allows the model to learn more abstract features and enhance generalization capabilities.

[0116] Target box transformation: In the target detection task, the position and size of the target box are updated accordingly according to the transformation of the image (such as scaling, rotation, cropping, etc.);

[0117] S3. Divide the processed random data into a training data set and a test data set. The test data set and the training data set are independent of each other and have the same data distribution. The training data set accounts for 70% to 80% and the test data set accounts for 20% to 30%. Use the training data set to train the candidate model. The training process is as follows:

[0118] Initialize the model: create a model instance and set hyperparameters such as learning rate, batch size, number of training rounds, etc.

[0119] Loss function and optimizer: select a loss function suitable for the task, such as cross entropy loss for classification problems and mean square error loss for regression problems, and select an optimization algorithm, such as Adam, SGD, etc., and set hyperparameters such as learning rate;

[0120] Training process: includes outer loop (epoch) and inner loop (batch);

[0121] Outer loop: traverses each round of training. Specific operations: batch load training data through the data loader and reset indicators such as total loss, number of iterations, etc.;

[0122] Inner loop: traverse each batch of data, specific operations:

[0123] Forward propagation: pass the input data through the model and calculate the output;

[0124] Calculate loss: Use the loss function to calculate the error between the model output and the target value;

[0125] Back propagation: calculate gradients, clear gradients, and perform back propagation;

[0126] Update parameters: Update model parameters through the optimizer;

[0127] Record indicators: accumulated loss, number of update iterations, etc.

[0128] Record logs: After each round, calculate the average loss and training time, and record log information;

[0129] S4. Use the test data set to evaluate each candidate model, weigh the accuracy and computational complexity of the model, select the best candidate model as the final model, and optimize and adjust it, including adjusting hyperparameters, adding regularization terms, using dropout and other methods to prevent overfitting, and adjusting the model structure, such as adding convolutional layers, changing the size of the convolution kernel, etc., to further improve the performance of the model;

[0130] In this embodiment, the model evaluation process is as follows: the accuracy (such as precision, recall, F1 score, etc.) and computational complexity (such as model size, inference time, etc.) of the model are calculated, and the evaluation results of each model, including accuracy and computational complexity, are recorded; and based on the evaluation results, the model accuracy and computational complexity are weighed, as follows:

[0131] Accuracy trade-off: Determine the threshold of accuracy requirements based on actual needs, and select models with accuracy exceeding the threshold as candidate models;

[0132] Computational complexity trade-off: Considering the resource constraints in actual application scenarios (such as computing resources, storage resources, etc.), among the candidate models, the model with lower computational complexity and meeting the accuracy requirements is selected as the final model;

[0133] Comprehensive evaluation: If there are multiple models with similar performance in terms of accuracy and computational complexity, other factors such as model stability and interpretability will be further considered to select the best candidate model as the final model;

[0134] S5. Deploy the final model to actual application scenarios, such as image classification systems, target detection systems, etc., to ensure that the model can correctly receive input images and output prediction results. Then, continuously monitor the performance of the model by collecting and evaluating data in actual application scenarios. Evaluation indicators include accuracy, response time, etc., and update and optimize according to actual needs, including adding new training data, adjusting model structure (such as the number of layers, number of neurons, etc.) or parameters (such as learning rate, batch size, etc.), introducing new adversarial defense mechanisms, etc.

[0135] like Figure 2 As shown, a high-precision image recognition system based on deep learning is applied to the high-precision image recognition algorithm based on deep learning as described above, including:

[0136] Image input module: responsible for obtaining image data from various devices (such as cameras, scanners, etc.) and passing it as raw input data to subsequent modules;

[0137] Data processing module: responsible for data labeling, preprocessing, and data enhancement of image data;

[0138] Alternative model generation module: responsible for selecting ResNet as the basic model, and integrating different adversarial sample defense mechanisms in turn to generate multiple alternative models;

[0139] Data partitioning module: responsible for dividing the processed random data into training data set and test data set;

[0140] Model training module: responsible for training the candidate model using the training data set;

[0141] Final model generation module: responsible for evaluating each candidate model using the test data set, calculating the accuracy and computational complexity of the model, and weighing the accuracy and computational complexity of the model based on the evaluation results, and selecting the best candidate model as the final model;

[0142] Optimization and adjustment module: responsible for performing optimization operations such as hyperparameter tuning and model structure adjustment on the final model to improve the performance of the model;

[0143] Model deployment module: responsible for deploying the optimized final model to actual application scenarios, such as intelligent monitoring, human-computer interaction, autonomous driving and other fields;

[0144] Monitoring and update module: responsible for continuously monitoring the performance of the model, including indicators such as accuracy and response time. If the model performance deteriorates, the model will be updated and optimized according to changes in actual needs to adapt to new application scenarios and data distribution.

[0145] The embodiments of the present invention are given for the purpose of illustration and description, and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.

Claims

1. A high-precision image recognition algorithm based on deep learning, characterized in that: Includes steps: S1. Select ResNet as the basic model and integrate different adversarial sample defense mechanisms in turn to generate multiple candidate models; S2, collect and annotate image data, preprocess and enhance the data; S3, dividing the processed random data into a training data set and a test data set, and using the training data set to train the candidate model; S4. Use the test data set to evaluate each candidate model, weigh the accuracy and computational complexity of the model, select the best candidate model as the final model, and optimize and adjust it; S5. Deploy the final model to the actual application scenario, continuously monitor the performance of the model, and update and optimize it according to actual needs.

2. According to the high-precision image recognition algorithm based on deep learning in claim 1, it is characterized in that: In step S1, each candidate model adopts a different defense mechanism or a combination of defense mechanisms, and the adversarial sample defense mechanism includes but is not limited to adversarial training, gradient masking, and adversarial sample detection; The adversarial training: by introducing adversarial samples during the training process, the model learns how to resist these adversarial attacks. The implementation method is: using an adversarial attack algorithm to generate adversarial samples, and sending the generated adversarial samples together with the original samples to the ResNet model for training, and then using the test data set to evaluate the performance of the model, including indicators such as accuracy and recall rate; The gradient masking: by modifying the activation function or other parameters of the model, it is difficult for the attacker to use the gradient information of the model to construct adversarial samples. The implementation method is: based on the ResNet model, the activation function or other parameters are modified to achieve gradient masking, and the modified model is used for training, and the performance of the model is evaluated using the test data set, including indicators such as the success rate of the adversarial attack; The adversarial sample detection: identifies whether an input sample is an adversarial sample by introducing a detection algorithm and taking corresponding defense measures. The implementation method is: a detector is trained using adversarial samples and normal samples to identify whether the input sample is an adversarial sample, and the trained detector is integrated into the ResNet model to detect the input sample. According to the output result of the detector, the input identified as an adversarial sample is rejected, warned or corrected.

3. The high-precision image recognition algorithm based on deep learning according to claim 2, characterized in that: The adversarial attack algorithm selects FGSM or PGD, wherein FGSM generates adversarial samples by adding a tiny perturbation with the same sign as the gradient to the original input sample, and PGD generates adversarial samples by applying the FGSM attack multiple times and projecting the adversarial samples back to the feasible set of the original data after each iteration.

4. The high-precision image recognition algorithm based on deep learning according to claim 1, characterized in that: The step S2 specifically includes the following steps: S21. Data collection: collected image data, and the collected image data must be able to cover various situations in actual application scenarios; S22, data annotation: assign a correct label to each image; S23, data preprocessing: including image normalization, resizing, and image denoising, as follows: Normalization: Use minimum-maximum normalization or Z-score normalization to scale the pixel values ​​of the image to a fixed range; Resizing: resize images to the same size by scaling, cropping, and padding; Image denoising: Use median filtering and Gaussian filtering to remove noise from images; S24, data enhancement: Generate new training samples by randomly transforming the original image, and the random transformation specifically includes: Rotation: Rotate the image at a random angle; Translation: Random translation of the image in the horizontal or vertical direction to simulate the situation where the object appears in different positions; Scaling: Randomly adjust the size or ratio of the image; Flip: flip the image horizontally or vertically; Cropping: Randomly select an area from the image for cropping; Color transformation: adjust the image's brightness, contrast, saturation and other color attributes to simulate different lighting conditions or shooting environments; Noise addition: Add random noise to the image to simulate the impact of different sensor noises on image quality and increase the robustness of the model.

5. The high-precision image recognition algorithm based on deep learning according to claim 4, characterized in that: In step S24, data enhancement also includes the following method: Random erasing: Randomly cover an area on the image so that the pixel value of the area becomes a constant or random value; Mixed image: linearly combine two images to form a new image; Target box transformation: In the object detection task, the position and size of the target box are updated accordingly according to the transformation of the image.

6. The high-precision image recognition algorithm based on deep learning according to claim 1, characterized in that: In step S3, the test data set and the training data set are independent of each other, and the data distribution is consistent, and the training data set accounts for 70% to 80% and the test data set accounts for 20% to 30%.

7. The high-precision image recognition algorithm based on deep learning according to claim 1, characterized in that: In step S3, the operation process of model training is as follows: Initialize the model: create a model instance and set hyperparameters; Loss function and optimizer: select a loss function suitable for the task, select an optimization algorithm, and set hyperparameters; Training process: includes outer loop and inner loop; The outer loop: traverses each round of training, and specifically performs the following operations: batch loading of training data through the data loader and resetting the indicators; The inner loop: traverses each batch of data, and the specific operations are: Forward propagation: pass the input data through the model and calculate the output; Calculate loss: Use the loss function to calculate the error between the model output and the target value; Back propagation: calculate gradients, clear gradients, and perform back propagation; Update parameters: Update model parameters through the optimizer; Record indicators: accumulated loss, update iteration number; Record logs: After each round, calculate the average loss and training time, and record log information.

8. The high-precision image recognition algorithm based on deep learning according to claim 1, characterized in that: In step S4, the model evaluation process: calculate the accuracy and computational complexity of the model, record the evaluation results of each model, including the accuracy and computational complexity; and weigh the model accuracy and computational complexity based on the evaluation results, as follows: Accuracy trade-off: Determine the threshold of accuracy requirements based on actual needs, and select models with accuracy exceeding the threshold as candidate models; Computational complexity trade-off: Considering the resource constraints in actual application scenarios, among the candidate models, the model with lower computational complexity and meeting the accuracy requirements is selected as the final model; Comprehensive evaluation: If there are multiple models with similar performance in terms of accuracy and computational complexity, other factors are further considered and the best candidate model is selected as the final model.

9. The high-precision image recognition algorithm based on deep learning according to claim 1, characterized in that: In step S5, continuous monitoring is achieved by collecting and evaluating data in actual application scenarios, and the evaluation indicators include accuracy and response time. The means of model updating and optimization include adding new training data, adjusting model structure or parameters, and introducing new adversarial defense mechanisms.

10. A high-precision image recognition system based on deep learning, applied to the high-precision image recognition algorithm based on deep learning as claimed in any one of claims 1 to 9, characterized in that: include: Image input module: responsible for obtaining image data and passing it as raw input data to subsequent modules; Data processing module: responsible for data labeling, preprocessing, and data enhancement of image data; Alternative model generation module: responsible for selecting ResNet as the basic model, and integrating different adversarial sample defense mechanisms in turn to generate multiple alternative models; Data partitioning module: responsible for dividing the processed random data into training data set and test data set; Model training module: responsible for training the candidate model using the training data set; Final model generation module: responsible for evaluating each candidate model using the test data set, calculating the accuracy and computational complexity of the model, and weighing the accuracy and computational complexity of the model based on the evaluation results, and selecting the best candidate model as the final model; Optimization and adjustment module: responsible for tuning the hyperparameters and adjusting the model structure of the final model; Model deployment module: responsible for deploying the optimized final model to the actual application scenario; Monitoring and update module: responsible for continuously monitoring the performance of the model. If the model performance deteriorates, the model will be updated and optimized according to changes in actual needs to adapt to new application scenarios and data distribution.