An adversarial attack method, system and storage medium in license plate text recognition

By obtaining the posterior probability matrix of the license plate image in the field of license plate recognition, determining and modifying the character confidence of the target sequence, and generating the attack target matrix, the problems of many iterations and poor attack effects in the existing technology are solved, and a more efficient adversarial attack is achieved.

CN114202678BActive Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111397376.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-06-13
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

The prior art has many iterations of adversarial attack methods in the field of license plate recognition, high calculation cost, poor attack effect, and difficult to reduce attack difficulty and improve attack rate.

Method used

By obtaining the license plate image to be attacked, input the CRNN model to obtain the posterior probability matrix, determine the target sequence where the character confidence needs to be modified, modify the confidence of the first and second characters under the target sequence, to generate the attack target matrix, judge the attack success and generate the attack sample image.

Benefits of technology

Reduces the difficulty of attacks, reduces the number of iterations, increases the rate of counterattacks, and obtains more realistic attack effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202678B_ABST
    Figure CN114202678B_ABST
Patent Text Reader

Abstract

The present invention provides an adversarial attack method, system and storage medium for license plate character recognition. The method includes: obtaining a license plate image to be attacked; inputting the license plate image to be attacked into a CRNN model to obtain an N*T posterior probability matrix corresponding to the license plate image to be attacked; determining a target sequence for which the character confidence needs to be modified; modifying the confidences corresponding to the first character and the second character in the target sequence to generate an attack target matrix; wherein, the first character is the character corresponding to the maximum confidence value in the target sequence, and the second character is the character corresponding to the second largest confidence value in the target sequence; determining whether the attack on the license plate image to be attacked is successful based on the posterior probability matrix and the attack target matrix; and generating an attack sample image according to the attack target matrix in the case of a successful attack. This adversarial attack method reduces the difficulty of attacking license plate images, reduces the number of iterations, and improves the adversarial attack rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image and text recognition, and in particular to a method, system and storage medium for countering attacks in license plate text recognition. Background Art

[0002] Text recognition is an image recognition system that recognizes text in an image based on input image information. In the field of text recognition, mainstream recognition methods include CRNN (convolutional and recurrent neural network), target detection and other methods. CRNN is a deep learning model that extracts image features and text sequence features by combining CNN neural network and RNN neural network. It first uses CNN network to extract image features of text, and then uses RNN to continue to extract text sequence features based on the convolution features extracted by CNN. CNN is a multi-layer neural network used to process two-dimensional data. It extracts features such as shape and texture in images through convolution kernels and pooling activation operations. Based on the features extracted by CNN network, RNN (recurrent neural network) is used to extract sequence information from image features extracted by the previous layer, and the generated fixed-length sequence information can be output.

[0003] Since neural networks are extremely vulnerable to adversarial examples, attackers can cause the neural network to learn inaccurate features and ultimately obtain incorrect prediction results by adding some small perturbations to the image or detection target. This process is called "adversarial attack", in which the adversarial example is a modified version of a clean image that is deliberately disturbed (for example, by adding noise) to fool machine learning techniques, such as deep neural networks. According to the attack effect, adversarial attacks can be divided into targeted attacks and untargeted attacks. Targeted attacks are successful when the model's prediction results are fooled into a specified target, while untargeted attacks are successful when the model's prediction results are fooled into errors. In the prior art, adversarial attacks on neural networks are generally carried out iteratively based on algorithms such as FGSM and PGD. In the application scenario of license plate recognition, since the license plate image has a small set of characters, the model is relatively robust and not very sensitive to attacks. Therefore, the existing adversarial attack methods in the field of license plate recognition have the following disadvantages: (1) The number of iterations is large. The CTCloss used in the training stage has high computational cost. It is necessary to map the intermediate variables output by the CRNN network with the attack target, which takes a long time to process. (2) The attack effect is poor, which is manifested in that too many pixels of the image need to be modified, and the generated adversarial results are too different from the expected ones, such as the number of digits in the license plate is reduced. Therefore, how to reduce the difficulty of adversarial attacks on license plate recognition, reduce the number of iterations, and increase the rate of adversarial attacks are technical problems that need to be solved urgently. Summary of the invention

[0004] In view of this, the present invention provides an adversarial attack method, system and storage medium for license plate character recognition to solve one or more problems existing in the prior art.

[0005] According to one aspect of the present invention, the present invention discloses an adversarial attack method for license plate character recognition, the method comprising:

[0006] Obtain a license plate image to be attacked;

[0007] Input the license plate image to be attacked into a CRNN model to obtain an N*T posterior probability matrix corresponding to the license plate to be attacked, where N is the length of the character set predicted by the RNN network and T is the number of sequences predicted by the RNN network;

[0008] Determine a target sequence for which the confidence of the character needs to be modified;

[0009] Modify the confidence of the first character and the second character corresponding to the target sequence to generate an attack target matrix; wherein, the first character is the character corresponding to the maximum confidence value under the target sequence, and the second character is the character corresponding to the second largest confidence value under the target sequence;

[0010] Based on the posterior probability matrix and the attack target matrix, determine whether the attack on the license plate to be attacked is successful;

[0011] In the case of a successful attack, generate an attack sample image according to the attack target matrix.

[0012] In some embodiments of the present invention, the method further comprises:

[0013] In the case of an unsuccessful attack, update the license plate image to be attacked received by the CRNN model using the gradient backpropagation algorithm.

[0014] In some embodiments of the present invention, determining a target sequence for which the confidence of the character needs to be modified includes:

[0015] Obtain each third character corresponding to the maximum confidence value under each sequence in the posterior probability matrix;

[0016] Determine a sequence in which the third character is different from the third characters in the adjacent two sequences as a candidate sequence;

[0017] Calculate the differences between the confidence values of the first character and the confidence values of the second character under each candidate sequence;

[0018] Take the sequence corresponding to the minimum difference among the differences as the target sequence.

[0019] In some embodiments of the present invention, modifying the confidence levels corresponding to the first character and the second character under the target sequence includes:

[0020] Modifying the confidence level value of the first character to 0;

[0021] Modifying the confidence level value of the second character to 1.

[0022] In some embodiments of the present invention, determining whether the attack on the license plate to be attacked is successful based on the posterior probability matrix and the attack target matrix includes:

[0023] Calculating the loss value of the posterior probability matrix and the attack target matrix based on the loss function;

[0024] Determining whether the loss value is less than a preset threshold;

[0025] In the case where the loss value is less than the preset threshold, the attack on the license plate to be attacked is successful;

[0026] In the case where the loss value is greater than the preset threshold, the attack on the license plate to be attacked fails.

[0027] In some embodiments of the present invention, the loss function is a cross-entropy loss function, a mean squared error loss function, or a mean absolute error loss function.

[0028] In some embodiments of the present invention, the calculation formula of the gradient backpropagation algorithm is:

[0029] x t+1 = x t + ε·sign(J(f θ (x), y))

[0030] where x is the license plate image to be attacked, f θ (x) is the posterior probability matrix output by the CRNN model, ε is the attack intensity, y is the attack target matrix, J represents the gradient of the posterior probability matrix and the attack target matrix obtained by differentiating according to the loss function, and x t represents the license plate image with interference factors generated in the t-th update iteration.

[0031] In some embodiments of the present invention, generating an attack sample image according to the attack target matrix includes:

[0032] Adding interference pixels to the current license plate image to be attacked received by the CRNN model in the current iteration process based on the calculation result of the gradient backpropagation algorithm;

[0033] Using the license plate image with added interference pixels as the attack sample image.

[0034] According to another aspect of the present invention, a countermeasure attack system in license plate character recognition is also disclosed. The system includes a processor and a memory, and computer instructions are stored in the memory. The processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described in any of the foregoing embodiments.

[0035] According to still another aspect of the present invention, a computer-readable storage medium is also disclosed, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any of the foregoing embodiments are implemented.

[0036] The countermeasure attack method and system in license plate character recognition obtain an attack target by modifying the confidence levels corresponding to the first character and the second character under the target sequence in the intermediate layer information of the CRNN model, and determine whether the attack on the license plate image to be attacked is successful based on the relationship between the intermediate layer information of the CRNN model and the attack target. This method reduces the attack difficulty, reduces the number of iterations, and improves the countermeasure attack rate. In addition, the countermeasure attack method calculates a loss function using the intermediate layer information of the CRNN model, backpropagates the gradient, and iteratively trains to modify the pixels of the input image to achieve the countermeasure attack target, which not only further effectively improves the countermeasure attack rate of license plate character recognition, but also can obtain a more realistic attack effect.

[0037] The additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned from the practice of the present invention. The objects and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings.

[0038] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. The components in the drawings are not drawn to scale, but are only for showing the principles of the present invention. For the convenience of showing and describing some parts of the present invention, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to the present invention. In the drawings:

[0040] Figure 1 is a schematic flowchart of a countermeasure attack method in license plate character recognition according to an embodiment of the present invention.

[0041] Figure 2 Schematic diagram of the decoding process of the CRNN model in the recognition stage according to an embodiment of the present invention.

[0042] Figure 3 Schematic diagram of the process for generating the attack target matrix according to an embodiment of the present invention.

[0043] Figure 4 Schematic diagram of the process for the adversarial attack method in license plate character recognition according to another embodiment of the present invention. Detailed implementation manners

[0044] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0045] Herein, it should be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0046] It should be emphasized that the terms "include / comprise / have" when used herein refer to the presence of features, elements, steps or components, but do not exclude the presence or addition of one or more other features, elements, steps or components.

[0047] In the training stage of the neural network, the CTC algorithm is generally used as the loss function to calculate the loss, and the gradient is backpropagated to train the model; the advantage of using the CTC algorithm is that since the output of the RNN network is a fixed-length sequence, in the field of text recognition, the strings to be recognized are variable-length, so the CTC algorithm is used to map the variable-length target sequence and the fixed sequence output by the RNN layer during the training process, calculate the model loss, and update the model parameters for training. Further, after the model training is completed, a license plate image is input into the model, and the model outputs a text sequence. According to the merging rule, after merging the repeated characters, the decoded string is output to obtain the final result. When using the model to recognize the license plate image, in order to perform an adversarial attack on the license plate image, generally, an adversarial sample needs to be obtained through a certain method. The process of obtaining the adversarial sample not only requires minimizing the number of iterations in the training process as much as possible, but also ensuring a high adversarial attack rate. Therefore, the present invention constructs a special adversarial attack method to guide the adversarial attack process in the field of license plate recognition.

[0048] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0049] Figure 1The flowchart of the adversarial attack method in license plate character recognition according to an embodiment of the present invention is shown as Figure 1 shown. The adversarial attack method includes the following steps S10 to S60.

[0050] Step S10: Obtain the license plate image to be attacked.

[0051] The license plate image has a text sequence of a fixed length, such as: Jing X AABBC.

[0052] Step S20: Input the license plate image to be attacked into the CRNN model to obtain the N*T posterior probability matrix corresponding to the license plate to be attacked, where N is the length of the character set predicted by the RNN network, and T is the number of sequences predicted by the RNN network.

[0053] The CRNN model is used to recognize the characters in the license plate image. First, it extracts the image features of the license plate image to be attacked received by the CNN neural network. The RNN, as a recurrent neural network, further extracts sequence information from the image features extracted by the CNN and predicts the probability distribution of each character in the character set, thereby generating a sequence information of a fixed length. For example, when the feature map of the license plate image is extracted based on the CNN network, further extract the feature vector sequence required by the RNN network according to the feature map. Each vector in the feature vector sequence is regarded as an associated rectangular area (receptive field) in the license plate image. Then, each feature vector in the feature vector sequence is used as the input of the RNN network at a time step. Therefore, the number of feature vectors in the feature vector sequence can also be regarded as the number of sequences predicted by the RNN network. Assume that the number of feature vectors in the feature vector sequence of the license plate image extracted by the CNN network is 15, then T in the posterior probability matrix is 15. And for the length of the character set, it is the total number of all characters predicted by the RNN network. For example, the characters in the character set are ε (separator), Jing, Jin, A, B, C... X, Y, Z, then the value corresponding to N at this time is the total number of the aforementioned all characters.

[0054] Since the CRNN network model is generally a combination of CNN+RNN+TCT, but the inventor found in the experiment that due to the small number of characters in the license plate recognition character set and the strong robustness of the model, the model is not easily attacked. If the CTCLoss is directly used as the adversarial attack objective function, it is easy to cause unexpected results in the attack, such as the phenomenon that the number of digits in the license plate number decreases. Therefore, the inventor in the present invention is based on the output of the intermediate layer RNN network as the basic data for generating the attack target in the subsequent steps, that is, further calculating the loss based on the output vector crnnOutput[T][N] to obtain the expected attack sample.

[0055] Step S30: Determine the target sequence for modifying the character confidence.

[0056] In this step, further determine the sequence that needs to have its confidence value modified in the posterior probability matrix output by the CRNN network model as the target sequence. Exemplarily, if the number of sequences in the posterior probability matrix at the output end of the CRNN is 11, since each column of sequences represents a time step, the time step is also 11 at this time. The 11 sequences are 1, 2... 11 from left to right. Assuming that the sequence for which the character confidence needs to be modified is 8, it can be considered that changing the pixels in the receptive field on the license plate image corresponding to sequence 8 is likely to be successfully attacked.

[0057] In an embodiment of the present invention, determining the target sequence for which the character confidence needs to be modified specifically includes: obtaining each third character corresponding to the maximum confidence value under each sequence in the posterior probability matrix; determining the sequence where the third character is different from the third characters under the two adjacent sequences as the candidate sequence; calculating the differences between the confidence values of the first character and the confidence values of the second character under each candidate sequence; and taking the sequence corresponding to the smallest difference among the differences as the target sequence.

[0058] In this embodiment, the third character refers to the character corresponding to the maximum confidence under each sequence. For example, the character corresponding to the maximum confidence under sequence 1 may be A, and the character corresponding to the maximum confidence under sequence 2 may be B. At this time, character A is the third character under sequence 1, and character B is the third character under sequence 2; from this, it can also be known that each sequence has a third character, and there may be a situation where the third characters in two or more sequences in the posterior probability matrix are the same character. Further, compare and judge the third characters under two adjacent sequences, that is, judge whether the third character under a certain sequence is the same as the third characters under the two sequences adjacent to the left and right of this sequence. If it is judged that the third character under a certain sequence is different from the third characters under the two sequences adjacent to the left and right of this sequence, then judge that this sequence is the candidate sequence. Exemplarily, assume that the third character corresponding to sequence 4 is D; if the third character of sequence 3 is A and the third character of sequence 5 is E, then at this time, the third characters of sequence 3, sequence 4, and sequence 3 are all different, so sequence 4 is a candidate sequence; and if the third character of sequence 3 is A and the third character of sequence 5 is D, then at this time, the third characters of sequence 4 and sequence 5 are the same, so neither sequence 4 nor sequence 5 can be used as a candidate sequence.

[0059] Generally, multiple candidate sequences can usually be found from the posterior probability matrix corresponding to each license plate image to be attacked. Further, a target sequence needs to be found from the multiple candidate sequences. For example, the found posterior sequences are: sequence 6, sequence 7, and sequence 9. Finally, sequence 9 is found from the above three sequences as the target sequence. Specifically, when selecting the target sequence, the sequence corresponding to the minimum difference between the confidence value of the first character and the confidence value of the second character can be found in the candidate sequences as the target sequence. Among them, the first character is the character corresponding to the maximum confidence in the sequence, and the second character is the character corresponding to the second largest confidence in the sequence. Listing the sequence with the minimum difference between the confidence value of the first character and the confidence value of the second character as the target sequence is because the confidence represents the prediction probability of the character. In the target sequence, if the confidence difference between the first character and the second character is small, it means that the probability that the RNN misjudges the second character in this sequence as the first character is greater. Then, using this target sequence as the attack target will increase the success rate of the license plate image attack.

[0060] Step S40: Modify the confidence levels corresponding to the first character and the second character in the target sequence to generate an attack target matrix; where the first character is the character corresponding to the maximum confidence value in the target sequence, and the second character is the character corresponding to the second largest confidence value in the target sequence.

[0061] After determining the target sequence in the above steps, further modify the confidence levels of the first character and the second character in this target sequence. Similar to the above steps, in this sequence, the probability value corresponding to the first character is the highest, and the probability value corresponding to the second character is only less than the probability value of the first character. At this time, the matrix with the confidence levels of the first character and the second character in the target sequence modified is used as the attack target matrix. At this time, the attack target matrix and the posterior probability matrix output by the middle layer of the CRNN model only differ in the confidence values of the first character and the second character in the target sequence, while the confidence values corresponding to the characters in the remaining sequences are the same.

[0062] Specifically, modifying the confidence levels corresponding to the first character and the second character under the target sequence includes: modifying the confidence level value of the first character to 0; modifying the confidence level value of the second character to 1. Since the probability of the second character under the target sequence being predicted in the posterior probability matrix is second only to the prediction probability of the first character, the CRNN model is prone to misidentifying the second character under the target sequence as the first character. Therefore, modifying the confidence level of the second character to 1 in the attack target matrix means that it is most desired that the model recognize the character under this sequence as the second character during the license plate attack process, and the possibility of recognizing it as the first character is set to 0. It should be understood that the setting method of modifying the confidence level value of the first character to 0 and the confidence level value of the second character to 1 is only a preferred example. It can also modify the confidence level values of the first character and the second character to other values, such as modifying the confidence level of the first character to 0.1 and the confidence level of the second character to 0.9, etc., as long as it is ensured that the confidence level of the second character under the modified target sequence is greater than the confidence level of the first character.

[0063] Step S50: Determine whether the attack on the license plate to be attacked is successful based on the posterior probability matrix and the attack target matrix.

[0064] In this step, the posterior probability matrix is the intermediate layer data output by the CRNN based on the input license plate image to be attacked, and the attack target matrix is the matrix obtained after modifying the confidence levels of the first character and the second character under the target sequence in the posterior probability matrix. In this step, it is further determined whether the attack on the license plate to be attacked is successful. In the case of a successful attack, the attack ends; while if it is determined that the attack fails, the license plate image received by the CRNN model can be updated using the gradient backpropagation algorithm. Then, the CRNN model can output the posterior probability matrix corresponding to the updated license plate image based on the updated license plate image. At this time, since there are slight differences in the license plate images received by the CRNN during the two iterative processes, the two posterior probability matrices obtained during the two iterative processes are also different. If the first iterative process is regarded as the t-th iteration, then the subsequent iteration is the (t + 1)-th iteration. Then, the updated license plate image is further used as the input of the CRNN, and the above steps S20 - S40 are executed in a loop, and the attack target matrix of the (t + 1)-th iteration can be obtained.

[0065] In one embodiment, determining whether the attack on the license plate to be attacked is successful based on the posterior probability matrix and the attack target matrix includes: calculating the loss values of the posterior probability matrix and the attack target matrix based on a loss function; determining whether the loss value is less than a preset threshold; if the loss value is less than the preset threshold, the attack on the license plate to be attacked is successful; if the loss value is greater than the preset threshold, the attack on the license plate to be attacked fails.

[0066] In the above embodiment, the loss function can specifically be a cross-entropy loss function, a mean squared error loss function, an average absolute error loss function, etc.; and the preset threshold can be set according to actual requirements.

[0067] Step S60: If the attack is successful, generate an attack sample image according to the attack target matrix.

[0068] In this step, the posterior probability matrix corresponding to the license plate image to be attacked received by the CRNN model is used for comparison with the attack target matrix. Therefore, if the attack target matrix obtained after the first modification of the posterior probability matrix meets the attack requirements, the license plate image to be attacked can be directly pixel-modified based on this attack target matrix, and the modified license plate image to be attacked is the attack sample image. If the attack target matrix after multiple iterative updates meets the attack requirements, then the attack sample image is the license plate image generated after pixel-modifying the original license plate image to be attacked multiple times.

[0069] Specifically, generating an attack sample image according to the attack target matrix includes: adding interference pixels to the current license plate image to be attacked received by the CRNN model in the current iteration process based on the calculation result of the gradient backpropagation algorithm; using the license plate image with interference pixels added as the attack sample image.

[0070] Among them, the calculation formula of the gradient backpropagation algorithm is:

[0071] x t+1 = x t + ε·sign(J(f θ (x), y));

[0072] Among them, x is the license plate image to be attacked, f θ (x) is the posterior probability matrix output by the CRNN model, ε is the attack intensity, y is the attack target matrix, J represents the gradient of the posterior probability matrix and the attack target matrix obtained by taking the derivative according to the loss function, and x t represents the license plate image with interference factors generated in the t-th update iteration.

[0073] To further describe the advantages of the present invention, the adversarial attack method and system in license plate character recognition disclosed in the present application will be described below through a specific embodiment.

[0074] Generally, if we want to make the attack on the license plate character recognition CRNN+CTC loss model more efficient and minimize the degree of image modification, we need to design an adversarial attack loss function according to the decoding rules of the CTC loss. Among them, inputImage is defined as the license plate image to be predicted input to the CRNN model; label is defined as the text characters on the license plate image; target is defined as the attack target, that is, the text characters that are expected to be recognized by the CRNN model after being misled; crnnOutput[T][N] is defined as the output vector of the CRNN network, where T represents the predicted sequence length of the RNN network and N represents the character set length, and crnnOutput[x][y] represents the confidence that the character at position x in the character sequence may be the character y in the character set; crnnTarget is defined as the attack target designed for the output vector of the CRNN network; ctcOutput[T] is defined as the character vector decoded by the ctc loss algorithm, and ctcOutput[x] represents the character at position x in the final predicted character sequence; crossEntropyLoss is defined as the cross-entropy loss function.

[0075] Specifically, Figure 2 is a schematic diagram of the decoding process of the CRNN model in the recognition stage according to an embodiment of the present invention, as Figure 2 shown. First, the output of the CRNN network is crnnOutput[T][N]; through this vector, we select the character with the highest confidence at each position, extract a fixed-length sequence, and decode it into the final sequence according to the following rules:

[0076] Rule 1: All consecutive and identical characters in the sequence are decoded and merged into one such character.

[0077] Rule 2: ε is decoded as a separator and not output, but two identical characters separated by ε are not considered consecutive and will not be merged by Rule 1.

[0078] According to the above two rules, the sequence "JingJingXXAεABBCD" is decoded as "JingXAABCD".

[0079] To meet the requirements of speed and effectiveness in adversarial attacks, it is necessary to precisely set the attack targets. For example, the attack target for the license plate picture of "Jing XAABCD" is set to "Jing XAAB0D". This setting can effectively improve the attack speed and effectiveness. However, for large-scale batch attacks, it is impossible to manually select weaknesses for each picture, and the attack weaknesses understood manually are also different from the text features extracted by the neural network. Therefore, the objective function can be constructed through the following methods:

[0080] (1) During adversarial attacks, instead of using CTCloss as the objective function, directly use the output of the intermediate layer of CRNN and calculate the cross-entropy loss with the attack target.

[0081] (2) The CRNN network model outputs a vector of size N*T, where T is the output sequence length. According to the decoding rules described above, the attack target of the objective function only changes the vector value at one position. The character represented by the position with the highest confidence at this position needs to be different from the character represented by the position with the highest confidence at the adjacent position, that is, modifying this character does not cause the possibility of an increase in the number of digits of the finally decoded character. At the same time, among the positions that meet all the above conditions, find the position where the difference between the highest confidence and the second highest confidence is the smallest, and modify the confidence values of the two positions, that is, set the confidence value of the second highest confidence position to 1 and the confidence value of the highest confidence position to 0 as the attack target.

[0082] The calculation process of its attack target can be implemented through the following algorithm:

[0083]

[0084]

[0085] Exemplarily, the posterior probability matrix crnnOutput of the original CRNN output is shown in Table 1, and the modified attack target matrix is shown in Table 2.

[0086] Table 1. Posterior probability matrix of the original CRNN output

[0087] 1 2 3 4 5 6 7 8 9 10 11 ε … … … … … 0.90 … … … … … Beijing 0.91 0.95 … … … … … … … … … Tianjin … … … … … … … … … … … … A … … … … 0.88 … 0.90 … … … … B … … … … 0.10 … … 0.66 0.70 … … C … … … … … … 0.50 … … 0.55 … D … … … … … … … 0.30 … … 0.6 E … … … … … … … … … … … … … … … … … … … … … X … … 0.90 0.91 … … … … … … … … 0 … … … … … … … … … 0.45 0.31 …

[0088] In Table 1 above, based on the attack target selection rules, the selected candidate sequences are sequences 5, 7, 10, and 11 respectively. The characters represented by the positions with the highest confidence in the above candidate sequences are different from the characters represented by the positions with the highest confidence at the adjacent positions.

[0089] Furthermore, select the sequence with the smallest difference between the highest confidence and the second highest confidence among the candidate sequences 5, 7, 10, and 11. It is not difficult to find that among the confidence vectors of sequence 10, the difference between "C" and "0" is the smallest. Therefore, select this position as the target for modification. The modified attack target matrix is shown in Table 2.

[0090] Table 2. Modified Attack Target Matrix

[0091] 1 2 3 4 5 6 7 8 9 10 11 ε … … … … … 0.90 … … … … … Beijing 0.91 0.95 … … … … … … … … … Tianjin … … … … … … … … … … … … A … … … … 0.88 … 0.90 … … … … B … … … … 0.10 … … 0.66 0.70 … … C … … … … … … 0.50 … … 0.0 … D … … … … … … … 0.30 … … 0.6 E … … … … … … … … … … … … … … … … … … … … … X … … 0.90 0.91 … … … … … … … … 0 … … … … … … … … … 1 0.31 …

[0092] Specifically, Figure 3 is a schematic flowchart of the attack target matrix generation process according to an embodiment of the present invention. Refer to Figure 3 , the CRNN network inputs the license plate image to be attacked inputImage, and the CRNN network outputs crnnOutput. Subsequently, according to the attack target design scheme of the present invention, crnnTarget is obtained, that is, the selection of the first-stage attack target is completed. Further, based on the above-mentioned generated attack target, the cross-entropy loss function is used to train the interference pixels. Among them Figure 4 is a schematic flowchart of the adversarial attack method in license plate character recognition according to another embodiment of the invention. As Figure 4 shown, during the adversarial attack training process, (1) First, input the license plate image to be attacked inputImage into the CRNN network, and the network outputs the crnnOutput matrix;

[0093] (2) Use the cross-entropy loss function to calculate the loss value between crnnOutput and crnnTarget; (3) Determine whether the loss value is less than the preset threshold. If it is less, the attack is successful and the loop ends; if it is greater, the loop continues to execute, and the gradient is calculated based on the gradient backpropagation algorithm; (4) Modify inputImage based on the gradient calculation result, and update the license plate image to be attacked input into the CRNN network model, and then loop through steps (1) to (4). Finally, obtain the adversarial attack image with interference pixels added as the adversarial attack sample image.

[0094] Through the above embodiments, it can be found that the adversarial attack method in license plate character recognition reduces the amount of calculation, and by selecting the attack weak point position of the CRNN model as the attack target position, reduces the attack difficulty, reduces the number of iterations, and ensures the attack effect.

[0095] Correspondingly, the present invention also discloses an adversarial attack system in license plate character recognition. The system includes a processor and a memory, and is characterized in that the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described in any one of the above embodiments.

[0096] In addition, the present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0097] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0098] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0099] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0100] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adversarial attack method in license plate text recognition, characterized in that, the method includes: Obtain the license plate image to be attacked; Input the license plate image to be attacked into the CRNN model to obtain the N*T posterior probability matrix corresponding to the license plate image to be attacked, where N is the length of the character set predicted by the RNN network and T is the number of sequences predicted by the RNN network; Determine the target sequence for which the character confidence needs to be modified; Modify the confidences corresponding to the first character and the second character in the target sequence to generate an attack target matrix; wherein, the first character is the character corresponding to the maximum confidence value in the target sequence, and the second character is the character corresponding to the second largest confidence value in the target sequence; Based on the posterior probability matrix and the attack target matrix, determine whether the attack on the license plate to be attacked is successful; In the case of a successful attack, generate an attack sample image according to the attack target matrix; Determining the target sequence for which the character confidence needs to be modified includes: Obtain each third character corresponding to the maximum confidence value in each sequence of the posterior probability matrix; Determine the sequence where the third character is different from the third characters in the adjacent two sequences as the candidate sequence; Calculate the differences between the confidence values of the first character and the confidence values of the second character in each candidate sequence; Take the sequence corresponding to the minimum difference among the differences as the target sequence.

2. The adversarial attack method in license plate text recognition according to claim 1, characterized in that, the method further includes: In the case of an attack failure, use the gradient backpropagation algorithm to update the license plate image to be attacked received by the CRNN model.

3. The adversarial attack method in license plate text recognition according to claim 1, characterized in that, Modifying the confidences corresponding to the first character and the second character in the target sequence includes: Modify the confidence value of the first character to 0; Modify the confidence value of the second character to 1.

4. The adversarial attack method in license plate text recognition according to claim 1, characterized in that, Based on the posterior probability matrix and the attack target matrix, determining whether the attack on the license plate to be attacked is successful includes: Calculate the loss value of the posterior probability matrix and the attack target matrix based on the loss function; Determine whether the loss value is less than a preset threshold; In the case where the loss value is less than the preset threshold, the attack on the license plate to be attacked is successful; In the case where the loss value is greater than the preset threshold, the attack on the license plate to be attacked fails.

5. The adversarial attack method in license plate text recognition according to claim 4, characterized in that, The loss function is a cross-entropy loss function, a mean square error loss function or an average absolute error loss function.

6. The adversarial attack method in license plate text recognition according to claim 2, characterized in that, The calculation formula of the gradient backpropagation algorithm is: x t+1 = x t + ε·sign(J(f θ (x), y)); Among them, x is the license plate image to be attacked, and f θ (x) is the posterior probability matrix output by the CRNN model, ε is the attack intensity, y is the attack target matrix, J represents the gradient of the posterior probability matrix and the attack target matrix obtained by differentiating according to the loss function, and x t represents the license plate image with interference factors generated in the t-th update iteration.

7. The adversarial attack method in license plate text recognition according to claim 2, characterized in that, Generating an attack sample image according to the attack target matrix includes: Add interference pixels to the current license plate image to be attacked received by the CRNN model during the current iteration based on the calculation result of the gradient backpropagation algorithm; Use the license plate image with interference pixels added as the attack sample image.

8. An adversarial attack system in license plate character recognition, the system includes a processor and a memory, characterized in that, computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described in any one of claims 1 to 7.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • License plate attack generation method based on adversarial attack

    CN108446700A

  • Method for improving verifiable defensive performance of model under maximum stochastic smoothing

    CN112766336A