An indefinite-length general digital and alphabetic verification code recognition system and method
By generating verification code pictures of various lengths, fonts and noises, and building a deep convolutional neural network model for processing, the problem of poor recognition effect of the existing verification code recognition system is solved, and efficient and accurate verification code recognition is achieved.
Patent Information
- Application Number
- CN202111606854.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The existing digital letter verification code recognition system is limited to fixed length, fixed font and fixed noise mode during recognition, resulting in poor recognition effect and inability to effectively process characters with high edge bonding, resulting in slow recognition rate and low accuracy.
By using the Captcha third-party library to generate digital letter verification code pictures of various lengths, fonts and noises, and construct a deep convolutional neural network model, convolution, pooling and overfitting the verification code pictures, split edge-bonded characters, and process indefinite-length verification codes through one-hot encoding.
It realizes efficient identification of various types of verification codes, improves the recognition rate and accuracy, and ensures the accuracy and consistency of verification code recognition.
Smart Images

Figure CN114329414B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a variable-length general digital and alphabetic verification code recognition system and method. Background Technique
[0002] A verification code is a public fully automated program that differentiates between users who are computers and those who are humans, and can prevent malicious password cracking, ticket stuffing, and forum spamming. It can effectively prevent a hacker from continuously attempting to log in to a specific registered user using a specific program in a brute-force cracking manner. Therefore, accurate recognition of the verification code can avoid situations where users are unable to use the program normally due to unclear recognition in daily life.
[0003] When existing digital and alphabetic verification code recognition systems recognize verification codes, they are all solutions for fixed lengths, fixed fonts, and fixed noise patterns, resulting in a narrow applicable range of the recognition system, poor usage effects, and when constructing a convolutional neural network model to recognize verification codes, unable to split characters with high edge adhesion, resulting in a slow model processing rate, and when training the constructed convolutional neural network, unable to ensure that the verification code accuracy reaches the maximum value, further reducing the recognition accuracy of the recognition system. Summary of the Invention
[0004] The purpose of the present invention is to provide a variable-length general digital and alphabetic verification code recognition system and method to solve the problems raised in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solution: A variable-length general digital and alphabetic verification code recognition method, the method comprising the following steps:
[0006] Step 1: Randomly generate digital and alphabetic verification code pictures of various lengths, fonts, and noises by using the captcha third-party library;
[0007] Step 2: Construct a deep convolutional neural network model based on the digital and alphabetic verification code pictures generated in Step 1;
[0008] Step 3: Train the deep convolutional neural network model constructed in Step 2 until a training model with higher accuracy is obtained;
[0009] Step 4: Use the deep convolutional neural network model trained in Step 3 to recognize the uploaded verification code pictures.
[0010] Further, the specific method for randomly generating digital and alphabetic verification code pictures of various lengths, fonts, and noises by using the captcha third-party library in Step 1 is:
[0011] Step 1 (I). Traverse all the fonts in the captcha third-party library, randomly select font sizes and string types, and generate multiple single-character images;
[0012] Step 2 (II). Randomly rotate, scale, and distort the multiple single-character images, and randomly select from the processed single-character images to combine and generate verification code images of variable lengths;
[0013] Step 3 (III). Randomly add different types of noise to the verification code images generated in Step 2 (II) to generate digital and alphabetic verification code images of various lengths, fonts, and noises.
[0014] Furthermore, the specific method for constructing the deep convolutional neural network model in Step 2 is as follows:
[0015] Step 2 (1). Select the digital and alphabetic verification code images generated in Step 1 as the input, perform convolution processing on the input verification code images, perform pooling and overfitting processing on the convolution-processed verification code images, and output the pooling- and overfitting-processed verification code images. Performing convolution processing on the verification code images is used to detect the character edges in the verification code images, which is beneficial for feature extraction. Performing pooling processing on the convolution-processed verification code images is used to reduce the redundant information in the convolution-processed images, thereby reducing the output value of the verification code images and improving work efficiency. Performing overfitting processing on the pooling-processed verification code images is used to avoid stagnation or decline when the model accuracy reaches the highest value;
[0016] Step 2 (2). Repeat the convolution, pooling, and overfitting processing on the verification code images output in Step 2 (1) until the verification code string can be recognized while minimizing the output value of the verification code images;
[0017] Step 2 (3). Split the strings in the verification code images output in Step 2 (2) into single-character images according to the distribution of pixel values after the last convolution processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, then the multi-character is split as a single-character image, and the single-character images are numbered according to their specific positions in the verification code images. Splitting the multi-character with glued edge positions as a single-character image can avoid the final output verification code being different from the actual verification code;
[0018] Step 2 (4). If there is a situation where a multi-character with glued edge positions is split as a single-character image, then calculate the loss value of the single-character image, and re-split the single-character image according to the calculation result of the loss value, and number the multi-characters according to their specific positions in the single-character image. The specific calculation formula for the loss value L is:
[0019]
[0020] Among them, m represents three character types of 0-9, a-z, and A-Z. j = 1, 2, 3. When j = 1, it represents the character type of 0-9. When j = 2, it represents the character type of a-z. When j = 3, it represents the character type of A-Z. y j represents the output value of the j-th character type. When j = 1, y 1 represents the output value of the 1st character type. At this time, y 1 = -1. When j = 2, y 2 represents the output value of the 2nd character type. At this time, y 2 = 0. When j = 3, y 3 represents the output value of the 3rd character type. At this time, y 3 = 1, p j represents the probability that the judged character is of the j-th character type. y j *p j represents the probability that the character is of the j-th character type. When it means that the probability of the character type of A-Z is the largest. At this time, the edge adhesion is split into the character pictures connected to it to the greatest extent. When it means that the probability of the character type of 0-9 is the largest. At this time, the edge adhesion is split into the character pictures connected to it to the greatest extent. When it means that the probability of the character type of a-z is the largest. At this time, the edge adhesion is split into the character pictures connected to it to the greatest extent;
[0021] Step 2(5). Input the numbered single-character pictures into the letter-channel convolutional neural network and the number-channel convolutional neural network according to the character types respectively. Perform convolution, pooling, and overfitting processing on the single-character pictures input into the letter-channel convolutional neural network and the number-channel convolutional neural network again until the characters on the single-character pictures can be read, and generate the verification code string by combining the output single characters according to the number order.
[0022] Furthermore, the specific method for training the deep convolutional neural network model constructed in Step 2 until a training model with a higher accuracy is obtained is as follows:
[0023] Step 3(1). Cut the verification code pictures into a training set and a validation set according to a ratio. Then, reduce the too-large pictures in the training set and the validation set to the standard specification size, and enlarge the too-small pictures according to the height and center them at the center position of the standard specification white background;
[0024] Step 3(2). Perform pixel grayscale processing on the segmented verification code images in Step 3(1), count the sizes of the color blocks in the processed images, fill the easily distinguishable small color blocks with the background color after counting. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating-point number between 0 and 1. Then, input the processed verification code image into the deep convolutional neural network model for recognition;
[0025] Step 3(3). Based on the one-hot algorithm, encode the recognition result into a 62-bit one-hot encoding. To process variable-length verification codes, use the 62+1 method to finally encode the recognition result into a 63-bit one-hot encoding;
[0026] Step 3(4). Calculate the accuracy of the variable-length verification code obtained after encoding. If the calculated accuracy changes significantly after multiple iterations, halve the initial learning rate set by the model and calculate again. If the calculated accuracy changes slightly after multiple iterations, stop the calculation to obtain the trained model. The specific accuracy calculation formula Q is as follows:
[0027]
[0028] where n represents the number of characters in the verification code, i = 1, 2, 3, 4, 5, 6, 7, 8 represents the i-th character in the verification code, X represents the 63-bit one-hot encoding, x represents the number of specific digits or letters in each character, represents the sum of the ratios of specific digits or letters in n characters to the 63-bit one-hot encoding. When there are no empty characters in the verification code, x = 8. When Q < 98.4%, it means that the result output by the model at this time does not match the actual situation. When Q = 98.4%, it means that the result output by the model at this time matches the actual situation. When Q > 98.4%, it means that there are empty characters in the output verification code. When the model accuracy is 98.4%, the model training ends.
[0029] Furthermore, the specific method for using the deep convolutional neural network model trained in Step 3 to recognize the uploaded verification code image in Step 4 is as follows:
[0030] Step 4(1). Open the trained model in the web service, and the UI automation test scenario uploads the verification code image file through the web interface;
[0031] Step 4(2). Open the verification code image file in the training model. In the training model, adjust the verification code image to the standard specification size and make the verification code image located in the middle of the background. Then, perform pixel grayscale processing on the verification code image, count the size of the color blocks in the processed image, fill the distinguishable small color blocks after statistics with the background color. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, input the processed verification code image into the training model. Based on the one-hot algorithm, encode the output result of the training model into a 63-bit one-hot encoding;
[0032] Step 4(3). Perform argmax calculation on the 63-bit one-hot encoding, convert the encoding into a recognizable string, and return the recognition result to the UI automation test scenario through the web interface.
[0033] An indefinite-length general digital and alphabetic verification code recognition system, including a verification code picture generation module, a deep convolutional neural network model construction module, a training module, and an identification module;
[0034] The verification code picture generation module is used to generate a large number of digital and alphabetic verification code pictures with various lengths, fonts, and noises required for training the deep convolutional neural network training model through the captcha third-party library, and transmit the generated digital and alphabetic verification code pictures to the deep convolutional neural network model construction module and the training module;
[0035] The deep convolutional neural network model construction module is used to receive the digital and alphabetic verification code pictures transmitted by the verification code picture generation module. The deep convolutional neural network model construction module uses tenoflow and kera to construct a deep convolutional neural network model, and transmits the constructed deep convolutional neural network model to the training module;
[0036] The training module is used to receive the deep convolutional neural network model constructed by the deep convolutional neural network model and the verification code pictures transmitted by the verification code picture generation module, train the deep convolutional neural network model based on the verification code pictures, and transmit the trained network model to the identification module;
[0037] The identification module is used to receive the training model transmitted by the training module. The training model encodes and recognizes the verification code pictures uploaded through the web interface, and then converts the encoding into a string through argmax calculation to obtain the recognition result.
[0038] Further, the verification code picture generation module includes a single-character picture generation unit, a verification code picture generation unit, and a picture noise generation unit;
[0039] The single-character picture generation unit traverses the system fonts, generates multiple single-character pictures with random font sizes and string types, and transmits the generated multiple single-character pictures to the verification code picture generation unit;
[0040] The verification code picture generation unit receives the multiple single-character pictures transmitted by the single-character picture generation unit, randomly rotates, scales, and distorts the received pictures, randomly combines the processed multiple single-character pictures to generate a verification code picture, and transmits the generated verification code picture to the picture noise generation unit;
[0041] The picture noise generation unit receives the verification code picture transmitted by the verification code picture generation unit, randomly adds noise to the received verification code picture, generates digital and letter verification code pictures with various lengths, fonts, and noises, saves the generated digital and letter verification code pictures to the disk, and transmits them to the deep convolutional neural network model.
[0042] Further, the deep convolutional neural network model construction module includes an input unit, a processing unit, and a deep convolutional neural network model construction unit;
[0043] The input unit receives the digital and letter verification code pictures transmitted by the picture noise generation unit and inputs the received digital and letter verification code pictures into the convolutional neural network;
[0044] The processing unit performs convolution, pooling, and overfitting processing on the digital and letter verification code pictures input into the convolutional neural network, repeats the convolution, pooling, and overfitting processing on the digital and letter verification code pictures until the output value of the verification code picture is minimized while being able to recognize the verification code string. Then, the processed verification code picture is split into single-character pictures according to the distribution of pixel values after the last convolution processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, the multi-character is regarded as a single-character picture for splitting, the single-character pictures are numbered according to their specific positions in the verification code picture, then the loss value of the single-character pictures containing multi-characters is calculated, the single-character pictures containing multi-characters are split according to the calculation results, and the multi-characters are numbered according to their specific positions in the single-character pictures. The numbered single-character pictures are respectively input into the letter-channel convolutional neural network and the digit-channel convolutional neural network according to the character type, and the single-character pictures are subjected to convolution, pooling, and overfitting processing again until the output value of the single-character picture is minimized while being able to recognize the character, and the output single-characters are combined into a verification code string according to the numbering order, and the entire process of obtaining the verification code string is transmitted to the deep convolutional neural network model construction unit;
[0045] The deep convolutional neural network model construction unit receives the entire process of obtaining the verification code string transmitted by the processing unit, constructs a deep convolutional neural network model according to the received content, and transmits the constructed deep convolutional neural network model to the training module.
[0046] Further, the training module includes a verification code image preprocessing unit, a verification code image denoising processing unit, a verification code label processing unit, and an accuracy calculation unit;
[0047] The verification code image preprocessing unit receives the verification code images transmitted by the verification code image generation module, divides the verification code images into a training set and a validation set according to a ratio, then reduces the too-large images in the training set and the validation set to the standard specification size, enlarges the too-small images according to the height and centers them in the middle of the standard specification white background, and transmits the preprocessed verification code images to the verification code image denoising processing unit;
[0048] The verification code image denoising processing unit receives the verification code images transmitted by the verification code image preprocessing unit, performs pixel grayscale processing on the received verification code images, counts the sizes of the color blocks in the processed images, fills the easily distinguishable small color blocks with the background color, and after filling, divides each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, the processed verification code images are input into the deep convolutional neural network model for recognition, and the recognition results are transmitted to the verification code label processing unit. Filling the easily distinguishable small color blocks with the background color can reduce the pixel value type without affecting the content of the verification code images, further reducing the operation steps of converting the pixel values into floating point numbers between - and improving the processing efficiency. Converting the pixel values into floating point numbers between - is used to avoid the influence of other interference factors on the image content and improve the image accuracy;
[0049] The verification code label processing unit receives the recognition results transmitted by the verification code image denoising processing unit, encodes the recognition results into 62-bit one-hot encoding based on the one-hot algorithm, and finally encodes the recognition results into 63-bit one-hot encoding in the 62 + 1 manner to handle variable-length verification codes. The encoded variable-length verification codes are transmitted to the accuracy calculation unit. The 62-bit one-hot encoding represents including 0-9, a-z, A-Z. The deep convolutional neural network model constructed by the deep convolutional neural network model construction module outputs one more character compared to the bits formed after the initial encoding by the one-hot algorithm. The added character is used to prevent the situation where the characters in the variable-length verification code are empty;
[0050] The accuracy calculation unit receives the encoded variable-length verification code transmitted by the verification code label processing unit, calculates the accuracy of the encoded variable-length verification code. If the calculated accuracy varies greatly after multiple iterations, the initial learning rate set by the model is halved and recalculated. If the calculated accuracy varies little after multiple iterations, the calculation is stopped, and the deep convolutional neural network model obtained by training at this time is saved to the disk and transmitted to the recognition module.
[0051] Further, the recognition module includes a web processing unit, a picture preprocessing unit, and a picture string recognition unit;
[0052] The web processing unit receives the trained model transmitted by the accuracy calculation unit, opens the received trained model in the web service. The UI automation test scenario uploads the verification code image file through the web interface, and transmits the trained model and the uploaded verification code image file to the picture preprocessing unit;
[0053] The picture preprocessing unit receives the trained model and the uploaded verification code image file transmitted by the web processing unit, opens the verification code image file in the trained model, adjusts the verification code image to the standard specification size and makes the verification code image located in the middle of the background in the trained model, then performs pixel grayscale processing on the verification code image, counts the size of the color blocks in the processed image, fills the easily distinguishable small color blocks with the background color after statistics. After filling, each pixel value is divided by, and the integer pixel value is converted into a floating point number between - and. Then the processed verification code image is input into the trained model. Based on the one-hot algorithm, the recognition result is encoded into a -bit one-hot code, and the -bit one-hot code is transmitted to the picture string recognition unit;
[0054] The picture string recognition unit receives the -bit one-hot code transmitted by the picture preprocessing unit, performs argmax calculation on the -bit one-hot code, converts the code into a recognizable string, and returns the recognition result to the UI automation test scenario through the web interface.
[0055] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0056] 1. By using randomly generated digital and letter picture verification codes with various lengths, various fonts, and containing various noises as inputs to construct a deep convolutional neural network model, and adding a character in the constructed deep convolutional neural network to prevent the situation where the characters of the variable-length verification code are empty, the present invention is applicable to the recognition of various types of verification codes, further improving the usage effect of the recognition system.
[0057] 2. When building a deep convolutional neural network model, the present invention calculates the loss value between multi-characters with high edge adhesion, splits the multi-characters into single characters according to the calculated loss value, and the processing rate of the model for single characters is relatively high compared to multi-characters, thereby improving the recognition rate of the verification code by the system and reducing the data processing volume.
[0058] 3. By training the constructed deep convolutional neural network, the present invention calculates the accuracy of the model that is continuously iteratively processed, and finds a training model with higher accuracy according to the calculation result, ensuring that the recognized verification code is basically the same as the standard verification code, and further improving the recognition accuracy of the recognition system. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0060] Figure 1 It is a schematic structural diagram of the working process of an indefinite-length general digital and alphabetic verification code recognition system and method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] Please refer to Figure 1 , the present invention provides a technical solution: an indefinite-length general digital and alphabetic verification code recognition method, and the method includes the following steps:
[0063] Step 1: Randomly generate digital and alphabetic verification code pictures with various lengths, fonts, and noises by using the captcha third-party library. The specific method is as follows:
[0064] Step 1 (Ⅰ). Traverse all fonts in the captcha third-party library, randomly select font sizes and string types, and generate multiple single-character pictures;
[0065] Step 2 (Ⅱ). Randomly rotate, scale, and distort multiple single-character pictures, and randomly select from the processed multiple single-character pictures to combine and generate verification code pictures with indefinite lengths;
[0066] Step Three (III). Randomly add different types of noise to the verification code images generated in Step Two (II) to generate digital and alphabetic verification code images with various lengths, fonts, and noises;
[0067] Step Two: Based on the digital and alphabetic verification code images generated in Step One, construct a deep convolutional neural network model. The specific method for constructing the model is as follows:
[0068] Step Two (1). Select the digital and alphabetic verification code images generated in Step One as the input, perform convolutional processing on the input verification code images, perform pooling and overfitting processing on the convolution-processed verification code images, and output the pooling- and overfitting-processed verification code images. Performing convolutional processing on the verification code images is used to detect the character edges in the verification code images, which is beneficial for feature extraction. Performing pooling processing on the convolution-processed verification code images is used to reduce the redundant information in the convolution-processed images, thereby reducing the output value of the verification code images and improving work efficiency. Performing overfitting processing on the pooling-processed verification code images is used to avoid stagnation or decline when the model accuracy reaches the highest value;
[0069] Step Two (2). Repeat the convolutional, pooling, and overfitting processing on the verification code images output in Step Two (1) until the output value of the verification code images is minimized while being able to recognize the verification code string;
[0070] Step Two (3). Split the strings in the verification code images output in Step Two (2) into single-character images according to the distribution of pixel values after the last convolutional processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, then treat the multi-character as a single-character image for splitting, and number the single-character images according to the specific positions of each single character in the verification code image. Treating the multi-character with glued edge positions as a single-character image for splitting can avoid the final output verification code being different from the actual verification code;
[0071] Step Two (4). If there is a situation where a multi-character with glued edge positions is treated as a single-character image for splitting, then calculate the loss value of the single-character image, and re-split the single-character image according to the calculation result of the loss value, and number the multi-characters according to the specific positions of each multi-character in the single-character image. The specific calculation formula for the loss value L is as follows:
[0072]
[0073] where m represents three character types: 0-9, a-z, A-Z, j = 1, 2, 3. When j = 1, it represents the character type 0-9. When j = 2, it represents the character type a-z. When j = 3, it represents the character type A-Z, and y j represents the output value of the jth character type. When j = 1, y1 Represents the output value of the first character type, where y 1 = -1. When j = 2, y 2 Represents the output value of the second character type, where y 2 = 0. When j = 3, y 3 Represents the output value of the third character type, where y 3 = 1, p j Represents the probability that the judged character is the j-th character type, y j *p j Represents the probability that the character is the j-th character type. When it indicates that the probability of the character type being A-Z is the highest. At this time, the edge bonding part is split into the character pictures connected to it to the greatest extent. When it indicates that the probability of the character type being 0-9 is the highest. At this time, the edge bonding part is split into the character pictures connected to it to the greatest extent. When it indicates that the probability of the character type being a-z is the highest. At this time, the edge bonding part is split into the character pictures connected to it to the greatest extent;
[0074] Step 2(5). Input the numbered single-character pictures into the letter-channel convolutional neural network and the digital-channel convolutional neural network according to the character types respectively, and perform convolution, pooling, and overfitting processing on the single-character pictures input into the letter-channel convolutional neural network and the digital-channel convolutional neural network again until the characters on the single-character pictures can be read, and generate the verification code string by combining the output single characters according to the number order;
[0075] Step 3: Train the deep convolutional neural network model constructed in Step 2 until a training model with a relatively high accuracy is obtained. The specific model training method is as follows:
[0076] Step 3(1). Cut the verification code pictures into a training set and a validation set according to a ratio. Then, reduce the too-large pictures in the training set and the validation set to the standard specification size, and enlarge the too-small pictures according to the height and center them in the center of the standard specification white background;
[0077] Step 3(2). Perform pixel grayscale processing on the verification code pictures cut in Step 3(1), count the size of the color blocks in the processed pictures, fill the easily distinguishable small color blocks after statistics with the background color. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, input the processed verification code pictures into the deep convolutional neural network model for recognition;
[0078] Step 3(3). Based on the one-hot algorithm, encode the recognition result into a 62-bit one-hot encoding. To process variable-length verification codes, use the 62 + 1 method to finally encode the recognition result into a 63-bit one-hot encoding;
[0079] Step 3(4). Calculate the accuracy of the variable-length verification code obtained after encoding. If the calculated accuracy changes significantly after multiple iterations, halve the initial learning rate set by the model and calculate again. If the calculated accuracy changes little after multiple iterations, stop the calculation to obtain the trained model. The specific accuracy calculation formula Q is as follows:
[0080]
[0081] Where n represents the number of characters in the verification code, i = 1, 2, 3, 4, 5, 6, 7, 8 represents the i-th character in the verification code, X represents the 63-bit one-hot encoding, and x represents the number of specific digits or letters in each character. represents the sum of the ratios of specific digits or letters in n characters to the 63-bit one-hot encoding. When there is no empty character in the verification code, x = 8. When Q < 98.4%, it means that the result output by the model at this time does not match the actual situation. When Q = 98.4%, it means that the result output by the model at this time matches the actual situation. When Q > 98.4%, it means that there is an empty character in the output verification code. When the model accuracy is 98.4%, the model training ends;
[0082] Step 4: Use the deep convolutional neural network model trained in Step 3 to recognize the uploaded verification code image. The specific recognition method is as follows:
[0083] Step 4(1). Open the trained model in the web service, and the UI automation test scenario uploads the verification code image file through the web interface;
[0084] Step 4(2). Open the verification code image file in the trained model, adjust the verification code image to the standard specification size in the trained model and make the verification code image located in the middle of the background. Then perform pixel grayscale processing on the verification code image, count the size of the color blocks in the processed image, fill the easily distinguishable small color blocks with the background color after statistics. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then input the processed verification code image into the trained model, and based on the one-hot algorithm, encode the output result of the trained model into a 63-bit one-hot encoding;
[0085] Step 4(3). Perform argmax calculation on the 63-bit one-hot encoding, convert the encoding into a recognizable string, and return the recognition result to the UI automation test scenario through a web interface.
[0086] A variable-length general digital and alphabetic verification code recognition system, including a verification code image generation module S1, a deep convolutional neural network model construction module S2, a training module S3, and a recognition module S4;
[0087] The verification code image generation module S1 is used to generate various lengths, fonts, and noisy digital and alphabetic verification code images required for training a deep convolutional neural network training model through the captcha third-party library, and transmit the generated digital and alphabetic verification code images to the deep convolutional neural network model construction module S2 and the training module S3; the verification code image generation module S1 includes a single-character image generation unit S11, a verification code image generation unit S12, and an image noise generation unit S13;
[0088] The single-character image generation unit S11 traverses the system fonts, generates multiple single-character images with random font sizes and string types, and transmits the generated multiple single-character images to the verification code image generation unit S12;
[0089] The verification code image generation unit S12 receives the multiple single-character images transmitted by the single-character image generation unit S11, randomly rotates, scales, and distorts the received images, randomly combines the processed multiple single-character images to generate a verification code image, and transmits the generated verification code image to the image noise generation unit S13;
[0090] The image noise generation unit S13 receives the verification code image transmitted by the verification code image generation unit S12, performs random noise addition processing on the received verification code image, generates various lengths, fonts, and noisy digital and alphabetic verification code images, saves the generated digital and alphabetic verification code images to the disk, and transmits them to the deep convolutional neural network model S2.
[0091] The deep convolutional neural network model construction module S2 is used to receive the digital and alphabetic verification code images transmitted by the verification code image generation module S1. The deep convolutional neural network model construction module S2 uses tensoflow and keras to construct a deep convolutional neural network model, and transmits the constructed deep convolutional neural network model to the training module S3; the deep convolutional neural network model construction module S2 includes an input unit S21, a processing unit S22, and a deep convolutional neural network model construction unit S23;
[0092] The input unit S21 receives the digital and alphabetic verification code images transmitted by the image noise generation unit S13, and inputs the received digital and alphabetic verification code images into the convolutional neural network;
[0093] The processing unit S22 performs convolution, pooling, and overfitting processing on the digital and alphabetic verification code pictures input into the convolutional neural network, and repeats the convolution, pooling, and overfitting processing on the digital and alphabetic verification code pictures until the output value of the verification code picture is minimized while being able to recognize the verification code string. Then, the processed verification code picture is split into single-character pictures according to the distribution of pixel values after the last convolution processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, the multi-character is regarded as a single-character picture for splitting, and the single-character pictures are numbered according to their specific positions in the verification code picture. Then, the loss value of the single-character pictures containing multi-characters is calculated, the single-character pictures containing multi-characters are split according to the calculation results, and the multi-characters are numbered according to their specific positions in the single-character pictures. The numbered single-character pictures are respectively input into the alphabet channel convolutional neural network and the digital channel convolutional neural network according to the character type, and convolution, pooling, and overfitting processing are performed on the single-character pictures again until the output value of the single-character picture is minimized while being able to recognize the character, and the output single-characters are combined into a verification code string according to the number order, and the entire process of obtaining the verification code string is transmitted to the deep convolutional neural network model construction unit S23;
[0094] The deep convolutional neural network model construction unit S23 receives the entire process of obtaining the verification code string transmitted by the processing unit S22, constructs a deep convolutional neural network model according to the received content, and transmits the constructed deep convolutional neural network model to the training module S3.
[0095] The training module S3 is used to receive the deep convolutional neural network model constructed by the deep convolutional neural network model S2 and the verification code pictures transmitted by the verification code picture generation module S1, train the deep convolutional neural network model based on the verification code pictures, and transmit the trained network model to the recognition module S4; the training module S3 includes a verification code picture preprocessing unit S31, a verification code picture denoising processing unit S32, a verification code label processing unit S33, and an accuracy calculation unit S34;
[0096] The verification code picture preprocessing unit S31 receives the verification code pictures transmitted by the verification code picture generation module S1, splits the verification code pictures into a training set and a validation set according to a ratio, then reduces the too-large pictures in the training set and the validation set to the standard specification size, enlarges the too-small pictures according to the height and places them centered in the middle of the standard specification white background, and transmits the preprocessed verification code pictures to the verification code picture denoising processing unit S32;
[0097] The verification code image denoising processing unit S32 receives the verification code image transmitted by the verification code image preprocessing unit S31, performs pixel grayscale processing on the received verification code image, counts the size of the color blocks in the processed image, fills the distinguishable small color blocks with the background color after statistics. After filling, each pixel value is divided by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, the processed verification code image is input into the deep convolutional neural network model for recognition, and the recognition result is transmitted to the verification code label processing unit S33. Filling the distinguishable small color blocks with the background color can reduce the pixel value type without affecting the content of the verification code image, further reducing the operation steps of converting the pixel value into a floating point number between 0 and 1, improving the processing efficiency. Converting the pixel value into a floating point number between 0 and 1 is used to avoid the influence of other interference factors on the image content and improve the image accuracy;
[0098] The verification code label processing unit S33 receives the recognition result transmitted by the verification code image denoising processing unit S32. Based on the one-hot algorithm, the recognition result is encoded into a 62-bit one hot encoding. To process the variable-length verification code, the recognition result is finally encoded into a 63-bit one hot encoding in the way of 62+1, and the encoded variable-length verification code is transmitted to the accuracy calculation unit S34. The 62-bit one hot encoding represents including 0-9, a-z, A-Z. The output of the deep convolutional neural network model constructed by the deep convolutional neural network model construction module S2 is 63, which is one character more than the 62-bit after the initial encoding by the one-hot algorithm. The added character is used to prevent the situation that the characters of the variable-length verification code are empty.
[0099] The accuracy calculation unit S34 receives the encoded variable-length verification code transmitted by the verification code label processing unit S33, calculates the accuracy of the encoded variable-length verification code. If the calculated accuracy changes greatly after multiple iterations, the initial learning rate set by the model is halved and calculated again. If the calculated accuracy changes little after multiple iterations, the calculation is stopped, and the trained deep convolutional neural network model at this time is saved to the disk and transmitted to the recognition module S4.
[0100] The recognition module S4 is used to receive the training model transmitted by the training module S3. The training model encodes and recognizes the verification code image uploaded through the web interface, and then converts the encoding into a string through argmax calculation to obtain the recognition result. The recognition module S4 includes a web processing unit S42, an image preprocessing unit S42, and an image string recognition unit S43;
[0101] The web processing unit S42 receives the training model transmitted by the accuracy calculation unit S34, opens the received training model in the web service, the UI automation test scenario uploads the verification code image file through the web interface, and transmits the training model and the uploaded verification code image file to the picture preprocessing unit S42;
[0102] The picture preprocessing unit S42 receives the training model and the uploaded verification code image file transmitted by the web processing unit S42, opens the verification code image file in the training model, adjusts the verification code image to the standard specification size and makes the verification code image located in the middle of the background in the training model, then performs pixel grayscale processing on the verification code image, counts the size of the color blocks in the processed image, fills the distinguishable small color blocks after statistics with the background color, after filling, divides each pixel value by 255, converts the integer pixel value into a floating point number between 0 and 1, then inputs the processed verification code image into the training model, based on the one-hot algorithm, encodes the output result of the training model into a 63-bit one hot encoding, and transmits the 63-bit one hot encoding to the picture string recognition unit S43;
[0103] The picture string recognition unit S43 receives the 63-bit one hot encoding transmitted by the picture preprocessing unit S42, performs argmax calculation on the 63-bit one hot encoding, converts the encoding into a recognizable string, and returns the recognition result to the UI automation test scenario through the web interface.
[0104] Embodiment 1:
[0105] captcha represents a verification code service suitable for high concurrency and easy integration;
[0106] tensoflow represents a symbolic mathematics system based on dataflow programming;
[0107] Keras represents an open-source artificial neural network library written in Python, which can be used as a high-level application programming interface for Tensorflow to design, debug, evaluate, apply and visualize deep learning models;
[0108] The one-hot algorithm refers to using an N-bit status register to encode N states, each state has its own independent register bit, and only one bit is valid at any time;
[0109] The learning rate represents an important hyperparameter in supervised learning and deep learning. The learning rate determines whether the objective function can converge to the local minimum and when it converges to the minimum;
[0110] argmax is a function that finds the argument (set) of a function.
[0111] Example 2:
[0112] Suppose the string in the input verification code image is "1234567" and the length of the verification code characters is 8 bits. Then, after encoding the output result of the deep convolutional neural network model, it is as follows:
[0113] The first character: 010000000000000000000000000000000000000000000000000000000000000;
[0114] The second character: 001000000000000000000000000000000000000000000000000000000000000;
[0115] The third character: 000100000000000000000000000000000000000000000000000000000000000;
[0116] The fourth character: 000010000000000000000000000000000000000000000000000000000000000;
[0117] The fifth character: 000001000000000000000000000000000000000000000000000000000000000;
[0118] The sixth character: 000000100000000000000000000000000000000000000000000000000000000;
[0119] The seventh character: 000000010000000000000000000000000000000000000000000000000000000;
[0120] The eighth character: 000000000000000000000000000000000000000000000000000000000000000;
[0121] Since the verification code "1234567" is less than 8 bits, the last line of encoded representation is added to indicate that the character is empty, in order to achieve the purpose of using a unified length of encoded representation for variable-length verification codes.
[0122] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0123] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An identification method for variable-length general digital and alphabetic verification codes, characterized in that: The method includes the following steps: Step 1: Randomly generate digital and alphabetic verification code images of various lengths, fonts, and noises by using the captcha third-party library; The specific method for randomly generating digital and alphabetic verification code images of various lengths, fonts, and noises in Step 1 by using the captcha third-party library is: Step 1(I). Traverse all fonts in the captcha third-party library, randomly select font sizes and string types, and generate multiple single-character images; Step 2(II). Randomly rotate, scale, and distort multiple single-character images, and randomly select from the processed single-character images to combine and generate verification code images of variable lengths; Step 3(III). Randomly add different types of noises to the verification code images generated in Step 2(II) to generate digital and alphabetic verification code images of various lengths, fonts, and noises; Step 2: Build a deep convolutional neural network model based on the digital and alphabetic verification code images generated in Step 1; The specific method for building a deep convolutional neural network model in Step 2 is: Step 2(1). Select the digital and alphabetic verification code images generated in Step 1 as the input, perform convolutional processing on the input verification code images, perform pooling and overfitting processing on the convolutionally processed verification code images, and output the pooled and overfitted verification code images; Step 2(2). Repeatedly perform convolution, pooling, and overfitting processing on the verification code images output in Step 2(1) until the verification code string can be recognized while minimizing the output value of the verification code images; Step 2(3). Split the strings in the verification code images output in Step 2(2) into single-character images according to the distribution of pixel values after the last convolutional processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, then treat the multi-character as a single-character image for splitting, and number the single-character images according to their specific positions in the verification code images; Step 2(4). If there is a situation where a multi-character with glued edge positions is treated as a single-character image for splitting, then calculate the loss value of the single-character image, re-split the single-character image according to the calculation result of the loss value, and number the multi-characters according to their specific positions in the single-character image. The specific calculation formula for the loss value L is: Among them, m represents three character types: 0 - 9, a - z, and A - Z. j = 1, 2, 3. When j = 1, it represents the character type of 0 - 9. When j = 2, it represents the character type of a - z. When j = 3, it represents the character type of A - Z, y j represents the output value of the j - th character type. When j = 1, y 1 represents the output value of the 1 - st character type. At this time, y 1 = - 1. When j = 2, y 2 represents the output value of the 2 - nd character type. At this time, y 2 = 0. When j = 3, y 3 represents the output value of the 3 - rd character type. At this time, y 3 = 1, p j represents the probability that the judged character is of the j - th character type. y j* p j represents the probability that the character is of the j - th character type; Step 2(5). Input the numbered single-character images into the alphabet channel convolutional neural network and the digital channel convolutional neural network according to the character type respectively, and perform convolution, pooling, and overfitting processing on the single-character images input into the alphabet channel convolutional neural network and the digital channel convolutional neural network again until the characters on the single-character images can be read, and combine the output single-characters to generate the verification code string according to the numbering order; Step 3: Train the deep convolutional neural network model built in Step 2 until a training model with a relatively high accuracy is obtained; Step 4: Use the deep convolutional neural network model trained in Step 3 to identify the uploaded verification code images.
2. A method for recognizing variable-length general digital and alphabetic verification codes according to claim 1, characterized in that: The specific method of training the deep convolutional neural network model constructed in step 2 in step 3 until a training model with higher accuracy is obtained is: Step 3(1). Cut the verification code pictures into a training set and a validation set according to a ratio. Then, reduce the too-large pictures in the training set and the validation set to the standard specification size, and enlarge the too-small pictures according to the height and place them centered in the middle of the standard specification white background; Step 3(2). Perform pixel grayscale processing on the verification code pictures cut in step 3(1), count the sizes of the color blocks in the processed pictures, fill the easily distinguishable small color blocks after statistics with the background color. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, input the processed verification code pictures into the deep convolutional neural network model for recognition; Step 3(3). Based on the one-hot algorithm, encode the recognition result into a 62-bit one-hot encoding. To process variable-length verification codes, use the 62+1 method to finally encode the recognition result into a 63-bit one-hot encoding; Step 3(4). Calculate the accuracy of the variable-length verification code obtained after encoding. If the calculated accuracy changes greatly after multiple iterations, halve the initial learning rate set by the model and calculate again. If the calculated accuracy changes little after multiple iterations, stop the calculation to obtain the training model. The specific accuracy calculation formula Q is: Wherein, n represents the number of characters in the verification code, i = 1, 2, 3, 4, 5, 6, 7, 8 represents the i-th character in the verification code, X represents a 63-bit one-hot encoding, and x represents the number of specific digits or letters represented in each character. It represents the sum of the ratios of the specific digits or letters in n characters to the 63-bit one-hot encoding.
3. A method for recognizing variable-length general digital and alphabetic verification codes according to claim 2, characterized in that: The specific method of using the deep convolutional neural network model trained in step 3 to recognize the uploaded verification code pictures in step 4 is: Step 4(1). Open the training model in the web service, and the UI automation test scenario uploads the verification code image file through the web interface; Step 4(2). Open the verification code image file in the training model, adjust the verification code image to the standard specification size in the training model and make the verification code image located in the middle of the background. Then, perform pixel grayscale processing on the verification code image, count the sizes of the color blocks in the processed image, fill the easily distinguishable small color blocks after statistics with the background color. After filling, divide each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, input the processed verification code image into the training model. Based on the one-hot algorithm, encode the output result of the training model into a 63-bit one-hot encoding; Step 4(3). Perform argmax calculation on the 63-bit one-hot encoding, convert the encoding into a recognizable string, and return the recognition result to the UI automation test scenario through the web interface.
4. A variable-length general digital and alphabetic verification code recognition system, characterized in that: It includes a verification code picture generation module (S1), a deep convolutional neural network model construction module (S2), a training module (S3) and an identification module (S4); The verification code image generation module (S1) is used to generate various digital and alphabetic verification code images with different lengths, fonts, and noises required for training a deep convolutional neural network training model through the captcha third-party library, and transmit the generated digital and alphabetic verification code images to the deep convolutional neural network model construction module (S2) and the training module (S3); The verification code image generation module (S1) includes a single-character image generation unit (S11), a verification code image generation unit (S12), and an image noise generation unit (S13); The single-character image generation unit (S11) traverses the system fonts, generates multiple single-character images with random font sizes and string types, and transmits the generated multiple single-character images to the verification code image generation unit (S12); The verification code image generation unit (S12) receives the multiple single-character images transmitted by the single-character image generation unit (S11), randomly rotates, scales, and distorts the received images, randomly combines the processed multiple single-character images to generate verification code images, and transmits the generated verification code images to the image noise generation unit (S13); The image noise generation unit (S13) receives the verification code images transmitted by the verification code image generation unit (S12), randomly adds noise to the received verification code images, generates digital and alphabetic verification code images with different lengths, fonts, and noises, saves the generated digital and alphabetic verification code images to the disk, and transmits them to the deep convolutional neural network model (S2); The deep convolutional neural network model construction module (S2) is used to receive the digital and alphabetic verification code images transmitted by the verification code image generation module (S1). The deep convolutional neural network model construction module (S2) uses tensoflow and keras to construct a deep convolutional neural network model, and transmits the constructed deep convolutional neural network model to the training module (S3); The deep convolutional neural network model construction module (S2) includes an input unit (S21), a processing unit (S22), and a deep convolutional neural network model construction unit (S23); The input unit (S21) receives the digital and alphabetic verification code images transmitted by the image noise generation unit (S13), and inputs the received digital and alphabetic verification code images into the convolutional neural network; The processing unit (S22) performs convolution, pooling, and overfitting processing on the digital and alphabetic verification code pictures input into the convolutional neural network, and repeats the convolution, pooling, and overfitting processing on the digital and alphabetic verification code pictures until the output value of the verification code picture is minimized while being able to recognize the verification code string. Then, the processed verification code picture is split into single-character pictures according to the distribution of pixel values after the last convolution processing. If it is determined through pixel values that the edge positions of multiple characters are glued together, the multi-character is regarded as a single-character picture for splitting, and the single-character pictures are numbered according to their specific positions in the verification code picture. Then, the loss value of the single-character pictures containing multi-characters is calculated, the single-character pictures containing multi-characters are split according to the calculation results, and the multi-characters are numbered according to their specific positions in the single-character pictures. The numbered single-character pictures are respectively input into the alphabet channel convolutional neural network and the digital channel convolutional neural network according to the character type, and the single-character pictures are subjected to convolution, pooling, and overfitting processing again until the output value of the single-character picture is minimized while being able to recognize the character, and the output single-characters are combined into a verification code string according to the numbering order, and the entire process of obtaining the verification code string is transmitted to the deep convolutional neural network model construction unit (S23); The deep convolutional neural network model construction unit (S23) receives the entire process of obtaining the verification code string transmitted by the processing unit (S22), constructs a deep convolutional neural network model according to the received content, and transmits the constructed deep convolutional neural network model to the training module (S3); The training module (S3) is used to receive the deep convolutional neural network model constructed by the deep convolutional neural network model (S2) and the verification code pictures transmitted by the verification code picture generation module (S1), train the deep convolutional neural network model based on the verification code pictures, and transmit the trained network model to the recognition module (S4); The recognition module (S4) is used to receive the trained model transmitted by the training module (S3). The trained model encodes and recognizes the verification code pictures uploaded through the web interface, and then converts the encoding into a string through argmax calculation to obtain the recognition result.
5. An arbitrary-length general digital and alphabetic verification code recognition system according to claim 4, characterized in that: The training module (S3) includes a verification code picture preprocessing unit (S31), a verification code picture denoising processing unit (S32), a verification code label processing unit (S33), and an accuracy calculation unit (S34); The verification code picture preprocessing unit (S31) receives the verification code pictures transmitted by the verification code picture generation module (S1), splits the verification code pictures into a training set and a validation set according to a ratio, then shrinks the too-large pictures in the training set and the validation set to the standard specification size, enlarges the too-small pictures according to the height and places them centered in the middle of the standard specification white background, and transmits the preprocessed verification code pictures to the verification code picture denoising processing unit (S32); The verification code image denoising processing unit (S32) receives the verification code image transmitted by the verification code image preprocessing unit (S31), performs pixel grayscale processing on the received verification code image, counts the size of color blocks in the processed image, fills the distinguishable small color blocks after statistics with the background color, divides each pixel value by 255 after filling, converts the integer pixel value into a floating point number between 0 and 1, then inputs the processed verification code image into the deep convolutional neural network model for recognition, and transmits the recognition result to the verification code label processing unit (S33); The verification code label processing unit (S33) receives the recognition result transmitted by the verification code image denoising processing unit (S32), encodes the recognition result into a 62-bit one-hot encoding based on the one-hot algorithm, and finally encodes the recognition result into a 63-bit one-hot encoding in the 62+1 manner to process the variable-length verification code, and transmits the encoded variable-length verification code to the accuracy calculation unit (S34); The accuracy calculation unit (S34) receives the encoded variable-length verification code transmitted by the verification code label processing unit (S33), calculates the accuracy of the encoded variable-length verification code. If the calculated accuracy changes greatly after multiple iterations, the initial learning rate set by the model is halved and recalculated. If the calculated accuracy changes little after multiple iterations, the calculation is stopped, and the deep convolutional neural network model obtained by training at this time is saved to the disk and transmitted to the recognition module (S4).
6. An indeterminate-length general digital and alphabetic verification code recognition system according to claim 5, characterized in that: the recognition module (S4) includes a web processing unit (S42), an image preprocessing unit (S42) and an image string recognition unit (S43); the web processing unit (S42) receives the training model transmitted by the accuracy calculation unit (S34), opens the received training model in the web service, the UI automation test scenario uploads the verification code image file through the web interface, and transmits the training model and the uploaded verification code image file to the image preprocessing unit (S42); The described picture preprocessing unit (S42) receives the training model transmitted by the web processing unit (S42) and the uploaded verification code image file, opens the verification code image file in the training model, adjusts the verification code image to the standard specification size and makes the verification code image located in the middle of the background in the training model, then performs pixel grayscale processing on the verification code image, counts the size of the color blocks in the processed image, fills the distinguishable small color blocks after statistics with the background color, and after filling, divides each pixel value by 255 to convert the integer pixel value into a floating point number between 0 and 1. Then, the processed verification code image is input into the training model, and based on the one-hot algorithm, the output result of the training model is encoded into a 63-bit one-hot encoding, and the 63-bit one-hot encoding is transmitted to the picture string recognition unit (S43); The described picture string recognition unit (S43) receives the 63-bit one-hot encoding transmitted by the picture preprocessing unit (S42), performs argmax calculation on the 63-bit one-hot encoding, converts the encoding into a recognizable string, and returns the recognition result to the UI automation test scenario through the web interface.
Citation Information
Patent Citations
Verification code identification method and system based on deep learning
CN109933975A
Captcha identification method and apparatus, and computer device and storage medium
WO2020215573A1