Verification code automatic identification method based on large language model
By introducing large language models and multimodal learning methods into verification code recognition technology, a multi-platform and multi-model verification code automatic recognition framework is built, which solves the problem of low recognition rate of verification codes with more complex backgrounds and noise in the existing technology, and achieves higher accuracy, robustness and stronger adaptability verification code recognition effect.
Patent Information
- Application Number
- CN202510044185.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-05-16
AI Technical Summary
When existing verification code recognition technology deals with verification codes with complex backgrounds, distorted or noise, the recognition rate is low, and when the deep learning model faces new verification codes or specific design verification codes, the generalization ability is insufficient and the recognition effect is unstable.
A method of automatic verification code recognition based on large language models is proposed. By introducing multi-modal learning methods, a multi-platform, multi-modal verification code automatic recognition framework is built, and models of OpenAI and Google platforms are called through interfaces, reducing dependence on a large amount of labeled data, and improving the accuracy and robustness of verification code recognition.
It improves the accuracy and robustness of verification code recognition, enhances the adaptability to verification codes from different sources and styles, reduces the cost and complexity of training and maintenance, and achieves a more stable verification code recognition effect.
Smart Images

Figure CN120014655A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a verification code automatic recognition method based on a large language model. Background Art
[0002] CAPTCHA is a fully automated public Turing test for distinguishing computers from humans. It is a commonly used mechanism in network security that can verify the identity of users and prevent malicious operations by automated programs. With the development of artificial intelligence technology, computer vision and deep learning algorithms have made significant progress in identifying CAPTCHAs. Types of CAPTCHAs include text, images, sounds, and dynamic CAPTCHAs. CAPTCHA recognition technology involves computer vision, pattern recognition, traditional machine learning, and deep learning. With the advancement of technology, the design of CAPTCHAs is also constantly innovating to ensure network security while improving user experience.
[0003] It is a common method to use OCR technology (Optical Character Recognition) combined with image processing technology (such as OpenCV, etc.) to recognize verification codes. These technologies include pre-processing operations such as grayscale, binarization, and denoising of images to improve recognition accuracy. However, these methods are mainly suitable for simple text verification code recognition. When dealing with verification codes with complex backgrounds, distortions, or more noise, the recognition rate will decrease.
[0004] With the development of deep learning technology, CNN (Convolutional Neural Networks) and RNN (Recursive Neural Network) are widely used in more complex image CAPTCHA recognition. CNN can recognize patterns and shapes in images through its powerful image feature extraction capabilities, while RNN is good at processing sequence data and is suitable for CAPTCHA recognition problems of variable length. These models can improve the recognition rate of CAPTCHAs.
[0005] However, the CAPTCHA testing method based on deep models requires developers to train the model themselves, which makes it difficult to collect training data and the training process time-consuming and labor-intensive. In addition, even deep learning models may show vulnerability in the face of specially designed adversarial attacks such as noise injection, occlusion, deformation, etc. In addition, deep models perform well on specific types of CAPTCHAs, but when faced with CAPTCHAs from different sources and styles, their generalization ability is limited and their recognition ability decreases. This is because deep models usually require a large amount of labeled data for training, and CAPTCHAs from different sources may have different styles and characteristics, which limits the generalization ability of the model.
[0006] openai-captcha-detection is a tool that uses OpenAI for CAPTCHA recognition. The tool implements text recognition of complex CAPTCHA images by calling OpenAI's API interface, helping developers to automate operations in CAPTCHA processing scenarios. The simple and easy-to-use API interface makes the tool convenient for integration in other projects. The tool supports multiple types of CAPTCHA recognition and provides detailed usage examples and code. Its advantages are high efficiency and accuracy, and it can process complex CAPTCHAs containing text and images, which is difficult for many traditional OCR tools to achieve. In addition, the tool's high recognition accuracy means that it can provide reliable CAPTCHA recognition services in a variety of application scenarios, thereby reducing the need for manual intervention.
[0007] Although the openai-captcha-detection tool has shown significant advantages in the field of captcha recognition with its efficiency and accuracy, it still has some limitations. First, the tool faces the challenge of large-scale testing because it relies on the call of OpenAI's API interface, and will encounter performance bottlenecks and cost issues when processing a large number of requests. Secondly, for the constantly evolving and updated captcha technology, the tool's adaptability and generalization capabilities are limited, especially when faced with new captchas or specifically designed captchas, its recognition effect is very unstable. In addition, since the accuracy of captcha recognition is affected by image quality, complexity and diversity, it is difficult for the tool to always maintain a high accuracy rate in practical applications. Summary of the invention
[0008] In order to solve the above technical problems, the embodiments of the present application propose a method for automatic verification code recognition based on a large language model, which aims to improve the accuracy, robustness and generalization ability of verification code recognition technology by introducing a multimodal learning method, while reducing the dependence on a large amount of labeled data and reducing the cost and complexity of training and maintenance.
[0009] In order to achieve the above object, an embodiment of the present application proposes a verification code automatic recognition method based on a large language model, the method comprising:
[0010] Obtain several English text verification codes, Chinese text verification codes, image verification codes, and visual reasoning verification codes, and after labeling them, form a small-scale test set and a large-scale test set;
[0011] A small-scale test set was used to conduct preliminary tests on the three large language models, chatgpt-4o-latest, gpt-4-turbo, and gemini-1.5-flash, so that the three large language models could acquire basic verification code recognition capabilities.
[0012] Set up interface calls and build a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. The verification code automatic recognition framework supports calling the OpenAI platform and the Google platform through interfaces. The OpenAI platform is equipped with chatgpt-4o-latest and gpt-4-turbo, and the Google platform is equipped with gemini-1.5-flash.
[0013] After successfully calling the interface, we gradually use large-scale test sets to test the verification code automatic recognition framework, and evaluate the verification code recognition capabilities of the three large language models based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and obtain a mature verification code automatic recognition framework;
[0014] Input the verification code to be recognized into a mature verification code automatic recognition framework to obtain the recognition result.
[0015] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a verification code automatic recognition method based on a large language model as described above.
[0016] In order to achieve the above-mentioned purpose, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement a verification code automatic recognition method based on a large language model as described above.
[0017] The embodiment of the present application proposes a verification code automatic recognition method based on a large language model, and builds a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. Not only can chatgpt-4o-latest and gpt-4-turbo provided in the OpenAI platform be used for verification code automatic recognition, but also gemini-1.5-flash provided in the Google platform can be used for verification code automatic recognition, and the call of multiple platforms and multiple models is successfully realized. These large language models can process image information and text information at the same time, improve the accuracy and robustness of verification code recognition, and reduce the dependence on a large amount of annotated data, and reduce the cost and complexity of training and maintenance. For the multi-platform, multi-model, and multi-modal verification code automatic recognition framework, this application conducted two parts of testing. Before building the verification code automatic recognition framework, a small-scale test set was used to conduct preliminary tests on the three large language models, so that the three large language models obtained basic verification code recognition capabilities. The design of the preliminary test shortened the test time and saved test resources. After the verification code automatic recognition framework is built, it is gradually tested using large-scale test sets. The verification code recognition capabilities of the three large language models are evaluated based on the recognition accuracy until the verification code recognition capabilities of the three large language models meet the preset performance standards. In this way, an accurate, secure, robust and mature verification code automatic recognition framework can be obtained, thereby realizing the verification code automatic recognition excellently.
[0018] Optionally, obtain several English text verification codes, Chinese text verification codes, image verification codes, and visual reasoning verification codes, and after labeling, form a small-scale test set and a large-scale test set, including:
[0019] Crawl several English text verification codes, several Chinese text verification codes, several image verification codes, and several visual reasoning verification codes from different Internet websites, and obtain English text verification code dataset R, Chinese text verification code dataset C, image verification code dataset T, and visual reasoning verification code dataset S respectively;
[0020] Mark each verification code with its serial number and the correct recognition result as a label;
[0021] Randomly select M1 verification codes from each data set to form a small-scale test set D tests ;
[0022] Randomly select M2 verification codes from each data set to form a large-scale test set D testb , and the large-scale test set D testb Randomly divided into m test groups, Among them, M2>M1.
[0023] Optionally, use a small-scale test set to conduct preliminary tests on the three large language models of chatgpt-4o-latest, gpt-4-turbo, and gemini-1.5-flash, so that the three large language models can obtain basic verification code recognition capabilities, including:
[0024] Build a small client compatible with mainstream platforms and use the small-scale test set D tests The verification codes in the code are input into the three large language models chatgpt-4o-latest, gpt-4-turbo and gemini-1.5-flash for recognition;
[0025] Based on the comparison of the recognition results and corresponding labels output by the three large language models, we can get the recognition results of the three large language models on the small-scale test set D. tests The recognition accuracy rate;
[0026] If three large language models are trained on a small test set D tests If the recognition accuracy rates of the three language models meet the preset preliminary test standards, the preliminary test of the three language models is completed to determine that the three language models have acquired basic verification code recognition capabilities;
[0027] If there is at least one large language model for the small test set D tests If the recognition accuracy rate does not meet the preset preliminary test standard, a new small-scale test set is constructed, and the verification code in the new small-scale test set is input into the large language model that does not meet the preset preliminary test standard, until the three large language models all obtain basic verification code recognition capabilities.
[0028] Optionally, set up an interface call, including:
[0029] Get API access rights and obtain the API Base URL and API Key from the agency CloseAI. The API Key is the credential for accessing and using OpenAI services, and the API Base URL is the base address of the API server, which is used to construct a complete API request URL.
[0030] Configure environment variables. Specify the address and port of the proxy server by setting environment variables os.environ["http_proxy"] and os.environ["https_proxy"] to ensure that all API requests go through the specified proxy server.
[0031] Initialize the OpenAI client and use the Python library provided by OpenAI to create an OpenAI client instance for all subsequent API calls. When creating the OpenAI client, pass in the API Base URL and API Key.
[0032] After configuring the OpenAI client, perform an interface call test to verify whether it can successfully interact with the CloseAI interface. If it is verified that it cannot successfully interact with the CloseAI interface, reset the interface call until it is verified that it can successfully interact with the CloseAI interface.
[0033] Optionally, in the multi-platform, multi-model, and multi-modal verification code automatic recognition framework, the OpenAI platform accesses and performs verification code recognition through the following steps:
[0034] S21, environment configuration and initialization;
[0035] S211, set up proxy, configure HTTP and HTTPS proxy, point to local proxy server 127.0.0.1:7890;
[0036] S212, initialize the OpenAI client, use the API key and the proxy server address to initialize the OpenAI client, so as to call the OpenAI API interface later;
[0037] S22, image preprocessing and coding;
[0038] S221, image encoding, defines the encode_image function, uses the io.BytesIO() object to store the binary data of the image, converts the image file into a base64-encoded image string, and sends it to the model in the OpenAI platform;
[0039] S222, read the image list from the specified input folder path P in Read all image files and filter out files with png, jpg and jpeg as suffixes;
[0040] S223, batch image encoding, calling the encode_image function for each image file read, converting all image files into base64-encoded image strings, and storing them in the encoded_images list;
[0041] S23, calling a large language model to perform verification code recognition;
[0042] S231, model call, for each base64-encoded image string in the encoded_images list, use the OpenAI client to call chatgpt-4o-latest and gpt-4-turbo for verification code recognition;
[0043] S232, constructing a request message, constructing a message message including a text prompt and image data, wherein the text prompt instructs the large language model to recognize text in the image and output the recognition result in a specific format;
[0044] S233, processing the response, for each request, processing the response of the large language model, and extracting the recognized text;
[0045] S234, result recording, recording the path of each image file and the recognized text in the specified output folder path P out middle;
[0046] S235, Error handling and logging, when processing each image, error handling and logging are added to facilitate tracking of processing progress and identification of problems that occur during the process;
[0047] S236, performance evaluation, by analyzing the recognition results, evaluate the accuracy and robustness of ChatGPT-4O-latest and GPT-4-Turbo in recognizing verification codes of different modalities.
[0048] Optionally, in the multi-platform, multi-model, and multi-modal verification code automatic recognition framework, the Google platform accesses and performs verification code recognition through the following steps:
[0049] S31, environment configuration, setting up a proxy server to ensure that requests can be sent through the specified proxy server;
[0050] S32, API configuration, using the configure function to configure the Google API key, where the Google API key is the credential for accessing the Google Generative AI service;
[0051] S33, model initialization, initialize the Gemini model instance, specify the large language model gemini-1.5-flash;
[0052] S34, image preprocessing, from the specified input folder path P in In the process, all image files are read out according to the test group, and one test group is processed at a time;
[0053] S35, image recognition, traverse each image file, use gemini-1.5-flash's generate_content method to perform verification code recognition, the generate_content method accepts a list containing prompt text and image objects;
[0054] S36, result processing and recording, extracting the recognition result according to the response of gemini-1.5-flash;
[0055] S37, exception handling, during the recognition process, if an exception occurs, capture the exception and record the error information;
[0056] S38, completion notification, after processing all image files, print the completion notification and save the corresponding recognition results to the specified output folder path P out middle.
[0057] Optionally, the preset performance standards include accuracy performance standards, security performance standards and robustness performance standards. After successfully calling the interface, the verification code automatic recognition framework is gradually tested using a large-scale test set, and the verification code recognition capabilities of the three large language models are evaluated based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and a mature verification code automatic recognition framework is obtained, including:
[0058] After successfully calling the interface, we gradually used a large-scale test set to test the automatic verification code recognition framework and evaluate the accuracy of the three large language models in recognizing verification codes of different modalities.
[0059] Analyze the performance of the three large language models under different types of attacks, including noise injection, occlusion, and deformation, and evaluate their ability to resist attacks to determine the security of the three large language models;
[0060] Analyze the generalization ability of the three large language models and examine their adaptability to verification codes from different sources and modalities to determine the robustness of the three large language models;
[0061] Determine whether the accuracy, security, and robustness of the three large language models meet the accuracy performance standard, security performance standard, and robustness performance standard respectively;
[0062] If the accuracy, security and robustness of the three large language models meet the accuracy performance standard, security performance standard and robustness performance standard respectively, a mature verification code automatic recognition framework is obtained;
[0063] If the accuracy, security or robustness of at least one large language model does not meet the accuracy performance standard, security performance standard or robustness performance standard, a new large-scale test set is constructed, and the verification code automatic recognition framework is continued to be tested using the new large-scale test set until the accuracy, security and robustness of the three large language models meet the accuracy performance standard, security performance standard and robustness performance standard respectively.
[0064] Optionally, the accuracy is measured by single character recognition accuracy and overall recognition accuracy, which are calculated by the following formula:
[0065] SCAR=(N s / N t )×100%;
[0066] ASR=(N r / N a )×100%;
[0067] Among them, N s To correctly identify the number of characters, N t is the total number of characters, SCAR represents the single character recognition accuracy, N r To correctly identify the number of samples, N a is the total number of samples, and ASR represents the overall recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the drawings required for use in the embodiments of the present application or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0069] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements, and the implementation methods described in the following embodiments do not represent all implementation methods consistent with the present application. On the contrary, they are only examples consistent with some aspects of the present application.
[0070] Figure 1 It is a flowchart of a verification code automatic recognition method based on a large language model provided in one embodiment of the present application;
[0071] Figure 2 It is a structural diagram of a multi-platform, multi-model, and multi-modal verification code automatic recognition framework provided in one embodiment of the present application;
[0072] Figure 3It is a structural schematic diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. It will be appreciated by those skilled in the art that in the embodiments of the present application, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is for the convenience of description, and the specific implementation of the present application should not be limited in any way. The various embodiments can be combined and referenced with each other under the premise of no contradiction.
[0074] An embodiment of the present application proposes a verification code automatic recognition method based on a large language model, which is applied to an electronic device, wherein the electronic device can be a terminal or a server. This embodiment and the following embodiments are all described by taking the server as an example. The implementation details of the verification code automatic recognition method based on a large language model proposed in this embodiment are specifically described below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this solution.
[0075] The specific process of the verification code automatic recognition method based on a large language model proposed in this embodiment can be as follows: Figure 1 As shown, including:
[0076] S11, obtain several English text verification codes, Chinese text verification codes, image verification codes and visual reasoning verification codes, and after labeling, form a small-scale test set and a large-scale test set.
[0077] In the specific implementation, the server first needs to collect data and build a test set to test the three large language models. The server needs to obtain several English text verification codes, Chinese text verification codes, image verification codes, and visual reasoning verification codes, and after labeling, form a small-scale test set and a large-scale test set. Both the small-scale test set and the large-scale test set contain multi-modal verification code data.
[0078] In an example, the server can crawl several English text verification codes, several Chinese text verification codes, several image verification codes, and several visual reasoning verification codes from different Internet websites to obtain English text verification code datasets R = {R1, R2, R3, ..., R n}、Chinese text verification code dataset C={C1,C2,C3,…,C n}、Image verification code dataset T={T1,T2,T3,…,Tn}, and visual reasoning verification code dataset S = {S1, S2, S3, ..., S n}.
[0079] Afterwards, the server needs to label each verification code with its serial number and the correct recognition result as a label. For English text verification codes and Chinese text verification codes, the correct recognition result is the correct recognition of the characters. For image verification codes, the correct recognition result is the image content, such as a train or a bicycle. For visual reasoning verification codes, the correct recognition result is the visual reasoning action, such as a sliding action or a click sequence action.
[0080] Next, the server needs to randomly select M1 verification codes from each data set (R, C, T, S) to form a small-scale test set D tests M1 is usually set to 20, so D tests The size is 80.
[0081] Finally, the server needs to randomly select M2 verification codes from each data set (R, C, T, S) to form a large-scale test set D testb , and the large-scale test set D testb Randomly divided into m test groups, so D testb It can be expressed as Among them, M2>M1, M2 is generally set to 100, then D tests The size of D is 400. m is usually set to 7, and the size of the 7 test group can be different. testb The test groups are randomly divided into m groups because gemini-1.5-flash has a rate limit of 15RPM (requests per minute). If the number is too large, gemini-1.5-flash cannot process it in time.
[0082] S12, using a small-scale test set to conduct preliminary tests on the three large language models chatgpt-4o-latest, gpt-4-turbo and gemini-1.5-flash, so that the three large language models can obtain basic verification code recognition capabilities.
[0083] In the specific implementation, after the server completes the construction of the small-scale test set and the large-scale test set, it can first use the small-scale test set to conduct preliminary tests on the three large language models of chatgpt-4o-latest, gpt-4-turbo and gemini-1.5-flash, so that the three large language models can obtain basic verification code recognition capabilities.
[0084] In one example, the server needs to build a small client that is compatible with mainstream platforms and send a small test set Dtests The verification codes in the code are input into the three large language models chatgpt-4o-latest, gpt-4-turbo and gemini-1.5-flash for recognition. The recognition results output by the three large language models are compared with the corresponding labels to obtain the recognition results of the three large language models on the small-scale test set D. tests If the three large language models are trained on a small test set D tests If the recognition accuracy of the three large language models meets the preset preliminary test standards, the server completes the preliminary test of the three large language models and determines that the three large language models have obtained basic verification code recognition capabilities. tests If the recognition accuracy rate does not meet the preset preliminary test standard, the server needs to build a new small-scale test set and input the verification code in the new small-scale test set into the large language model that does not meet the preset preliminary test standard until the three large language models all obtain basic verification code recognition capabilities.
[0085] S13, set up interface calls, build a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. The verification code automatic recognition framework supports calling the OpenAI platform and the Google platform through interfaces. Chatgpt-4o-latest and gpt-4-turbo are set in the OpenAI platform, and gemini-1.5-flash is set in the Google platform.
[0086] In the specific implementation, after the server completes the preliminary test of the three major language models, you can set up interface calls to build a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. The built multi-platform, multi-model, and multi-modal verification code automatic recognition framework supports calling the OpenAI platform and the Google platform through interfaces. Chatgpt-4o-latest and gpt-4-turbo are set in the OpenAI platform, and gemini-1.5-flash is set in the Google platform.
[0087] In an example, the structure of the multi-platform, multi-model, and multi-modal verification code automatic recognition framework can be as follows: Figure 2 shown.
[0088] In one example, when setting up an interface call, the server needs to obtain API access rights and obtain the API Base URL and API Key from the agency CloseAI. Among them, the API Key is the credential for accessing and using the OpenAI service, and the API Base URL is the base address of the API server, which is used to construct a complete API request URL. Next, you need to configure the environment variables. By setting the environment variables os.environ["http_proxy"] and os.environ["https_proxy"], you can specify the address and port of the proxy server to ensure that all API requests go through the specified proxy server. Then initialize the OpenAI client, use the Python library provided by OpenAI, create an OpenAI client instance for all subsequent API calls, and pass in the API Base URL and API Key when creating the OpenAI client. Finally, after configuring the OpenAI client, perform an interface call test to verify whether it can successfully interact with the CloseAI interface. If it is verified that it cannot successfully interact with the CloseAI interface, you need to reset the interface call until it is verified that it can successfully interact with the CloseAI interface.
[0089] S14, after successfully calling the interface, gradually use a large-scale test set to test the verification code automatic recognition framework, and evaluate the verification code recognition capabilities of the three large language models based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and obtain a mature verification code automatic recognition framework.
[0090] In the specific implementation, after successfully calling the interface, the server needs to gradually use a large-scale test set to test the verification code automatic recognition framework, and evaluate the verification code recognition capabilities of the three large language models based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and obtain a mature verification code automatic recognition framework.
[0091] In one example, the verification code recognition capability of a large language model needs to be considered from three aspects: accuracy, security, and robustness. Correspondingly, the preset performance standards include accuracy performance standards, security performance standards, and robustness performance standards.
[0092] After successfully calling the interface, the server gradually uses the large-scale test set D testb The automatic verification code recognition framework is tested to evaluate the accuracy of three large language models in recognizing verification codes of different modalities.
[0093] Next, we analyze the performance of the three large language models when facing different types of attacks, including noise injection, occlusion, and deformation, and evaluate their ability to resist attacks to determine the security of the three large language models.
[0094] At the same time, the generalization ability of the three large language models is analyzed, and the adaptability of the three large language models to verification codes from different sources and modalities is examined to determine the robustness of the three large language models.
[0095] After completing the evaluation of the verification code recognition capability of the large language model, the server needs to determine whether the accuracy, security, and robustness of the three large language models meet the preset accuracy performance standards, security performance standards, and robustness performance standards respectively.
[0096] If the accuracy, security and robustness of the three large language models meet the accuracy performance standards, security performance standards and robustness performance standards respectively, a mature verification code automatic recognition framework can be obtained.
[0097] If the accuracy, security or robustness of at least one large language model does not meet the accuracy performance standard, security performance standard or robustness performance standard, it is necessary to construct a new large-scale test set and use the new large-scale test set to continue testing the verification code automatic recognition framework until the accuracy, security and robustness of the three large language models meet the accuracy performance standard, security performance standard and robustness performance standard respectively.
[0098] In an example, the accuracy of a large language model can be measured by the single character recognition accuracy and the overall recognition accuracy, which can be calculated by the following formula:
[0099] SCAR=(N s / N t )×100%;
[0100] ASR=(N r / N a )×100%;
[0101] Among them, N s To correctly identify the number of characters, N t is the total number of characters, SCAR represents the single character recognition accuracy, N r To correctly identify the number of samples, N a is the total number of samples, and ASR represents the overall recognition accuracy.
[0102] S15, inputting the verification code to be recognized into a mature verification code automatic recognition framework to obtain a recognition result.
[0103] In a specific implementation, after obtaining a mature verification code automatic recognition framework, the server can be deployed to an application scenario with demand, and the verification code to be recognized is input into the mature verification code automatic recognition framework to obtain a recognition result.
[0104] This embodiment proposes a verification code automatic recognition method based on a large language model, and builds a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. It can not only use chatgpt-4o-latest and gpt-4-turbo provided in the OpenAI platform to automatically recognize verification codes, but also use gemini-1.5-flash provided in the Google platform to automatically recognize verification codes, and successfully implements the call of multiple platforms and multiple models. These large language models can process image information and text information at the same time, improve the accuracy and robustness of verification code recognition, and reduce the dependence on a large amount of annotated data, reducing the cost and complexity of training and maintenance. For the verification code automatic recognition framework, this application conducted two parts of testing. Before building the verification code automatic recognition framework, a small-scale test set was used to conduct preliminary tests on the three large language models, so that the three large language models obtained basic verification code recognition capabilities. The design of the preliminary test shortened the test time and saved test resources. After the verification code automatic recognition framework is built, it is gradually tested using large-scale test sets. The verification code recognition capabilities of the three large language models are evaluated based on the recognition accuracy until the verification code recognition capabilities of the three large language models meet the preset performance standards. In this way, an accurate, secure, robust and mature verification code automatic recognition framework can be obtained, thereby realizing the verification code automatic recognition excellently.
[0105] In one embodiment, in a multi-platform, multi-model, and multi-modal verification code automatic recognition framework, the OpenAI platform accesses and performs verification code recognition through the following steps.
[0106] S21, environment configuration and initialization.
[0107] S211, set up proxy, configure HTTP and HTTPS proxy, point to local proxy server 127.0.0.1:7890.
[0108] S212, initialize the OpenAI client, use the API key and proxy server address to initialize the OpenAI client, so as to subsequently call the OpenAI API interface.
[0109] S22, image preprocessing and encoding.
[0110] S221, image encoding, defines the encode_image function, uses the io.BytesIO() object to store the binary data of the image, converts the image file into a base64-encoded image string, and sends it to the model in the OpenAI platform.
[0111] S222, read the image list from the specified input folder path P in Read all image files and filter out files with png, jpg and jpeg as suffixes.
[0112] S223, batch image encoding, calling the encode_image function for each image file read, converting all image files into base64-encoded image strings, and storing them in the encoded_images list.
[0113] S23, calling the large language model to perform verification code recognition.
[0114] S231, model call, for each base64-encoded image string in the encoded_images list, use the OpenAI client to call chatgpt-4o-latest and gpt-4-turbo for verification code recognition.
[0115] S232, construct a request message, construct a message message including a text prompt and image data, the text prompt instructs the large language model to recognize the text in the image and output the recognition result in a specific format.
[0116] S233, processing the response, for each request, processing the response of the large language model, and extracting the recognized text.
[0117] S234, result recording, recording the path of each image file and the recognized text in the specified output folder path P out middle.
[0118] S235, Error handling and logging, when processing each image, error handling and logging are added to facilitate tracking of processing progress and identification of problems that occur during the process.
[0119] S236, performance evaluation, by analyzing the recognition results, evaluate the accuracy and robustness of ChatGPT-4O-latest and GPT-4-Turbo in recognizing verification codes of different modalities.
[0120] In one embodiment, in a multi-platform, multi-model, and multi-modal verification code automatic recognition framework, the Google platform accesses and performs verification code recognition through the following steps.
[0121] S31, environment configuration, set up the proxy server, and ensure that requests can be sent through the specified proxy server.
[0122] S32, API configuration, use the configure function to configure the Google API key, where the Google API key is the credential for accessing the Google Generative AI service.
[0123] S33, model initialization, initializes the Gemini model instance, and specifies the large language model gemini-1.5-flash.
[0124] S34, image preprocessing, from the specified input folder path P in In the process, all image files are read out according to the test group, and one test group is processed at a time.
[0125] S35, image recognition, traverses each image file and uses the generate_content method of gemini-1.5-flash to identify the verification code. The generate_content method accepts a list containing prompt text and image objects.
[0126] S36, result processing and recording, extracting the recognition result according to the response of gemini-1.5-flash.
[0127] S37, exception handling, during the recognition process, if an exception occurs, capture the exception and record the error information.
[0128] S38, completion notification, after processing all image files, print the completion notification and save the corresponding recognition results to the specified output folder path P out middle.
[0129] It is worth noting that the verification code automatic recognition framework uses a more complex script than the openai-captcha-detection tool, which can process multiple image files. It also adds logging and exception handling for intermediate output to facilitate responding to errors that may occur when processing image files.
[0130] Compared with the openai-captcha-detection tool, which mainly targets specific types of CAPTCHAs, the CAPTCHA automatic recognition framework enhances its adaptability to CAPTCHAs from different sources and styles and improves its generalization ability by applying multiple platforms and multiple models.
[0131] In addition, the verification code automatic recognition framework sets up HTTP and HTTPS proxies and API base URLs to solve network connection problems. It also uses the io.BytesIO() object to store the binary data of image files and then performs base64 encoding. It can handle image files of different formats and has better compatibility.
[0132] The step division of the above methods is only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this application.
[0133] Another embodiment of the present application provides an electronic device, whose specific structure can be as follows: Figure 3 As shown, it includes: at least one processor M41; and a memory M42 that is communicatively connected to the at least one processor M41; wherein the memory M42 stores instructions that can be executed by the at least one processor M41, and the instructions are executed by the at least one processor M41 so that the at least one processor M41 can execute a verification code automatic recognition method based on a large language model as described in the above-mentioned method embodiments.
[0134] The memory and the processor may be connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus may also connect various other circuits such as peripherals, voltage regulators, and power management circuits together, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium.
[0135] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0136] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a verification code automatic recognition method based on a large language model as described in the above method embodiment.
[0137] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM (Read-Only Memory), RAM (Random Access Memory), disk or optical disk and other media that can store program codes.
[0138] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.
Claims
1. A verification code automatic recognition method based on a large language model, characterized in that: include: Obtain several English text verification codes, Chinese text verification codes, image verification codes, and visual reasoning verification codes, and after labeling them, form a small-scale test set and a large-scale test set; A small-scale test set was used to conduct preliminary tests on the three large language models, chatgpt-4o-latest, gpt-4-turbo, and gemini-1.5-flash, so that the three large language models could acquire basic verification code recognition capabilities. Set up interface calls and build a multi-platform, multi-model, and multi-modal verification code automatic recognition framework. The verification code automatic recognition framework supports calling the OpenAI platform and the Google platform through interfaces. The OpenAI platform is equipped with chatgpt-4o-latest and gpt-4-turbo, and the Google platform is equipped with gemini-1.5-flash. After successfully calling the interface, we gradually use large-scale test sets to test the verification code automatic recognition framework, and evaluate the verification code recognition capabilities of the three large language models based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and obtain a mature verification code automatic recognition framework; Input the verification code to be recognized into a mature verification code automatic recognition framework to obtain the recognition result.
2. The method for automatic verification code recognition based on a large language model as claimed in claim 1, characterized in that: Obtain several English text verification codes, Chinese text verification codes, image verification codes, and visual reasoning verification codes, and after labeling them, form a small-scale test set and a large-scale test set, including: Crawl several English text verification codes, several Chinese text verification codes, several image verification codes, and several visual reasoning verification codes from different Internet websites, and obtain English text verification code dataset R, Chinese text verification code dataset C, image verification code dataset T, and visual reasoning verification code dataset S respectively; Mark each verification code with its serial number and the correct recognition result as a label; Randomly select M1 verification codes from each data set to form a small-scale test set D tests ; Randomly select M2 verification codes from each data set to form a large-scale test set D testb , and the large-scale test set D testb Randomly divided into m test groups, Among them, M2>M1.
3. The method for automatic verification code recognition based on a large language model as claimed in claim 2, characterized in that: A small-scale test set was used to conduct preliminary tests on the three large language models, chatgpt-4o-latest, gpt-4-turbo, and gemini-1.5-flash, so that the three large language models can obtain basic verification code recognition capabilities, including: Build a small client compatible with mainstream platforms and use the small-scale test set D tests The verification codes in the code are input into the three large language models chatgpt-4o-latest, gpt-4-turbo and gemini-1.5-flash for recognition; Based on the comparison of the recognition results and corresponding labels output by the three large language models, we can get the recognition results of the three large language models on the small-scale test set D. tests The recognition accuracy rate; If three large language models are trained on a small test set D tests If the recognition accuracy rates of the three language models meet the preset preliminary test standards, the preliminary test of the three language models is completed to determine that the three language models have acquired basic verification code recognition capabilities; If there is at least one large language model for the small test set D tests If the recognition accuracy rate does not meet the preset preliminary test standard, a new small-scale test set is constructed, and the verification code in the new small-scale test set is input into the large language model that does not meet the preset preliminary test standard, until the three large language models all obtain basic verification code recognition capabilities.
4. The method for automatic verification code recognition based on a large language model as claimed in claim 2, characterized in that: Set up the interface call, including: Get API access rights and obtain the API Base URL and API Key from the agency CloseAI. The API Key is the credential for accessing and using OpenAI services, and the API Base URL is the base address of the API server, which is used to construct a complete API request URL. Configure environment variables. Specify the address and port of the proxy server by setting environment variables os.environ["http_proxy"] and os.environ["https_proxy"] to ensure that all API requests go through the specified proxy server. Initialize the OpenAI client and use the Python library provided by OpenAI to create an OpenAI client instance for all subsequent API calls. When creating the OpenAI client, pass in the API Base URL and API Key. After configuring the OpenAI client, perform an interface call test to verify whether it can successfully interact with the CloseAI interface. If it is verified that it cannot successfully interact with the CloseAI interface, reset the interface call until it is verified that it can successfully interact with the CloseAI interface.
5. The method for automatic verification code recognition based on a large language model as claimed in claim 4, characterized in that: In the multi-platform, multi-model, and multi-modal verification code automatic recognition framework, the OpenAI platform accesses and performs verification code recognition through the following steps: S21, environment configuration and initialization; S211, set up proxy, configure HTTP and HTTPS proxy, point to local proxy server 127.0.0.1:7890; S212, initialize the OpenAI client, use the API key and the proxy server address to initialize the OpenAI client, so as to call the OpenAI API interface later; S22, image preprocessing and coding; S221, image encoding, defines the encode_image function, uses the io.BytesIO() object to store the binary data of the image, converts the image file into a base64-encoded image string, and sends it to the model in the OpenAI platform; S222, read the image list from the specified input folder path P in Read all image files and filter out files with png, jpg and jpeg as suffixes; S223, batch image encoding, calling the encode_image function for each image file read, converting all image files into base64-encoded image strings, and storing them in the encoded_images list; S23, calling a large language model to perform verification code recognition; S231, model call, for each base64-encoded image string in the encoded_images list, use the OpenAI client to call chatgpt-4o-latest and gpt-4-turbo for verification code recognition; S232, constructing a request message, constructing a message message including a text prompt and image data, wherein the text prompt instructs the large language model to recognize text in the image and output the recognition result in a specific format; S233, processing the response, for each request, processing the response of the large language model, and extracting the recognized text; S234, result recording, recording the path of each image file and the recognized text in the specified output folder path P out middle; S235, Error handling and logging, when processing each image, error handling and logging are added to facilitate tracking of processing progress and identification of problems that occur during the process; S236, performance evaluation, by analyzing the recognition results, evaluate the accuracy and robustness of ChatGPT-4O-latest and GPT-4-Turbo in recognizing verification codes of different modalities.
6. A verification code automatic recognition method based on a large language model as claimed in claim 5, characterized in that: In the multi-platform, multi-model, and multi-modal verification code automatic recognition framework that has been built, the Google platform accesses and performs verification code recognition through the following steps: S31, environment configuration, setting up a proxy server to ensure that requests can be sent through the specified proxy server; S32, API configuration, using the configure function to configure the Google API key, where the Google API key is the credential for accessing the Google Generative AI service; S33, model initialization, initialize the Gemini model instance, specify the large language model gemini-1.5-flash; S34, image preprocessing, from the specified input folder path P in In the process, all image files are read out according to the test group, and one test group is processed at a time; S35, image recognition, traverse each image file, use gemini-1.5-flash's generate_content method to perform verification code recognition, the generate_content method accepts a list containing prompt text and image objects; S36, result processing and recording, extracting the recognition result according to the response of gemini-1.5-flash; S37, exception handling, during the recognition process, if an exception occurs, capture the exception and record the error information; S38, completion notification, after processing all image files, print the completion notification and save the corresponding recognition results to the specified output folder path P out middle.
7. The method for automatic verification code recognition based on a large language model as claimed in claim 2, characterized in that: The preset performance standards include accuracy performance standards, security performance standards, and robustness performance standards. After successfully calling the interface, the verification code automatic recognition framework is gradually tested using a large-scale test set. The verification code recognition capabilities of the three large language models are evaluated based on the recognition accuracy, until the verification code recognition capabilities of the three large language models meet the preset performance standards, and a mature verification code automatic recognition framework is obtained, including: After successfully calling the interface, we gradually used a large-scale test set to test the automatic verification code recognition framework and evaluate the accuracy of the three large language models in recognizing verification codes of different modalities. Analyze the performance of the three large language models under different types of attacks, including noise injection, occlusion, and deformation, and evaluate their ability to resist attacks to determine the security of the three large language models; Analyze the generalization ability of the three large language models and examine their adaptability to verification codes from different sources and modalities to determine the robustness of the three large language models; Determine whether the accuracy, security, and robustness of the three large language models meet the accuracy performance standard, security performance standard, and robustness performance standard respectively; If the accuracy, security and robustness of the three large language models meet the accuracy performance standard, security performance standard and robustness performance standard respectively, a mature verification code automatic recognition framework is obtained; If the accuracy, security or robustness of at least one large language model does not meet the accuracy performance standard, security performance standard or robustness performance standard, a new large-scale test set is constructed, and the verification code automatic recognition framework is continued to be tested using the new large-scale test set until the accuracy, security and robustness of the three large language models meet the accuracy performance standard, security performance standard and robustness performance standard respectively.
8. The method for automatic verification code recognition based on a large language model as claimed in claim 7, characterized in that: Accuracy is measured by single character recognition accuracy and overall recognition accuracy, which are calculated using the following formula: SCAR=(N s / N t )×100%; <h2 style=";text-align:left;direction:ltr">ASR=(N<h2 style=";text-align:left;direction:ltr"> r <h2 style=";text-align:left;direction:ltr"> / N<h2 style=";text-align:left;direction:ltr"> a <h2 style=";text-align:left;direction:ltr"> )×100%; Among them, N s To correctly identify the number of characters, N t is the total number of characters, SCAR represents the single character recognition accuracy, N r To correctly identify the number of samples, N a is the total number of samples, and ASR represents the overall recognition accuracy.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a verification code automatic recognition method based on a large language model as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a verification code automatic recognition method based on a large language model as described in any one of claims 1 to 8.
Citation Information
Cited By
Verification code management method based on multimodality and reinforcement learning and computer system
CN120654224A
Captcha management method and computer system based on multi-modal and reinforcement learning
CN120654224B