Model training method and device, click verification code identification method and device, equipment and storage medium

Through the collaborative training framework of the target detection model and the click point prediction model, the problems of low click-to-captcha recognition accuracy and high training cost are solved, and efficient and accurate click-to-captcha recognition is achieved, which is suitable for a variety of captcha types and scenarios.

CN120708023APending Publication Date: 2025-09-26CTRIP TRAVEL NETWORK TECH SHANGHAI0
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510787154.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing technology has low accuracy in recognizing verification codes by clicking on them, and the training model requires a large amount of labeled data, which is costly and inefficient to train. It is difficult to process mixed verification codes, and when relying on large models, the accuracy is low and the training time is long.

Method used

A collaborative training framework of the target detection model and the click point prediction model is adopted. By sharing parameters and optimizing the joint loss function, the target detection and click point prediction tasks are integrated to reduce the complexity of data collection and labeling, and improve recognition accuracy and generalization ability.

Benefits of technology

It significantly reduces the cost and complexity of manual labeling, improves the recognition accuracy and efficiency of point-to-point verification codes, adapts to various verification code types, and supports different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708023A_ABST
    Figure CN120708023A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, a click verification code recognition method and device, equipment and a storage medium, a target detection model is trained based on a target detection data set, the target detection data set comprises a plurality of click verification code pictures and corresponding object labeling information, a click point prediction model is trained based on a click point prediction data set, and a click verification code recognition model is trained based on the click point prediction model. The click point prediction data set comprises a source object sub-picture and a target object sub-picture, the source object sub-picture and the target object sub-picture are extracted from a source region and a target region in the click verification code picture respectively, and the target detection model and the click point prediction model share click verification code picture resources in a training stage. Repeated work of data acquisition and labeling is reduced, and training efficiency is improved through cooperative training. In an identification stage, through a cooperation mechanism of target detection and click point prediction, click point position information of an object to be clicked in an original point selected verification code picture is accurately positioned, and reliability and accuracy of verification of the verification code are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and storage medium for model training and point-to-point verification code recognition. Background Art

[0002] With the increasing demand for information security, various types of verification codes have become widely used in human-machine authentication. Click verification codes, a typical image-based verification code, require users to identify the location of the clickable object in the image and click on it to complete the authentication process. Compared to traditional text-based verification codes, click verification codes are more interactive and secure, but they also increase the complexity of automated recognition and reduce recognition accuracy.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0004] In response to the problems in the prior art, the object of the present invention is to provide a model training, point-to-point verification code recognition method, device, equipment and storage medium, which overcomes the difficulties of the prior art and is used to solve the technical problem of low accuracy in point-to-point verification code recognition in the prior art.

[0005] A first aspect of the present disclosure provides a training method for a click verification code recognition model, comprising:

[0006] Training a target detection model based on a target detection dataset, the target detection dataset comprising a plurality of click verification code images and corresponding object annotation information, the object annotation information corresponding to each of the click verification code images comprising annotated object categories and object location information, the object categories comprising source objects and target objects, the source objects and target objects respectively corresponding to source regions and target regions in the click verification code images, the object location information comprising location information of each of the source objects and target objects in the click verification code images, the target detection model being configured to input the click verification code images and output detected object categories and object location information;

[0007] Training a click point prediction model based on a click point prediction dataset, the click point prediction dataset including a source object sub-image and a target object sub-image, the source object sub-image and the target object sub-image being extracted from the source region and the target region of the click verification code image, respectively, the click point prediction model being configured to input the click point prediction dataset and output a target object matching the source object as an object to be clicked;

[0008] Based on the object position information detected by the target detection model, the click point position information of the object to be clicked in the click verification code picture is obtained, so as to perform a simulated click operation on the terminal based on the click point position information.

[0009] Optionally, when the click verification code image includes multiple source objects, the click verification code image includes multiple objects to be clicked that match the multiple source objects respectively, and the object annotation information also includes a click order for the multiple objects to be clicked, so as to perform simulated click operations on the multiple objects to be clicked in sequence based on the click order.

[0010] Optionally, the target detection model is a YOLO model.

[0011] Optionally, the click point prediction model is a deep neural network with an encoder-decoder structure.

[0012] Optionally, the training method of the click verification code recognition model further includes:

[0013] Before training a click point prediction model based on a click point prediction dataset, positive samples are extracted from a pre-generated first folder, and negative samples are extracted from a pre-generated second folder. The positive samples and negative samples constitute the click point prediction dataset. The positive samples include the source object sub-picture and the matching target object sub-picture, and the negative samples include the source object sub-picture and the unmatched target object sub-picture. The positive samples and negative samples corresponding to the same source object sub-picture have the same file name.

[0014] Optionally, the target detection model and the click point prediction model constitute an end-to-end network, which takes the click verification code image as input and the object to be clicked as output; the end-to-end network jointly trains the target detection model and the click point prediction model based on a joint loss function, and the joint loss function is composed of a weighted fusion of the target detection loss function and the click point prediction loss function.

[0015] Optionally, the target detection model and the click point prediction model share at least part of the feature extraction layer.

[0016] A second aspect of the present disclosure provides a method for identifying a click verification code, comprising:

[0017] Receive and click on the verification code picture;

[0018] Inputting the click verification code image into a trained object detection model and outputting object categories and object location information, wherein the object categories include source objects and target objects, the source objects and target objects respectively corresponding to the source area and target area in the click verification code image, and the object location information includes the location information of each of the source objects and target objects in the click verification code image;

[0019] Segmenting the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area respectively;

[0020] Inputting the source object sub-image and the target object sub-image into a trained click point prediction model, and outputting a target object matching the source object as an object to be clicked;

[0021] Based on the object position information, the click point position information of the object to be clicked in the click verification code picture is obtained, so as to perform a simulated click operation on the terminal based on the click point position information.

[0022] Optionally, the target detection model is a YOLO model.

[0023] Optionally, the click point prediction model is a deep neural network with an encoder-decoder structure.

[0024] Optionally, the click point prediction model is a twin network.

[0025] Optionally, obtaining the click point position information of the object to be clicked in the click verification code image based on the object position information includes:

[0026] Locating a bounding box in the click verification code image based on the object position information, and selecting an area with the largest pixel value in the bounding box as a primary click area;

[0027] In the main click area, a center point or a maximum response point is sampled as at least one click point coordinate as the click point position information.

[0028] Optionally, the target detection model and the click point prediction model constitute an end-to-end network, and the end-to-end network takes the click verification code image as input and the object to be clicked as output.

[0029] Optionally, the object detection model and the click point prediction model share at least part of the feature extraction layer.

[0030] Optionally, when there are multiple source objects, the target detection model further outputs a click order; and performing a simulated click operation on the terminal based on the click point position information includes:

[0031] Based on the click order, simulated click operations are sequentially performed on the multiple objects to be clicked that respectively match the multiple source objects.

[0032] A third aspect of the present disclosure provides a training device for a click verification code recognition model, comprising:

[0033] A first training module trains a target detection model based on a target detection dataset, wherein the target detection dataset includes a plurality of click verification code images and corresponding object annotation information, wherein the object annotation information corresponding to each of the click verification code images includes an annotated object category and object location information, wherein the object category includes a source object and a target object, wherein the source object and the target object correspond to a source area and a target area in the click verification code image, respectively, and the object location information includes location information of each of the source object and the target object in the click verification code image, and the target detection model is configured to input the click verification code image and output the detected object category and object location information;

[0034] a second training module, training a click point prediction model based on a click point prediction dataset, wherein the click point prediction dataset includes a source object sub-image and a target object sub-image, wherein the source object sub-image and the target object sub-image are extracted from the source region and the target region of the click verification code image, respectively; and the click point prediction model is configured to input the click point prediction dataset and output a target object matching the source object as an object to be clicked;

[0035] The positioning module obtains the click point position information of the object to be clicked in the click verification code image based on the object position information detected by the target detection model, so as to perform a simulated click operation on the terminal based on the click point position information.

[0036] A fourth aspect of the present disclosure provides a device for identifying a click verification code, comprising:

[0037] Receiving module, receiving and clicking verification code pictures;

[0038] A first prediction module inputs the click verification code image into a trained object detection model and outputs object categories and object location information, wherein the object categories include source objects and target objects, the source objects and target objects respectively corresponding to source areas and target areas in the click verification code image, and the object location information includes location information of each of the source objects and target objects in the click verification code image;

[0039] a segmentation module, which segments the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area respectively;

[0040] A second prediction module inputs the source object sub-image and the target object sub-image into a trained click point prediction model, and outputs a target object that matches the source object as the object to be clicked;

[0041] The simulation operation module obtains the click point position information of the object to be clicked in the click verification code picture based on the object position information, and performs a simulation click operation on the terminal based on the click point position information.

[0042] A third aspect of the present disclosure provides an electronic device, comprising:

[0043] processor;

[0044] a memory storing executable instructions for the processor;

[0045] The processor is configured to execute the executable instructions to perform the steps of the method for training a click verification code recognition model or the method for clicking verification code recognition described in any of the above embodiments.

[0046] A fourth aspect of the present disclosure provides a computer-readable storage medium for storing a program, which, when executed, implements the training method of the click verification code recognition model or the steps of the click verification code recognition method described in any embodiment.

[0047] The model training and verification code recognition method, device, equipment, and storage medium provided by the embodiments of the present disclosure have the following advantages:

[0048] During the training phase, the target detection model and the click point prediction model share the resources of the click verification code image, reducing the duplication of data collection and annotation, significantly reducing the complexity and cost of manual annotation, and improving training efficiency through collaborative training. During the recognition phase, the collaborative mechanism of target detection and click point prediction is used to segment the object based on its location information and accurately locate the click point location information of the clicked object in the original click verification code image, ensuring the reliability and accuracy of the verification code verification. This implementation can support multiple click verification code types to adapt to different application scenarios.

[0049] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Other features, objectives and advantages of the present invention will become more apparent upon reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0051] Figure 1 A flowchart showing a method for training a click verification code recognition model provided by an embodiment of the present disclosure;

[0052] Figure 2-Figure 6 Schematic diagram showing different ways to click on verification code images;

[0053] Figure 7 One of the flow charts of a method for identifying a click verification code provided by an embodiment of the present disclosure is shown;

[0054] Figure 8 A flowchart showing a method for training a target detection model provided by an embodiment of the present disclosure;

[0055] Figure 9 A flowchart showing a method for training a click point prediction model provided by an embodiment of the present disclosure;

[0056] Figure 10 A second flowchart showing a method for identifying a click verification code provided by an embodiment of the present disclosure is shown;

[0057] Figure 11 A schematic diagram showing the module structure of a training device for a click verification code recognition model provided by an embodiment of the present disclosure;

[0058] Figure 12 A schematic diagram showing the module structure of a click verification code recognition device provided by an embodiment of the present disclosure is shown;

[0059] Figure 13 A schematic diagram showing the structure of an electronic device provided by an embodiment of the present disclosure;

[0060] Figure 14 It is a schematic diagram of the structure of a computer program product according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0061] In order to make the technical solution of the present invention clearer and more specific, the exemplary embodiments of the present invention are now described in detail with reference to the accompanying drawings. It should be understood that the following embodiments are only used to illustrate the technical implementation path of the present invention, rather than to limit the scope of protection of the present invention. Those skilled in the art should be able to make various modifications and adjustments based on the basic ideas disclosed in the present invention, and these modifications should also be considered to be included in the scope of protection of the present invention.

[0062] It should be noted that the specific implementations provided in this specification can be implemented in various forms, and the examples given should not be construed as limiting the present invention in any way. The structural features, module divisions, and processing steps in the implementations can be flexibly combined and adjusted according to the system architecture, deployment platform, and operating environment to better adapt to the needs of the target application scenario.

[0063] In addition, the accompanying drawings are used to illustrate the principle structure and implementation steps of the present invention and are not necessarily drawn to physical scale. To simplify the description of the drawings, the same reference numerals in the drawings represent the same or functionally similar components or modules, so these same elements are not repeatedly described in multiple figures. The functional entities or processing flows shown in the drawings can be implemented in a variety of ways and are not limited to a specific hardware structure or logical division. The specific implementation can be achieved using software functions, API interfaces, script logic, microservice processes or hardware modules, and can also be completed on a single computing node or in a distributed system.

[0064] Among related technologies, CAPTCHA recognition technology is based on image classification and target detection methods, but it faces many challenges. For example, the objects to be clicked in CAPTCHAs often have diverse categories and complex interference elements, and the granularity of click point annotation far exceeds the object location information annotation requirements required for traditional target detection, which places higher demands on the model's accuracy and generalization capabilities. At the same time, existing methods often train target detection and click point prediction as independent tasks, which has the following shortcomings:

[0065] (1) Training the model requires a large amount of labeled data, and the labeling cost is high;

[0066] (2) The separation of target detection and click point prediction leads to low training efficiency and long recognition time;

[0067] (3) Lack of versatility, making it difficult to process mixed verification codes with text and icons;

[0068] (4) Relying on large models to recognize distorted text or complex backgrounds has low accuracy, is time-consuming, costly, and has limited access.

[0069] To solve the above problems, the disclosed embodiments propose an efficient and accurate model training and click verification code recognition solution, and provide an efficient joint training framework that can integrate target detection and click point prediction tasks, effectively integrate the information flows of the two, and improve the overall recognition accuracy and generalization ability through shared parameters and joint loss optimization strategy.

[0070] The present disclosure provides a method for training a click verification code recognition model. Figure 1 The following steps are shown as an example, including but not limited to:

[0071] Step 110: Training an object detection model based on an object detection dataset, wherein the object detection dataset includes a plurality of click verification code images and corresponding object annotation information, wherein the object annotation information corresponding to each of the click verification code images includes an annotated object category and object location information, wherein the object category includes a source object and a target object, wherein the source object and the target object correspond to a source region and a target region in the click verification code image, respectively, and the object location information includes location information of each of the source object and the target object in the click verification code image, wherein the object detection model is configured to input the click verification code image and output the detected object category and object location information;

[0072] Step 120: Training a click point prediction model based on a click point prediction dataset, wherein the click point prediction dataset includes a source object sub-image and a target object sub-image, wherein the source object sub-image and the target object sub-image are extracted from the source region and the target region of the click verification code image, respectively. The click point prediction model is configured to input the click point prediction dataset and output a target object that matches the source object as the object to be clicked;

[0073] Step 140: Based on the object position information detected by the target detection model, obtain the click point position information of the object to be clicked in the click verification code image, so as to perform a simulated click operation on the terminal based on the click point position information.

[0074] In this embodiment, the target detection model identifies the object category and object location information of the source object and the target object, thereby enabling regional classification and positioning of the object in the click verification code image. The subsequent click point prediction model uses sub-images extracted from the source and target areas to further accurately match the source and target objects, accurately identify the object to be clicked that matches the source object, and further predict the click point location information of the object to be clicked in the click verification code image. This combination of step-by-step positioning (region positioning + click point positioning) effectively reduces the misrecognition rate, improves the accuracy of the click position, and meets the high-precision requirements of the click verification code.

[0075] Based on this collaborative mechanism, the object detection dataset and the click point prediction dataset share the resources of the click verification code image, reducing the duplication of data collection. The regional localization results (object location information) of the object detection model provide a constraint range for the click point prediction model. The click point prediction dataset only needs to be segmented based on the object location information within this range to construct the click point prediction dataset, significantly reducing the complexity and cost of manual annotation. Simultaneously, the object detection model and the click point prediction model are trained separately for specific tasks, optimizing resource allocation and improving training efficiency.

[0076] The trained object detection model can handle a variety of object categories, while the click prediction model supports flexible click prediction through sub-image matching. This modular design allows the model to adapt to different types of click verification codes (such as text or icons), improving the method's versatility and adaptability to diverse verification scenarios.

[0077] In the disclosed embodiment, the click verification code is a human-machine verification mechanism that distinguishes humans from automated programs by requiring users to click a specified object (such as "apple") in the click verification code image. Figure 2 As shown, it displays "Please click "hand", "low", and "public" in sequence", and displays "hand", "low", and "public" waiting to be clicked objects and interference objects such as "coarse", "sign", and "shock" in the verification code picture, and each text is displayed in a different form. Only when the user clicks the correct objects to be clicked in sequence can the authentication be complete and correct.

[0078] In this embodiment, when the click verification code image includes multiple source objects, the click verification code image includes multiple objects to be clicked that match the multiple source objects respectively, and the object annotation information also includes a click order for the multiple objects to be clicked, so as to perform simulated click operations on the multiple objects to be clicked in sequence based on the click order.

[0079] In another embodiment, the object annotation information may not include the click order of the multiple objects to be clicked. In this case, the simulated click operation may be randomly performed on the multiple objects to be clicked.

[0080] In an embodiment of the present disclosure, the target detection dataset includes a plurality of click verification code images. Figure 2 The picture shows a text-based verification code. Figure 3-Figure 5 There are many kinds of click verification code pictures with different backgrounds, fonts and sizes. Figure 6 An icon class shown clicks on a verification code image, displays multiple icons, and instructs the user to click the correct icon objects in the image in sequence.

[0081] In the embodiment of the present disclosure, professional tools can be used to mark the object category and object location information for each click verification code picture. Figure 2 As shown in the figure, the object class table is divided into source and target. Source represents the corresponding object in the source area, while target represents the object in the target area. The object location information can be defined by a bounding box, and the location information of each bounding box is recorded to represent the object location information. The bounding box can be expressed in the format of (x_min, y_min, x_max, y_max), defining the spatial extent of each object.

[0082] The click point prediction dataset consists of sub-images obtained by segmenting each object in the verification code image based on the object location information, such as Figure 2 The object position information is segmented to generate a source object sub-image and a corresponding target object sub-image.

[0083] In one embodiment, the target object sub-image in the click point prediction dataset can be annotated with similarity matching information with the source object sub-image. If a match is found, it is marked as a match; otherwise, it is marked as a mismatch, thereby achieving supervised training of the click point prediction model.

[0084] Among them, the source object sub-image and the matching target object sub-image constitute positive samples, while the source object sub-image and the unmatched target object sub-image constitute negative samples. In this way, the click point prediction dataset constitutes a positive and negative sample set.

[0085] Considering that each source object sub-image corresponds to only one matching target object sub-image, but may also correspond to many unmatched target object sub-images, resulting in a large difference in the amount of data between positive and negative samples, we can replicate the positive samples multiple times to balance the difference in data size between them and negative samples.

[0086] In one embodiment, before training a click point prediction model based on a click point prediction dataset, positive samples are extracted from a pre-generated first folder and negative samples are extracted from a pre-generated second folder. The positive samples and negative samples constitute the click point prediction dataset. The positive samples include the source object sub-picture and the matching target object sub-picture, and the negative samples include the source object sub-picture and the unmatched target object sub-picture. The first folder and the second folder corresponding to the same source object sub-picture have the same file name.

[0087] At this time, the positive and negative samples are pre-generated. By naming the positive and negative samples corresponding to the same source object sub-image the same file name, the corresponding positive and negative samples can be extracted immediately, improving training efficiency.

[0088] In addition, the name of each sub-image can be mapped to a corresponding position in the original point selection verification code image, so that the position of each sub-image in the original point selection verification code image can be located.

[0089] In this embodiment, the target detection dataset and the click point prediction dataset can share image resources, form a collaborative mechanism, and reduce data collection costs.

[0090] In an embodiment of the present disclosure, before training the target detection model, the verification code image can be preprocessed, including normalizing the image size and applying data enhancement techniques (such as random flipping and brightness adjustment) to improve the robustness of the model.

[0091] As described above, the object position information may be a spatial range defined by a bounding box, and the click point position information may be any valid position within the bounding box.

[0092] As an optional embodiment, the target detection model can be a deep neural network, such as the YOLO v10 model; the click point prediction model is a deep neural network with an encoder-decoder structure, where the encoder extracts features and the decoder generates outputs.

[0093] As an implementation, the click point prediction model can use a Siamese network, which consists of two or more sub-networks with shared parameters. Each sub-network receives a source object sub-image or a target object sub-image as input, uses a ResNet-50 encoder to extract features, maps the input to a low-dimensional feature space, and uses a loss function to measure the similarity or difference between the two inputs.

[0094] In the embodiment of the present disclosure, the process of training the target detection model includes:

[0095] Model initialization, setting the network structure and initial parameters, and accelerating convergence based on pre-trained weights;

[0096] Data input: input the click verification code image of the target detection dataset into the model;

[0097] Feature extraction, generating multi-scale feature maps through convolutional layers to capture features of objects of different sizes;

[0098] Output the predicted object location information and object category, and use the target detection loss function (such as classification loss and positioning loss) to measure the difference between the predicted object location information and object category and the labeled value;

[0099] Parameter optimization, adjust parameters through optimizers (such as gradient descent), and iterate training until convergence.

[0100] In the disclosed embodiment, the process of training the click point prediction model includes:

[0101] Input the click point prediction dataset into the model to be trained;

[0102] Extract feature information of each sub-image through the encoder;

[0103] The decoder outputs the target object sub-image that matches the source object sub-image as a similarity matching result, specifically a similarity probability map;

[0104] The click point prediction loss function is used to measure the difference between the similarity matching result and the actual matching degree. The optimizer adjusts the parameters and iterative training is used to improve the accuracy.

[0105] In the disclosed embodiment, an object detection model and a click prediction model form an end-to-end network that takes a verification code image as input and outputs the object to be clicked. This end-to-end network jointly trains the two models based on a joint loss function, which is a weighted fusion of the object detection loss function and the click prediction loss function. By optimizing the weighted balance objectives, the performance of both models is simultaneously improved, avoiding the iterative overhead of training each model individually. The weighted mechanism balances the coarse positioning (object detection) and fine positioning (click prediction) tasks, ensuring stable and efficient training.

[0106] Optionally, the object detection model and the click prediction model share at least some feature extraction layers, reducing computational overhead by reusing parameters. Combining feature extraction layer sharing with joint training enables an efficient joint training framework, integrating information flows and improving recognition accuracy and generalization. During the joint training phase, training is terminated when the joint loss function converges below a set threshold for several consecutive rounds, or when there is no significant improvement in object detection accuracy or click prediction accuracy on the validation set.

[0107] During joint training, the object detection model and the click prediction model share the convolutional weights of the feature extraction layer, reducing parameter redundancy and optimizing training efficiency. This shared weighting reduces computational complexity, and the joint loss function ensures collaborative optimization of the two tasks, reducing costs and improving collaborative performance. The final CAPTCHA recognition model, which includes both the object detection model and the click prediction model, is suitable for the recognition and automatic execution of various CAPTCHAs.

[0108] The disclosed embodiment also provides a method for identifying click verification codes, which aims to complete user click verification in an automated manner, supporting network security verification and automated testing scenarios. The target detection model and click point prediction model can be used Figure 1 Train using the training method shown, or use other training methods.

[0109] The execution subject of this point-to-point verification code recognition method is the user terminal and / or the backend server. Figure 7 As shown, the method for clicking the verification code to identify the code includes but is not limited to the following steps:

[0110] Step 710: Receive a verification code image click;

[0111] Step 720: Input the click verification code image into a trained object detection model, and output object categories and object location information. The object categories include source objects and target objects, and the source objects and target objects correspond to source areas and target areas in the click verification code image, respectively. The object location information includes location information of each of the source objects and target objects in the click verification code image.

[0112] Step 730: Segment the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area respectively;

[0113] Step 740: Input the source object sub-image and the target object sub-image into a trained click point prediction model, and output a target object that matches the source object as the object to be clicked;

[0114] Step 750: Based on the object position information, obtain the click point position information of the object to be clicked in the click verification code image, so as to perform a simulated click operation on the terminal based on the click point position information.

[0115] In this implementation, the target detection model quickly locates the source and target objects, segments the original click verification code image based on object category and location information, extracts the source and target object sub-images, and narrows the processing scope of the click point prediction model. The click point prediction model focuses on similarity matching to determine the object to be clicked. Through the coordinated mechanism of target detection and click point prediction, the click point location information is precisely located to ensure the reliability and accuracy of verification. This implementation supports multiple click verification code types (such as text and icon types) to adapt to different scenarios.

[0116] Therefore, this embodiment can improve the verification efficiency and accuracy of clicking on the verification code.

[0117] In one embodiment, a verification code image to be processed is received and input data is provided for recognition, ensuring efficient acquisition and standardized processing, and supporting real-time verification.

[0118] Click on the verification code image for reference Figure 2-Figure 6 There are multiple types shown (text or icon), and the types are not limited.

[0119] The receiving process includes image acquisition and preprocessing. JPEG or PNG format images can be obtained from verification scenarios (such as website login pages) through the API. In the application scenario, when the user opens the terminal interface, the backend server pushes the verification code image and prompts the user to click the correct object, triggering the automatic recognition mechanism.

[0120] Preprocessing includes image size normalization, pixel value normalization to the range of [0,1], and noise removal and / or contrast enhancement to adapt to complex backgrounds.

[0121] In the disclosed embodiment, the trained object detection model is used to identify the object category and object location information, which are expressed in the format of (x_min, y_min, x_max, y_max) to provide coarse positioning for click point prediction.

[0122] The object detection model is a deep neural network that extracts multi-scale features through convolutional layers, predicts object categories and object location information, and outputs results containing multiple object location information and corresponding categories.

[0123] Therefore, the model distinguishes the source and target areas through multi-scale feature extraction, identifies the object position information of each object, achieves accurate classification and coarse positioning, and optimizes recognition efficiency.

[0124] As an implementation method, the target detection model may be a YOLO v10 model.

[0125] In this disclosed embodiment, the original click verification code image is segmented based on the object position information output by the object detection model, extracting the source object sub-image and the target object sub-image. Each sub-image is cropped in the (x_min, y_min, x_max, y_max) format based on the object position information, reducing the processing range, eliminating background interference, and improving the efficiency of click point prediction.

[0126] In the disclosed embodiment, the click point prediction model performs similarity matching among multiple target object sub-pictures based on each source object sub-picture, determines the object to be clicked with the highest similarity, and achieves accurate matching.

[0127] The click point prediction model is a deep neural network with an encoder-decoder structure, and a twin network can be used.

[0128] Optionally, the primary click area is located based on the object's position information. The area with the largest pixel value within the primary click area is selected, and the center point or the point with the largest response within this area is sampled as the click point coordinates. Object position information constrains the click range and provides fine-grained location distribution. The area with the largest pixel value ensures coordinate accuracy, while the center point is suitable for uniform distribution and the point with the largest response is suitable for concentrated distribution, improving flexibility.

[0129] In the disclosed embodiment, the click point coordinates are output as click point location information, which is used for terminal simulation click to complete automated verification.

[0130] Simulated clicks use automated scripts (such as Selenium) to simulate user click behavior and complete verification code verification.

[0131] The output process involves formatting the coordinates into a standard JSON format (e.g., {"x":150, "y":200}) and passing it to the automated system via an API. The script then parses the coordinates and executes the click on the terminal (e.g., a browser or mobile device), supporting multi-platform deployment.

[0132] This implementation achieves seamless automated verification through standardized coordinate output, simplifies interface design compared to related technologies, ensures efficiency through script execution, and is suitable for scenarios with high request volumes.

[0133] In an optional embodiment of the present disclosure, when there are multiple source objects, the object detection model further outputs a click order; and performing a simulated click operation on the terminal based on the click point position information includes:

[0134] Based on the click order, simulated click operations are sequentially performed on the multiple objects to be clicked that respectively match the multiple source objects.

[0135] In an optional embodiment of the present disclosure, the target detection model and the click point prediction model constitute an end-to-end network, and the end-to-end network takes the click verification code image as input and the object to be clicked as output.

[0136] In another embodiment, the object detection model and the click point prediction model share at least part of the feature extraction layer.

[0137] like Figure 8 As shown, the embodiment of the present disclosure also provides a flowchart of a training method for a click verification code recognition model, which specifically includes the following steps.

[0138] First, prepare the target detection dataset and train the YOLO v10 model (corresponding to the target detection model above), which includes the following steps:

[0139] Step 810: Download multiple verification code images and define two object categories: source and target (corresponding to the source object and target object above) to reduce the number of labeled samples. Use LabelImg software to label the object category and object location information for each verification code image. The object location information is recorded in the format of (x_min, y_min, x_max, y_max), for example Figure 2 The object position information of the "hand" object is (50, 20, 100, 50).

[0140] Step 820: Set parameters and train: Configure YOLO v10 parameters, including the number of iterations, image input size, pre-loaded ImageNet weights, and object categories, and perform training to generate a preliminary model YOLOv10_click_select.pt.

[0141] Step 830: Convert the preliminary model YOLO v10_click_select.pt to the general format yolov10_click_select.onnx as the target detection model.

[0142] Step 840: Load the yolov10_click_select.onnx model, obtain the weight parameter matrix value, input the above multiple original click verification code images into YOLO v10 in sequence, and output the object category and object location information in each click verification code image;

[0143] Step 850: Segment the objects in the click verification code image based on the object position information, divide them into source object sub-images and target object sub-images according to the object category and save them to the source and target folders. The sub-image naming rule includes the name of the original click verification code image, such as "image1_source_1.jpg".

[0144] Steps 810-850 are used to train the object detection model and further generate a click point prediction dataset consisting of source object sub-images and target object sub-images. The following describes in detail the training method of the click point prediction model.

[0145] Further, if Figure 9 As shown, prepare the click point prediction dataset and train the Siamese model, which includes the following steps:

[0146] Step 910: Obtain positive and negative sample sets. Traverse each sub-image in the source and target folders, call the large model queen-vl-plus to identify and match, and manually verify, saving the results to the matched (positive samples, such as "hand" and "hand") (corresponding to the first folder above) and unmatched (negative matching samples, such as "hand" and "thick") folders (corresponding to the second folder above).

[0147] During this process, the number of positive and negative samples is balanced. Since each source object sub-image corresponds to one matching target sub-image, but may correspond to multiple non-matching target sub-images, the amount of positive and negative sample data varies greatly. Positive samples are replicated to match the number of negative samples.

[0148] Step 920: Reconstruct the data loading logic and improve the original data loading logic (each file has ≥3 similar pictures, and 1 picture is taken from different files to form a sample). To adapt to the scenario with only 2 similar pictures, it is changed to randomly taking corresponding pictures from the matched and unmatched folders each time to form positive and negative samples.

[0149] Step 930: Set parameters and train. Configure the Siamese network parameters, use ResNet-50 as the encoder, train the generated model siamese_click_select.pth, output the similarity probability map, and convert the model to the universal format siamese_click_select.onnx.

[0150] Before training the Siamese model, extract positive samples from the pre-generated first folder (matched) and negative samples from the second folder (unmatched). The file names are consistent to ensure the correspondence.

[0151] Using the target detection model and click point prediction model trained by the above method, a specific click verification code recognition solution is implemented, as follows: Figure 10 As shown, the following steps are included.

[0152] Step 1010: Obtain the click verification code image through the API to support real-time verification requirements.

[0153] Step 1020: Input the click verification code image into the trained YOLO v10 model (yolov10_click_select.onnx), and output the object category (source and target) and object location information;

[0154] Step 1030: Segment the click verification code image to obtain a source object sub-image and a target object sub-image;

[0155] Step 1040: Pair each sub-image in the source set with a sub-image in the target set, input the Siamese model (siamese_click_select.onnx), and output a similarity probability map. The image with the highest similarity is the object to be clicked. The source set is sorted from left to right to ensure sequential processing of multi-object scenes. Based on the Siamese model, a similarity probability map is obtained between the source object sub-image and the target object sub-image, and the matching object to be clicked is determined based on the similarity probability map.

[0156] Step 1050: Locate and select the main click area in the verification code image based on the object position information, sample the center point or the maximum response point in the main click area, and generate the click point coordinates as the click point position information.

[0157] Step 1060: Execute simulated click. Format the click point location information into JSON coordinates (e.g., {"x":150, "y":200}), pass it to an automated script (e.g., Selenium) via an API, and simulate a click on the terminal to complete verification.

[0158] This implementation utilizes a combined YOLO v10 and Siamese deep learning training model to effectively identify different types of verification codes and meet the needs of high-request business scenarios.

[0159] Figure 11 Schematic diagram of a module of an embodiment of a training device for a click verification code recognition model provided by an embodiment of the present disclosure. Figure 11 As shown, the training device 1100 of the click verification code recognition model disclosed herein includes but is not limited to:

[0160] A first training module 1110 trains a target detection model based on a target detection dataset, wherein the target detection dataset includes a plurality of click verification code images and corresponding object annotation information, wherein the object annotation information corresponding to each of the click verification code images includes an annotated object category and object location information, wherein the object category includes a source object and a target object, wherein the source object and the target object correspond to a source region and a target region in the click verification code image, respectively, and the object location information includes location information of each of the source object and the target object in the click verification code image, wherein the target detection model is configured to input the click verification code image and output the detected object category and object location information;

[0161] A second training module 1120 trains a click point prediction model based on a click point prediction dataset, wherein the click point prediction dataset includes a source object sub-image and a target object sub-image, wherein the source object sub-image and the target object sub-image are extracted from the source region and the target region of the click verification code image, respectively. The click point prediction model is configured to input the click point prediction dataset and output a target object that matches the source object as an object to be clicked;

[0162] The positioning module 1130 obtains the click point position information of the object to be clicked in the click verification code image based on the object position information detected by the target detection model, so as to perform a simulated click operation on the terminal based on the click point position information.

[0163] In an optional embodiment, before training the click point prediction model based on the click point prediction dataset, the second training module 1120 is further configured to:

[0164] Positive samples are extracted from a pre-generated first folder, and negative samples are extracted from a pre-generated second folder. The positive samples and negative samples constitute the click point prediction dataset. The positive samples include the source object sub-picture and the matching target object sub-picture, and the negative samples include the source object sub-picture and the unmatched target object sub-picture. The positive samples and negative samples corresponding to the same source object sub-picture have the same file name.

[0165] The training device of the above-mentioned click verification code recognition model provides a collaborative training mechanism for the target detection model and the click point prediction model. The target detection dataset and the click point prediction dataset share the click verification code image resources, reducing the duplication of data collection and annotation. The regional positioning result (object location information) of the target detection model provides a constraint range for the click point prediction model. It is only necessary to intercept the source object sub-image and the target object sub-image in the click verification code image within this range to construct the click point prediction dataset, which significantly reduces the complexity and cost of manual annotation. At the same time, the target detection model and the click point prediction model are trained for specific tasks respectively to optimize resource allocation and improve training efficiency.

[0166] The modules of this device can be interconnected through a data interface to collaboratively implement the training of the click verification code recognition model. The device can be embedded in the front-end development tool as an independent software module, or it can be used as a cloud service to provide an interface for developers to call remotely.

[0167] An embodiment of the present invention further provides a training device for a click-through verification code recognition model, comprising a processor and a memory storing executable instructions for the processor. The processor is configured to execute the executable instructions to perform the steps of a method for training a click-through verification code recognition model.

[0168] As shown above, the training device of the click verification code recognition model of the present invention can provide a collaborative training mechanism for the target detection model and the click point prediction model. The target detection data set and the click point prediction data set share the click verification code image resources, reducing the duplication of data collection and annotation. The regional positioning result (object location information) of the target detection model provides a constraint range for the click point prediction model. It is only necessary to intercept the source object sub-image and the target object sub-image in the click verification code image within this range to construct the click point prediction data set, which significantly reduces the complexity and cost of manual annotation. At the same time, the target detection model and the click point prediction model are trained for specific tasks respectively, optimizing resource allocation and improving training efficiency.

[0169] Figure 12 This is a module diagram of an embodiment of the point-to-point verification code recognition device provided by the present disclosure. Figure 12 As shown, the click verification code recognition device 1200 of the present disclosure includes but is not limited to:

[0170] Receiving module 1210 receives a verification code image;

[0171] The first prediction module 1220 inputs the click verification code image into a trained object detection model and outputs object categories and object location information, wherein the object categories include source objects and target objects, the source objects and target objects respectively corresponding to the source area and target area in the click verification code image, and the object location information includes the location information of each of the source objects and target objects in the click verification code image;

[0172] A segmentation module 1230 segments the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area, respectively;

[0173] The second prediction module 1240 inputs the source object sub-image and the target object sub-image into a trained click point prediction model, and outputs a target object that matches the source object as the object to be clicked;

[0174] The simulation operation module 1250 obtains the click point position information of the object to be clicked in the click verification code image based on the object position information, and performs a simulation click operation on the terminal based on the click point position information.

[0175] In an optional embodiment, when there are multiple source objects, the target detection model further outputs the click order; the simulation operation module 1250 is specifically used to:

[0176] Based on the click order, simulated click operations are sequentially performed on a plurality of objects to be clicked that respectively match a plurality of source objects.

[0177] This click verification code recognition device quickly locates the source object and target object through the target detection model, and then segments the original click verification code image to obtain the source object sub-image and multiple target object sub-images. This can extract and narrow the processing range of the subsequent click point prediction model, so that the click point prediction model can focus on image similarity matching to match the object to be clicked. In this way, through the collaborative mechanism of target detection and click point prediction, the click point position information of the object to be clicked in the original click verification code image is accurately located, ensuring the reliability and accuracy of the verification code verification. This embodiment can support multiple click verification code types and adapt to different application scenarios.

[0178] The modules of this device can be interconnected through a data interface to collaboratively realize the point-to-point verification code recognition device. The device can be embedded in the front-end development tool as an independent software module, or it can be used as a cloud service to provide an interface for developers to call remotely.

[0179] An embodiment of the present invention further provides a device for identifying a verification code by clicking on it, comprising a processor and a memory storing executable instructions of the processor. The processor is configured to execute the steps of the verification code identification method by executing the executable instructions.

[0180] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "platforms."

[0181] Figure 13 This is a schematic diagram of the structure of the electronic device of the present invention. Figure 13 An electronic device 1300 according to this embodiment of the present invention will be described. Figure 13 The electronic device 1300 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0182] like Figure 13 As shown, electronic device 1300 is implemented as a general-purpose computing device. Components of electronic device 1300 may include, but are not limited to, at least one processing unit 1310, at least one storage unit 1320, a bus 1330 connecting various platform components (including storage unit 1320 and processing unit 1310), and a display unit 1340.

[0183] The storage unit stores program codes, which can be executed by the processing unit 1310, so that the processing unit 1310 executes the steps according to various exemplary embodiments of the present invention described in the training method of the click verification code recognition model or the click verification code recognition method section of this specification. For example, the processing unit 1310 can execute the following steps: Figure 1 or Figure 7 Follow the steps shown in .

[0184] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1321 and / or a cache memory unit 1322 , and may further include a read-only memory unit (ROM) 1323 .

[0185] The storage unit 1320 may also include a program / utility 1324 having a set (at least one) of program modules 1325, such program modules 1325 including but not limited to: a processing system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0186] Bus 1330 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0187] The electronic device 1300 can also communicate with one or more external devices 1301 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1300, and / or any device that enables the electronic device 1300 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 1350. Furthermore, the electronic device 1300 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1360. The network adapter 1360 can communicate with other modules of the electronic device 1300 via a bus 1330. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 1300, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0188] The present disclosure also provides a computer-readable storage medium for storing a program that implements the training method for a click-to-verify code recognition model or the steps of the click-to-verify code recognition method when the program is executed. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned training method for a click-to-verify code recognition model or the click-to-verify code recognition method section of this specification.

[0189] like Figure 14 As shown, a computer program product 1400 for implementing the above method according to an embodiment of the present invention can be implemented in a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the computer program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0190] The computer program product can employ any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0191] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0192] The program code for performing the processes of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0193] In summary, the object of the present invention is to provide a method, apparatus, device and storage medium for model training and point-to-point verification code recognition. The target detection model and the click point prediction model share the click-to-point verification code image resources during the training phase, reducing the duplication of data collection and annotation, significantly reducing the complexity and cost of manual annotation, and improving training efficiency through collaborative training. In the recognition phase, through the collaborative mechanism of target detection and click point prediction, the click point position information of the object to be clicked in the original point-to-point verification code image is accurately located to ensure the reliability and accuracy of verification code verification. This embodiment can support multiple types of click-to-point verification codes and adapt to different application scenarios.

[0194] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A training method for a click verification code recognition model, characterized in that: include: Training a target detection model based on a target detection dataset, the target detection dataset comprising a plurality of click verification code images and corresponding object annotation information, the object annotation information corresponding to each of the click verification code images comprising annotated object categories and object location information, the object categories comprising source objects and target objects, the source objects and target objects respectively corresponding to source regions and target regions in the click verification code images, the object location information comprising location information of each of the source objects and target objects in the click verification code images, the target detection model being configured to input the click verification code images and output detected object categories and object location information; Training a click point prediction model based on a click point prediction dataset, the click point prediction dataset including a source object sub-image and a target object sub-image, the source object sub-image and the target object sub-image being extracted from the source region and the target region of the click verification code image, respectively, the click point prediction model being configured to input the click point prediction dataset and output a target object matching the source object as an object to be clicked; Based on the object position information detected by the target detection model, the click point position information of the object to be clicked in the click verification code picture is obtained, so as to perform a simulated click operation on the terminal based on the click point position information.

2. The training method of the click verification code recognition model according to claim 1, characterized in that: When the click verification code image includes multiple source objects, the click verification code image includes multiple objects to be clicked that match the multiple source objects respectively, and the object annotation information also includes a click order for the multiple objects to be clicked, so as to perform simulated click operations on the multiple objects to be clicked in sequence based on the click order.

3. The training method of the click verification code recognition model according to claim 1, characterized in that: The target detection model is the YOLO model.

4. The training method of the click verification code recognition model according to claim 1, characterized in that: The click point prediction model is a deep neural network with an encoder-decoder structure.

5. The training method of the click verification code recognition model according to claim 1, characterized in that: The training method of the click verification code recognition model also includes: Before training a click point prediction model based on a click point prediction dataset, positive samples are extracted from a pre-generated first folder, and negative samples are extracted from a pre-generated second folder. The positive samples and negative samples constitute the click point prediction dataset. The positive samples include the source object sub-picture and the matching target object sub-picture, and the negative samples include the source object sub-picture and the unmatched target object sub-picture. The positive samples and negative samples corresponding to the same source object sub-picture have the same file name.

6. The method for training a click verification code recognition model according to claim 1, wherein: The target detection model and the click point prediction model constitute an end-to-end network, which takes the click verification code image as input and the object to be clicked as output; the end-to-end network jointly trains the target detection model and the click point prediction model based on a joint loss function, and the joint loss function is composed of a weighted fusion of the target detection loss function and the click point prediction loss function.

7. The training method of the click verification code recognition model according to claim 1, characterized in that: The object detection model and the click point prediction model share at least part of the feature extraction layer.

8. A method for identifying a click verification code, characterized in that: include: Receive and click on the verification code picture; Inputting the click verification code image into a trained object detection model and outputting object categories and object location information, wherein the object categories include source objects and target objects, the source objects and target objects respectively corresponding to the source area and target area in the click verification code image, and the object location information includes the location information of each of the source objects and target objects in the click verification code image; Segmenting the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area respectively; Inputting the source object sub-image and the target object sub-image into a trained click point prediction model, and outputting a target object matching the source object as an object to be clicked; Based on the object position information, the click point position information of the object to be clicked in the click verification code picture is obtained, so as to perform a simulated click operation on the terminal based on the click point position information.

9. The method for identifying a click verification code according to claim 8, wherein: The target detection model is the YOLO model.

10. The method for identifying a click verification code according to claim 8, wherein: The click point prediction model is a deep neural network with an encoder-decoder structure.

11. The method for identifying a click verification code according to claim 10, wherein: The click point prediction model is a twin network.

12. The method for identifying a click verification code according to claim 8, wherein: The obtaining, based on the object position information, the click point position information of the object to be clicked in the click verification code image includes: Locating a bounding box in the click verification code image based on the object position information, and selecting an area with the largest pixel value in the bounding box as a primary click area; In the main click area, a center point or a maximum response point is sampled as at least one click point coordinate as the click point position information.

13. The method for identifying a click verification code according to claim 8, wherein: The target detection model and the click point prediction model constitute an end-to-end network, and the end-to-end network takes the click verification code image as input and the object to be clicked as output.

14. The method for identifying a click verification code according to claim 8, wherein: The object detection model and the click point prediction model share at least part of the feature extraction layer.

15. The method for identifying a click verification code according to claim 8, wherein: When there are multiple source objects, the object detection model further outputs a click order; The performing a simulated click operation on the terminal based on the click point position information includes: Based on the click order, simulated click operations are sequentially performed on the multiple objects to be clicked that respectively match the multiple source objects.

16. A training device for a click verification code recognition model, characterized in that: include: A first training module trains a target detection model based on a target detection dataset, wherein the target detection dataset includes a plurality of click verification code images and corresponding object annotation information, wherein the object annotation information corresponding to each of the click verification code images includes an annotated object category and object location information, wherein the object category includes a source object and a target object, wherein the source object and the target object correspond to a source area and a target area in the click verification code image, respectively, and the object location information includes location information of each of the source object and the target object in the click verification code image, and the target detection model is configured to input the click verification code image and output the detected object category and object location information; a second training module, training a click point prediction model based on a click point prediction dataset, wherein the click point prediction dataset is composed of a source object sub-image and a target object sub-image, wherein the source object sub-image and the target object sub-image are extracted from the source region and the target region of the click verification code image, respectively; and the click point prediction model is configured to input the click point prediction dataset and output a target object matching the source object as an object to be clicked; The positioning module obtains the click point position information of the object to be clicked in the click verification code image based on the object position information detected by the target detection model, so as to perform a simulated click operation on the terminal based on the click point position information.

17. A device for identifying a click verification code, characterized in that: include: Receiving module, receiving and clicking verification code pictures; A first prediction module inputs the click verification code image into a trained object detection model and outputs object categories and object location information, wherein the object categories include source objects and target objects, the source objects and target objects respectively corresponding to source areas and target areas in the click verification code image, and the object location information includes location information of each of the source objects and target objects in the click verification code image; a segmentation module, which segments the click verification code image based on the object category and object location information to obtain a source object sub-image and a target object sub-image corresponding to the source area and the target area respectively; A second prediction module inputs the source object sub-image and the target object sub-image into a trained click point prediction model, and outputs a target object that matches the source object as the object to be clicked; The simulation operation module obtains the click point position information of the object to be clicked in the click verification code picture based on the object position information, and performs a simulation click operation on the terminal based on the click point position information.

18. An electronic device, characterized in that: include: processor; a memory storing executable instructions for the processor; The processor is configured to execute the steps of the method for training a click verification code recognition model according to any one of claims 1 to 7, or the steps of the method for clicking a verification code recognition according to any one of claims 8 to 15, by executing the executable instructions.

19. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the training method of the click verification code recognition model described in any one of claims 1 to 7 or the steps of the click verification code recognition method described in any one of claims 8 to 15 are implemented.