Security verification method and device, storage medium and terminal

By cropping the original image and generating misleading images, and by leveraging the differences between the human visual system and the machine recognition system, the problem of automated programs bypassing security verification is solved, thus achieving effective verification of user identity.

CN120850271APending Publication Date: 2025-10-28ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511028212.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, automated programs can easily bypass security verification strategies, making network attacks difficult to defend against, especially with the development of artificial intelligence technology, which makes attack methods more intelligent and covert.

Method used

By cropping the original image, a misleading image is generated that matches the preset similarity to the sub-image. This misleading image is then displayed in the interactive interface, instructing the user to select the correct sub-image. The user's identity is verified by utilizing the difference between the human visual system and the machine recognition system.

Benefits of technology

It effectively distinguishes between real users and automated programs, increasing the difficulty and accuracy of security verification and enhancing the system's security and resistance to attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850271A_ABST
    Figure CN120850271A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a security verification method and device, a storage medium and a terminal. Obtaining an original image in response to a security verification request triggered by user operation; the original image is cut to obtain at least one sub-image, at least one misleading image corresponding to each sub-image is obtained, and each sub-image and each misleading image meet a preset similarity condition; displaying the cut original image, each sub-image and each misleading image in an interactive interface, and indicating a user to select a sub-image corresponding to the cut original image from each sub-image and each misleading image; and verifying the identity security of the user based on the selection result of the user. An original image is cut to obtain a plurality of sub-images, and a plurality of misleading images similar to the sub-images are generated according to the sub-images, so that the images are displayed to a user, and the user is indicated to select correct sub-images corresponding to the original image. And the identity security of the current user can be effectively verified according to the selection result of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a security verification method, apparatus, storage medium, and terminal. Background Technology

[0002] In today's internet environment, various cyberattacks are rampant, among which the most common type involves attackers using bots to simulate human user actions. The core of this type of attack lies in the bot's ability to bypass security checks and carry out malicious activities by mimicking human behavior. The impact of such attacks on systems is profound, easily leading to data breaches, functional impairments, service interruptions, and system resource exhaustion. Furthermore, with the development of artificial intelligence technology, attack methods will become more intelligent and covert. Therefore, a comprehensive security verification solution is urgently needed to address the significant challenges posed by attacks using bots simulating human users. Summary of the Invention

[0003] This specification provides a security verification method, apparatus, storage medium, and terminal, which can solve the technical problem in related technologies that security verification strategies are easily bypassed by automated attacks.

[0004] Firstly, embodiments of this specification provide a security verification method, the method comprising: In response to a security verification request triggered by a user action, obtain the original image; The original image is cropped to obtain at least one sub-image, and at least one misleading image corresponding to each sub-image is obtained. Each sub-image and each misleading image satisfy a preset similarity condition. The interactive interface displays the cropped original image, each sub-image, and each misleading image, instructing the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image. The identity security of the users is verified based on their selection results.

[0005] In one possible implementation, obtaining at least one misleading image corresponding to each sub-image includes: inputting each sub-image into an image generation model, determining at least one misleading image corresponding to each sub-image generated by the image generation model; the image generation model is obtained by training a basic generation model using sample images.

[0006] In one possible implementation, the method further includes: periodically collecting images from real-world scenes and / or generating images using an image generation model, and storing all collected images in an image library; obtaining the original image includes: randomly selecting an image from the image library as the original image.

[0007] In one possible implementation, cropping the original image to obtain at least one sub-image includes: randomly selecting at least one cropping position in the original image, and cropping the original image according to each cropping position to obtain at least one sub-image.

[0008] In one possible implementation, the sub-images and misleading images are arranged and displayed in a random order.

[0009] In one possible implementation, the above-mentioned verification of the user's identity security based on the user's selection result includes: if the user's selection result is correct, then determining that the user's identity is secure; if the user's selection result is incorrect, then determining that the user's identity is at risk, and triggering an advanced verification process for the user.

[0010] Secondly, embodiments of this specification provide a security verification device, which includes: The process triggering module is used to obtain the original image in response to a security verification request triggered by a user operation; The image acquisition module is used to crop the original image to obtain at least one sub-image, and to acquire at least one misleading image corresponding to each sub-image, wherein each sub-image and each misleading image satisfy a preset similarity condition. The image display module is used to display the cropped original image, each sub-image, and each misleading image in the interactive interface, and to instruct the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image; The security verification module is used to verify the identity security of the users based on their selection results.

[0011] In one possible implementation, the image acquisition module is further configured to input each sub-image into the image generation model and determine at least one misleading image generated by the image generation model corresponding to each sub-image; the image generation model is obtained by training the basic generation model using sample images.

[0012] In one possible implementation, the security verification device further includes: an image library management module, used to periodically collect images from real-world scenes and / or generate images through an image generation model, and store all collected images in an image library; the process triggering module is also used to randomly select an image from the image library as the original image.

[0013] In one possible implementation, the image acquisition module is further configured to randomly select at least one cropping position in the original image and crop the original image according to each cropping position to obtain at least one sub-image.

[0014] In one possible implementation, the sub-images and misleading images are arranged and displayed in a random order.

[0015] In one possible implementation, the security verification module is further configured to determine that the user's identity is secure if the user's selection result is correct, and to determine that the user's identity is at risk if the user's selection result is incorrect, and to trigger an advanced verification process for the user.

[0016] Thirdly, embodiments of this specification provide a computer program product containing instructions that, when run on a computer or processor, cause the computer or processor to perform the steps of the method described above.

[0017] Fourthly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the steps of the method described above.

[0018] Fifthly, embodiments of this specification provide a terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is adapted to be loaded by the processor and to execute the steps of the method described above.

[0019] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following: This specification provides a security verification method that, in response to a security verification request triggered by a user operation, acquires an original image; crops the original image to obtain at least one sub-image; acquires at least one misleading image corresponding to each sub-image; and ensures that each sub-image and each misleading image satisfy a preset similarity condition. The cropped original image, each sub-image, and each misleading image are displayed in an interactive interface, instructing the user to select the sub-image corresponding to the cropped original image from among the sub-images and misleading images. The user's identity security is verified based on the user's selection. In this embodiment, when a user triggers a security verification request, an original image is cropped to obtain multiple sub-images, and multiple misleading images similar to the sub-images are generated. When these images are displayed to the user, instructing them to select the correct sub-image corresponding to the original image, the user's recognition ability can be verified based on their selection. If the process is automated, it is more likely to be misled by similar features in the images; if the user observes with their naked eye, they are more likely to directly select the correct sub-image, thus effectively verifying the current user's identity security. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 An exemplary system architecture diagram of a security verification method provided in the embodiments of this specification; Figure 2 A flowchart illustrating a security verification method provided in an embodiment of this specification; Figure 3 An example of the image display effect of a security verification method provided in the embodiments of this specification; Figure 4 A flowchart illustrating a security verification method provided in an embodiment of this specification; Figure 5 A structural block diagram of a security verification device provided in the embodiments of this specification; Figure 6 This is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. Detailed Implementation

[0022] To make the features and advantages of the embodiments of this specification more apparent and understandable, the technical solutions of the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the embodiments of this specification.

[0023] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those in this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments in this specification as detailed in the appended claims. Furthermore, in the description of the embodiments in this specification, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, in the description of the embodiments in this specification, "multiple" refers to two or more.

[0024] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0025] In today's internet environment, various cyberattacks are rampant, among which the use of bots to simulate human user actions is particularly common. The core of these attacks lies in the attackers' use of various methods to build automated programs that mimic human behavior to bypass security detection. Basic attack types in "mimicry attacks" include malicious registration, spam delivery, brute-force account attacks, DDoS attacks, data scraping, traffic manipulation, and web application attacks. Brute-force attacks use automated tools to try numerous account and password combinations, while DDoS attacks paralyze services through a large number of requests. With the development of artificial intelligence technology, attack methods have become more intelligent and covert. For example, mouse movement trajectory simulation can use machine learning to analyze human operating habits and generate trajectory data with random pauses and misoperations, significantly improving the CAPTCHA cracking rate. Session hijacking and deepfake techniques use browser automation frameworks (such as Selenium) to simulate the entire login process, and can even bypass behavioral analysis detection through UI automation. Attackers use these methods to gain illicit profits or damage system functions, posing a serious challenge to the security and stability of online platforms.

[0026] Currently, several common defense methods exist in the field of security verification. These include image verification, which requires users to select a valid image from multiple images. However, this approach is vulnerable to enumeration; attackers can annotate all images and then use the annotated results to pass verification during subsequent attacks. Another method combines image and text verification, such as requiring users to select specific Chinese characters in an image. However, with the increasing sophistication of image text recognition technology, attackers can develop scripts that automatically recognize and select text in images to bypass current security verification strategies.

[0027] Therefore, this specification provides a security verification method to solve the aforementioned technical problem that security verification strategies are easily bypassed by automated attacks.

[0028] Please see Figure 1 , Figure 1 This is an exemplary system architecture diagram of a security verification method provided in the embodiments of this specification.

[0029] like Figure 1As shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 serves as the medium for providing a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired or wireless communication links, such as wired communication links including fiber optic cables, twisted-pair cables, or coaxial cables, and wireless communication links including Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.

[0030] Terminal 101 can interact with server 103 via network 102 to receive messages from or send messages to server 103. Alternatively, terminal 101 can interact with server 103 via network 102 to receive messages or data sent to server 103 by other users. Terminal 101 can be hardware or software. When terminal 101 is hardware, it can be various electronic devices, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal 101 is software, it can be installed in the aforementioned electronic devices and can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module; no specific limitation is made here.

[0031] In the embodiments of this specification, terminal 101 first responds to a security verification request triggered by a user operation and obtains the original image. At this time, terminal 101 crops the original image to obtain at least one sub-image and obtains at least one misleading image corresponding to each sub-image. Each sub-image and each misleading image satisfy a preset similarity condition. Further, terminal 101 displays the cropped original image, each sub-image, and each misleading image in the interactive interface, instructing the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image. Finally, terminal 101 can verify the user's identity security based on the user's selection result.

[0032] Server 103 can be a business server providing various services. It should be noted that server 103 can be hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 103 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module; no specific limitations are made here.

[0033] Alternatively, the system architecture may not include server 103. In other words, server 103 may be an optional device in the embodiments of this specification. That is, the method provided in the embodiments of this specification can be applied to a system structure that only includes terminal 101. The embodiments of this specification do not limit this.

[0034] It should be understood that Figure 1 The number of terminals, networks, and servers shown is only illustrative; the number can be any number of terminals, networks, and servers depending on the implementation requirements.

[0035] Please see Figure 2 , Figure 2 This is a flowchart illustrating a security verification method provided in an embodiment of this specification. The execution entity in this embodiment can be a terminal performing security verification, a processor within the terminal performing the security verification method, or a security verification service within the terminal performing the security verification method. For ease of description, the following example uses a processor within a terminal as the execution entity to illustrate the specific execution process of the security verification method.

[0036] like Figure 2 As shown, a security verification method may include at least: S202. In response to a security verification request triggered by a user action, obtain the original image.

[0037] Optionally, to accurately intercept all risks, the system can trigger a security verification process before responding to certain user operations. By verifying the user's identity, the security of information in the system is ensured. Specifically, the system can monitor user behavior characteristics to promptly know the user's current operation and respond to the security verification process triggered by the user's operation. This allows the system to instruct the user to complete the security verification step in the interactive interface to achieve the target operation. Monitoring user operations can be implemented through a listener. The monitored user operations can be diverse, including but not limited to data download and upload, file access and reading, payment and withdrawal, etc. The target user operation that needs to trigger the security verification process can be determined according to the security verification requirements in the actual scenario. This specification does not specifically limit such user operations in the embodiments.

[0038] Optionally, when an image is cropped, a real user, when identifying a sub-image belonging to the cropped original image, often relies on visual observation of the relevance of the image content and the rationality of the image edge connections. An automated program, however, characterizes the image features by analyzing the pixel arrangement of both the original and sub-images, and then evaluates the correlation between them. In other words, a real user can observe the relationship between the original and sub-images through the visual representation of the image, but an automated program can only distinguish them through implicit image features. Therefore, the automated program's ability to distinguish between the original and sub-images is inferior to that of a real user. Based on this, in the embodiments of this specification, after cropping the original image, multiple confusing sub-image options are provided based on the cropped sub-images. The user is instructed to select the correct sub-image corresponding to the original image from all options. The user's selection result can determine the user's ability to recognize the image, thereby effectively distinguishing whether the user's action was triggered by a real user or an automated program.

[0039] Furthermore, the security verification process first requires acquiring the original image. The original image can be extracted from a pre-prepared image library or generated in real-time using a generative model. The type of the original image is not limited; it can include landscapes, people, animals, buildings, and various other subjects. Please refer to [link / reference]. Figure 3 , Figure 3 This is an example of the image display effect of a security verification method provided in the embodiments of this specification. As shown in Figure 3, the largest image at the top is the original image, and its type is a landscape image.

[0040] S204. Cropping the original image to obtain at least one sub-image, and obtaining at least one misleading image corresponding to each sub-image, wherein each sub-image and each misleading image satisfy a preset similarity condition.

[0041] Optionally, cropping the original image can yield at least one sub-image; see further. Figure 3 Taking the original image as an example of cropping it twice, two sub-images are obtained respectively. When the original image is cropped multiple times, the sub-images can be of the same size or different sizes.

[0042] Furthermore, for sub-images, multiple images need to be prepared as obfuscated options to be displayed to the user. For example... Figure 3 As shown, at least one misleading image corresponding to each sub-image can be obtained, which satisfies the preset similarity condition with the sub-image (taking 4 misleading images as an example). Since these misleading images are similar to but not exactly the same as the correct sub-image, the similarity between the misleading images and the sub-images can easily interfere with the judgment of the automated program, but real users can accurately distinguish all images.

[0043] In one feasible implementation, misleading images can be obtained by searching an image library from the same source as the original image, or by selecting from other pre-prepared image libraries. Alternatively, a neural network can be used to analyze the pixel features of the original image to generate misleading images similar to the sub-images. The goal of generating misleading images is to ensure that each misleading image reaches a set high similarity threshold with the original image or each sub-image at the level of automated program perception. In other words, the core of achieving human-distinguishable and machine-indistinguishable results lies in utilizing the perceptual differences between the human visual system (HVS) and machine recognition models.

[0044] Specifically, human vision prioritizes semantics, focusing on high-level features such as semantic information, structural outlines, and object categories, while being less sensitive to details, noise, and high-frequency texture changes. For example, in an image, simply changing the grass color from dark green to light green allows humans to recognize that both colors originate from the same scene; similarly, applying noise reduction to the edges of leaves does not affect the overall semantic understanding of the image. Machine learning models, however, rely on pixel features, specifically low-level features like edges, textures, and pixel distribution, or intermediate features extracted by pre-trained models. Therefore, even minor pixel perturbations can significantly alter the model's feature vector distance; preserving edge structure while changing semantic content can lead to misclassification of semantically different images as the same scene. Thus, pre-defined similarity conditions should be designed to suit the image analysis characteristics of automated programs, specifically by considering both semantic and pixel features. From a semantic perspective, the semantic similarity between the misleading image and each sub-image should not exceed a first threshold. This threshold can be obtained through sample training, and it should ensure that both the success rate for human identification and the failure rate for machine identification meet expectations. On the other hand, the pixel feature similarity between the misleading image and each sub-image needs to exceed a second threshold, so that when a machine identifies an image based on pixel features, it is easy to mistake the misleading image for the original image. In this way, the final misleading image that meets the preset similarity conditions can be quickly identified by humans through semantic information, while machines may be confused due to low-level features such as similar edges and colors, thus achieving the desired misleading effect.

[0045] S206. Display the cropped original image, each sub-image, and each misleading image in the interactive interface, and instruct the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image.

[0046] Please continue reading. Figure 3Once all images are prepared, the original cropped image, each sub-image, and each misleading image are displayed in the interactive interface. The user is then instructed to select the sub-image corresponding to the original cropped image from among the sub-images and misleading images. The sub-images and misleading images can be arranged and displayed in a random order within the interactive interface. This randomness ensures that the image arrangement presented to the user each time is accidental, increasing system security and making it difficult for attackers to predict and imitate.

[0047] Optionally, when displaying each sub-image and each misleading image, further processing such as random rotation, scaling, and perspective transformation can be applied to each sub-image and each misleading image. For example, rotating sub-image A by 15°, enlarging sub-image B by 1.3 times, etc., increases the mapping complexity between the sub-image and the original image, making it difficult for attackers to reverse locate the source of the sub-image through simple geometric rules.

[0048] In addition, the verification time can be adjusted according to the risk level of the user's operation. For example, the verification time for high-risk operations can be set to 10 seconds, and the verification time for low-risk operations can be set to 30 seconds. When instructing users to perform the corresponding verification, in addition to informing users of the verification rules in text, audio can also be played to explain the rules to users. Furthermore, in addition to supporting users to select operations with a keyboard and mouse, users can also complete the verification by voice commands (such as "select the third picture").

[0049] Another feasible implementation is to provide customizable verification difficulty options for different user groups. Based on the user's data profile, various verification processes with different difficulties can be developed. The appropriate verification process can be automatically used for each user group, or users can choose the verification difficulty themselves, thereby improving the user's verification experience in the system.

[0050] S208. Verify user identity security based on user selection results.

[0051] Optionally, after a user makes a selection, the system can further verify the user's identity security based on the selection result. It can also provide the user with immediate feedback on the verification result and offer additional assistance or simpler verification options if the user makes multiple incorrect selections.

[0052] Specifically, after generating all misleading images and determining the front-end layout displayed to the user, all sub-images and misleading images can be pre-encoded. Each code is unique among all current options. For example, when there are 6 options, the images for each option are sequentially encoded as "01", "02", up to "06". During encoding, the codes of the correct sub-images are stored as a correct set for use during verification. When the user selects all sub-images they believe belong to the original image from a randomly arranged set of sub-images through interactive operations such as clicking or checking, the set of codes for the selected images is recorded. The user's selection set is compared with the pre-determined correct set, requiring that the user's selection set must be completely consistent with the correct set. That is, the user must select all correct sub-images and cannot include any misleading images. If the two sets match exactly, the authentication is successful; otherwise, the authentication fails. By using a precise matching strategy of selecting all and only, the coherence of human understanding of the overall semantics of images is effectively utilized, raising the attack threshold for automated programs. Even if an automated program selects some of the correct sub-images, it is difficult to know that it has accurately hit all the correct sub-images, and it is very easy to mistakenly select misleading images.

[0053] In this embodiment, a security verification method is provided. In response to a security verification request triggered by a user operation, an original image is acquired; the original image is cropped to obtain at least one sub-image; at least one misleading image corresponding to each sub-image is acquired; each sub-image and each misleading image satisfy a preset similarity condition; the cropped original image, each sub-image, and each misleading image are displayed in an interactive interface, instructing the user to select the sub-image corresponding to the cropped original image from among the sub-images and misleading images; and the user's identity security is verified based on the user's selection. In this embodiment, when a user triggers a security verification request, an original image is cropped to obtain multiple sub-images, and multiple misleading images similar to the sub-images are generated. When these images are displayed to the user, instructing the user to select the correct sub-image corresponding to the original image, the user's recognition ability can be verified based on the user's selection. If it is an automated program, it is more likely to be misled by similar features in the images; if the user observes with their naked eye, it is easier to directly select the correct sub-image, thereby effectively verifying the current user's identity security.

[0054] Please see Figure 4 , Figure 4 This is a flowchart illustrating a security verification method provided in an embodiment of this specification.

[0055] like Figure 4 As shown, a security verification method may include at least: S402. In response to a security verification request triggered by a user operation, randomly select an image from the image library as the original image.

[0056] Optionally, to directly authenticate users, images from real-world scenarios and / or images generated using image generation models can be collected periodically. All collected images are stored in an image library. During verification, a single original image is randomly selected from the library. This random selection ensures the randomness of each verification process, thereby enhancing system security. The original image contains sufficient detail and rich visual elements for effective cropping. The images in the library can be a fusion of various image data sources, including mixed sampling of images of natural landscapes, everyday objects, art collections, animals, and buildings. Regularly updating the image library also ensures its richness, increases diversity, and prevents attackers from learning specific patterns within the library.

[0057] S404. Randomly select at least one cropping position in the original image, and crop the original image according to each cropping position to obtain at least one sub-image.

[0058] Optionally, when cropping sub-images from the original image, a preset program can be used to randomly select at least one cropping position in the original image, and crop the original image according to each cropping position to obtain at least one sub-image. The randomness of the cropping position ensures that each generated verification code is different, thus improving security.

[0059] In one feasible implementation, randomly cropping sub-images from the original image can be achieved through a dynamic fragmentation algorithm. Specifically, the number of randomly generated sub-images is first determined for the original image, typically at least one, usually 3-6. To ensure the reliability of verification, the number of random sub-images can be selected based on the total number of all required sub-images. Ideally, the number of correct random sub-images in the original image should not exceed 50% of all options. Besides determining the number, the size range also needs to be determined. When selecting the size, the focus is on whether it can contain identifiable and effective features, avoiding situations where sub-images of the original image lack effective features, making it difficult for users to distinguish between correct and incorrect sub-images. Furthermore, the positional distribution of the sub-images needs to be considered to ensure that the fragments cover key or random areas of the original image. Based on the cropping parameters, at least one original sub-image is cropped from the original image non-overlapping, making each original sub-image a unique local view of the original image. It should be noted that when cropping the original sub-images, considering the size of the original image and the amount of key information, a small amount of overlap between the sub-images is permissible in practical applications, as long as the overlap is not excessive enough to cause two sub-images to be obviously identical. This dynamic generation of cropping parameters for the original image enhances the randomness of the sub-images from different original images, making each generated fragment combination unique and greatly increasing the difficulty for attackers to predict.

[0060] S406. Input each sub-image into the image generation model and determine at least one misleading image generated by the image generation model corresponding to each sub-image; the image generation model is obtained by training the basic generation model using sample images.

[0061] Furthermore, when generating misleading images based on sub-images, a large-scale artificial intelligence model can be used. Specifically, a large-scale basic generation model is pre-trained using a large number of sample images, enabling it to output images with the expected similarity for each input image. After convergence, the basic generation model yields the image generation model in this embodiment. When deployed in a real-world application scenario, each sub-image can be input into the pre-trained model, which then generates at least one misleading image corresponding to each sub-image. The size of the misleading image can be the same as or different from the sub-image. The image generation model ensures that the generated small image has a certain visual similarity to the cropped small image, but also sufficient differences to test the user's recognition ability. By introducing random cropping and generating similar images, the CAPTCHA pattern becomes highly dynamic and complex, increasing the difficulty of cracking automated programs.

[0062] In one feasible implementation, the image generation model is also regularly updated and optimized using validation feedback from real-world scenarios to ensure the quality and diversity of the generated images, thereby preventing the model from generating easily recognizable image patterns.

[0063] S408. Display the cropped original image, each sub-image, and each misleading image in the interactive interface, and instruct the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image.

[0064] For details regarding step S408, please refer to the description in step S206; it will not be repeated here.

[0065] S410. If the user's selection result is correct, the user's identity is determined to be secure; if the user's selection result is incorrect, the user's identity is determined to be at risk, and an advanced verification process is triggered for the user.

[0066] As described in the above embodiments, if the user selects the correct sub-image, it can be confirmed that the current user is a real user, meaning the user's identity is secure, and further, a series of user operations can be supported. Conversely, if the user makes an incorrect selection, it indicates that the current user's identity is at risk, and in this case, additional assistance and advanced verification processes can be provided to the user.

[0067] In the embodiments of this specification, a security verification method is provided. In response to a security verification request triggered by a user operation, an image is randomly selected from an image library as the original image; at least one cropping position is randomly selected in the original image, and the original image is cropped according to each cropping position to obtain at least one sub-image; each sub-image is input into an image generation model to determine at least one misleading image corresponding to each sub-image generated by the image generation model; the image generation model is obtained by training a basic generation model using sample images; the cropped original image, each sub-image, and each misleading image are displayed in an interactive interface, instructing the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image; if the user's selection result is correct, the user's identity is determined to be secure; if the user's selection result is incorrect, the user's identity is determined to be at risk, and an advanced verification process is triggered for the user. This specification's embodiments randomly crop small images from the original image in real time. Randomization of the position ensures the uniqueness and unpredictability of each CAPTCHA generation. Through multi-layered randomness and dynamic generation during cropping, misleading image generation, and image arrangement, the complexity and diversity of the CAPTCHA are further increased, improving its security and reducing the risk of cracking. In practical applications, it also possesses flexible scalability, easily introducing new types of images or image generation algorithms to adapt to new security needs or user preferences, further enhancing the system's security and resistance to attacks.

[0068] Please see Figure 5 , Figure 5 This is a structural block diagram of a security verification device provided in an embodiment of this specification. Figure 5 As shown, the security verification device 500 includes: The process triggering module 510 is used to obtain the original image in response to a security verification request triggered by a user operation; The image acquisition module 520 is used to crop the original image to obtain at least one sub-image, and to acquire at least one misleading image corresponding to each sub-image, wherein each sub-image and each misleading image satisfy a preset similarity condition. The image display module 530 is used to display the cropped original image, each sub-image, and each misleading image in the interactive interface, and to instruct the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image; The security verification module 540 is used to verify the user's identity security based on the user's selection results.

[0069] Optionally, the image acquisition module 520 is further configured to input each sub-image into the image generation model and determine at least one misleading image generated by the image generation model corresponding to each sub-image; the image generation model is obtained by training the basic generation model using sample images.

[0070] Optionally, the security verification device 500 further includes: an image library management module, used to periodically collect images from real-world scenarios and / or generate images through an image generation model, and store all collected images in an image library; and a process triggering module 510, used to randomly select an image from the image library as the original image.

[0071] Optionally, the image acquisition module 520 is further configured to randomly select at least one cropping position in the original image and crop the original image according to each cropping position to obtain at least one sub-image.

[0072] Optionally, the sub-images and misleading images are arranged and displayed in a random order.

[0073] Optionally, the security verification module 540 is also used to determine that the user's identity is secure if the user's selection result is correct, and to determine that the user's identity is at risk if the user's selection result is incorrect, and to trigger an advanced verification process for the user.

[0074] In this embodiment, a security verification device is provided, comprising: a process triggering module for acquiring an original image in response to a security verification request triggered by a user operation; an image acquisition module for cropping the original image to obtain at least one sub-image and acquiring at least one misleading image corresponding to each sub-image, wherein each sub-image and each misleading image satisfy a preset similarity condition; an image display module for displaying the cropped original image, each sub-image, and each misleading image in an interactive interface, instructing the user to select the sub-image corresponding to the cropped original image from the sub-images and misleading images; and a security verification module for verifying the user's identity security based on the user's selection result. In this embodiment, when a user triggers a security verification request, an original image is cropped to obtain multiple sub-images, and multiple misleading images similar to the sub-images are generated based on the sub-images. When these images are displayed to the user, instructing the user to select the correct sub-image corresponding to the original image, the user's recognition ability can be verified based on the user's selection result. If it is an automated program, it is easier to be misled by similar features in the images; if the user observes with the naked eye, it is easier to directly select the correct sub-image, thereby effectively verifying the current user's identity security.

[0075] This specification provides a computer program product containing instructions that, when run on a computer or processor, cause the computer or processor to perform the steps of any of the methods described above.

[0076] This specification also provides a computer storage medium that can store multiple instructions adapted for loading by a processor and executing the steps of any of the methods described in the above embodiments.

[0077] See Figure 6 , Figure 6 This is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. Figure 6 As shown, terminal 600 may include: at least one terminal processor 601, at least one network interface 604, user interface 603, memory 605, and at least one communication bus 602.

[0078] The communication bus 602 is used to enable communication between these components.

[0079] The user interface 603 may include a display screen and a camera. Optionally, the user interface 603 may also include a standard wired interface and a wireless interface.

[0080] The network interface 604 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0081] The terminal processor 601 may include one or more processing cores. The terminal processor 601 connects to various parts within the terminal 600 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 605, and by calling data stored in the memory 605. Optionally, the terminal processor 601 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The terminal processor 601 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the terminal processor 601.

[0082] The memory 605 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 605 may include a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 605 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 605 may also be at least one storage device located remotely from the aforementioned terminal processor 601. Figure 6 As shown, the memory 605, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a security verification program.

[0083] exist Figure 6 In the terminal 600 shown, the user interface 603 is mainly used to provide an input interface for the user and to obtain the user's input data; while the terminal processor 601 can be used to call the security verification program stored in the memory 605 and specifically perform the following operations: In response to a security verification request triggered by a user action, obtain the original image; The original image is cropped to obtain at least one sub-image, and at least one misleading image corresponding to each sub-image is obtained. Each sub-image and each misleading image satisfy a preset similarity condition. The interactive interface displays the cropped original image, each sub-image, and each misleading image, instructing the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image. Verify user identity security based on user selection results.

[0084] In some embodiments, when the terminal processor 601 performs the step of acquiring at least one misleading image corresponding to each sub-image, it specifically performs the following steps: inputting each sub-image into the image generation model, determining at least one misleading image corresponding to each sub-image generated by the image generation model; the image generation model is obtained by training the basic generation model using sample images.

[0085] In some embodiments, the terminal processor 601 further performs the following steps: periodically collecting images from real-world scenes and / or generating images through an image generation model, and storing all collected images in an image library; when the terminal processor 601 performs the acquisition of an original image, it specifically performs the following steps: randomly selecting an image from the image library as the original image.

[0086] In some embodiments, when the terminal processor 601 performs cropping of the original image to obtain at least one sub-image, it specifically performs the following steps: randomly selecting at least one cropping position in the original image, and cropping the original image according to each cropping position to obtain at least one sub-image.

[0087] In some embodiments, the sub-images and misleading images are arranged and displayed in a random order.

[0088] In some embodiments, when the terminal processor 601 performs the following steps to verify the security of a user's identity based on the user's selection result: if the user's selection result is correct, then the user's identity is determined to be secure; if the user's selection result is incorrect, then the user's identity is determined to be at risk, and an advanced verification process is triggered for the user.

[0089] In the several embodiments provided in this specification, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0090] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0091] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0092] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0093] Furthermore, it should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the original images and user profiles involved in this specification were obtained with full authorization.

[0094] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0095] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0096] The above is a description of a security verification method, apparatus, storage medium, and terminal provided in the embodiments of this specification. For those skilled in the art, based on the ideas of the embodiments of this specification, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this specification.

Claims

1. A security verification method, the method comprising: In response to a security verification request triggered by a user action, obtain the original image; The original image is cropped to obtain at least one sub-image, and at least one misleading image corresponding to each sub-image is obtained. Each sub-image and each misleading image satisfy a preset similarity condition. The interactive interface displays the cropped original image, each sub-image, and each misleading image, instructing the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image; The user's identity security is verified based on the user's selection result.

2. The method according to claim 1, wherein obtaining at least one misleading image corresponding to each sub-image comprises: Each sub-image is input into the image generation model to determine at least one misleading image generated by the image generation model corresponding to each sub-image; The image generation model is obtained by training a basic large-scale generation model using sample images.

3. The method according to claim 1, further comprising: Regularly collect images from real-world scenes and / or generate images using image generation models, and store all collected images in an image library; The acquisition of the original image includes: One image is randomly selected from the image library as the original image.

4. The method according to claim 1, wherein cropping the original image to obtain at least one sub-image comprises: At least one cropping position is randomly selected from the original image, and the original image is cropped according to each cropping position to obtain at least one sub-image.

5. The method according to claim 1, wherein the sub-images and the misleading images are arranged and displayed in a random order.

6. The method according to claim 1, wherein verifying the user's identity security based on the user's selection result includes: If the user's selection result is correct, then the user's identity is determined to be a secure identity; If the user's selection result is incorrect, it is determined that the user's identity is at risk, and an advanced verification process is triggered for the user.

7. A security verification device, the device comprising: The process triggering module is used to obtain the original image in response to a security verification request triggered by a user operation; The image acquisition module is used to crop the original image to obtain at least one sub-image, and to acquire at least one misleading image corresponding to each sub-image, wherein each sub-image and each misleading image satisfy a preset similarity condition. The image display module is used to display the cropped original image, each sub-image, and each misleading image in the interactive interface, and to instruct the user to select the sub-image corresponding to the cropped original image from each sub-image and each misleading image; The security verification module is used to verify the user's identity security based on the user's selection result.

8. A computer program product comprising instructions that, when run on a computer or processor, causes the computer or processor to perform the steps of the method as claimed in any one of claims 1 to 6.

9. A computer storage medium storing a plurality of instructions adapted for loading by a processor and performing the steps of the method as claimed in any one of claims 1 to 6.

10. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as claimed in any one of claims 1 to 6.