Method and device for generating certificate image

Automatically generate multiple layer templates for certificate images through layer generation models, solving the problem of low efficiency in generating negative samples of certificate images, improving the recognition accuracy of the image recognition network, and achieving effective recognition of forged images.

CN120298519APending Publication Date: 2025-07-11ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359504.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art generates negative samples for training image recognition networks with low efficiency and is difficult to effectively identify document images generated by AIGC.

Method used

Through pre-trained layer generation models, based on the ID image and its layer information, image templates of multiple layers are automatically generated to display various types of content in the ID image, and these image templates are used to generate images as negative samples to train the image recognition network.

Benefits of technology

The efficiency of generating document image templates and the number of negative samples are improved, the accuracy of the generated images by the image recognition network is enhanced, and the accurate recognition of forged images is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298519A_ABST
    Figure CN120298519A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for generating a certificate image, and the method comprises the steps: obtaining a first image layer corresponding to first image layer information through an image layer generation model based on a first certificate image and the first image layer information, so as to obtain a first image template, the first layer information is used for indicating a first type of content in a first certificate image needing to be displayed in a layer needing to be generated; the first image template is used for generating an image and comprises a plurality of first image layers, the plurality of first image layers are respectively used for displaying various types of contents in the first certificate image, and the image layer generation model is obtained by training in advance based on each sample certificate image, each corresponding sample image layer information and each sample image template; and at least based on the first image template, a generated image serving as a negative sample of a training image recognition network is obtained, and the image recognition network is used for recognizing whether the input image is a real image, so that the efficiency of generating the negative sample for training the image recognition network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method and device for generating document images. Background Art

[0002] With the development and popularization of AIGC (Artificial Intelligence Generated Content) technology, a large number of generated images based on AIGC have emerged. In the network environment, people often cannot distinguish between real images and generated images.

[0003] In some scenarios, generated images based on AIGC may cause troubles to people's lives. For example, in the financial field scenario, there have been multiple cases of using AIGC technology to generate document images (hereinafter referred to as generated images) to attack the identity verification system of financial institutions. In order to defend against attacks using generated images obtained by AIGC technology, it is crucial to provide an image recognition network for distinguishing the authenticity of images. In order to ensure that the trained image recognition network can well identify and distinguish the authenticity of input images, it is necessary to generate a sufficient number of generated images as negative samples to train the image recognition network.

[0004] Then, how to provide a method for generating document images to more conveniently obtain negative samples for training an image recognition network has become an urgent problem to be solved. Summary of the Invention

[0005] One or more embodiments of this specification provide a method and device for generating document images to automatically generate an image template of a document image, so as to improve the efficiency of generating negative samples for training an image recognition network.

[0006] According to a first aspect, a method for generating a document image is provided, including:

[0007] Obtain a first document image and its corresponding first layer information, where the first layer information is used to indicate first-type content in the first document image to be displayed in a required generated layer;

[0008] Based on the first document image and the first layer information, through a layer generation model, obtain a first layer corresponding to the first layer information to obtain a first image template, where the first image template is used to generate an image, the first image template includes a plurality of first layers, and the plurality of first layers are respectively used to display various types of content in the first document image, and the layer generation model is pre-trained based on various sample document images, their corresponding various sample layer information, and various sample image templates;

[0009] Based on the first image template corresponding to the first image, a generated image is obtained, where the generated image serves as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

[0010] According to a second aspect, there is provided an apparatus for generating a certificate image, including:

[0011] An acquisition module configured to acquire a first certificate image and its corresponding first layer information, where the first layer information is used to indicate first-type content in the first certificate image to be displayed in a layer to be generated;

[0012] A first obtaining module configured to, based on the first certificate image and the first layer information, through a layer generation model, obtain a first layer corresponding to the first layer information to obtain a first image template, where the first image template is used to generate an image, the first image template includes a plurality of first layers, and the plurality of first layers are respectively used to display various types of content in the first certificate image, and the layer generation model is pre-trained based on respective sample certificate images, their corresponding respective sample layer information, and respective sample image templates;

[0013] A second obtaining module configured to obtain a generated image at least based on the first image template corresponding to the first image, where the generated image serves as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

[0014] According to a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in the first aspect.

[0015] According to a fourth aspect, there is provided a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.

[0016] A method and apparatus for generating a certificate image according to an embodiment of this specification, obtain a first certificate image and its corresponding first layer information, where the first layer information is used to indicate first type content in the first certificate image to be displayed in the layer to be generated; based on the first certificate image and the first layer information, through a layer generation model, obtain a first layer corresponding to the first layer information to obtain a first image template, where the first image template is used to generate an image, the first image template includes multiple first layers, and the multiple first layers are respectively used to display various types of content in the first certificate image, and the layer generation model is pre-trained based on each sample certificate image, its corresponding sample layer information, and each sample image template; then, obtain a generated image based at least on the first image template corresponding to the first image, where the generated image is used as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

[0017] In the above process, through a pre-trained layer generation model, based on the first certificate image and the first layer information indicating the first type content in the first certificate image to be displayed in the layer to be generated, automatically generate a first layer corresponding to the first layer information to obtain a first image template corresponding to the first certificate image, which includes multiple first layers for displaying various types of content in the first certificate image, so as to improve the efficiency of generating the image template of the certificate image. Further, obtain a generated image based at least on the first image template corresponding to the first image, so as to improve the efficiency and diversification of generating the negative sample of the image recognition network. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a schematic diagram of the implementation framework of an embodiment disclosed in this specification;

[0020] Figure 2 It is a schematic flowchart of a method for generating a certificate image provided in the embodiment;

[0021] Figure 3 It is a schematic structural diagram of a layer generation model provided in the embodiment;

[0022] Figure 4 It is a schematic block diagram of an apparatus for generating a certificate image provided in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The technical solutions of the embodiments of this specification will be described in detail below in conjunction with the accompanying drawings.

[0024] The embodiments of this specification disclose a method and device for generating certificate images. First, the application scenarios and technical concepts of the method will be introduced as follows:

[0025] As mentioned above, given that the current AIGC technology cannot directly generate complete certificate images, some black industries generally obtain forged certificate images based on the PSD templates of manually created certificate images. In view of the current methods of black industry forging images, in order to better and accurately identify the forged images by the black industry, it can be considered to also use the PSD templates of certificate images to generate new images, so as to use the generated images as negative samples to train the image recognition network, so that the trained image recognition network can better identify whether the input image is a generated image.

[0026] Currently, manually creating PSD templates results in low efficiency in generating negative samples for training the image recognition network. Then, how to provide a method for generating certificate images to more conveniently obtain negative samples for training the image recognition network has become an urgent problem to be solved.

[0027] In view of this, the inventor proposes a method for generating certificate images. Figure 1 The schematic diagram of the implementation scenario according to an embodiment disclosed in this specification is shown. In this implementation scenario, a trained layer generation model is obtained, where the layer generation model is pre-trained based on each sample certificate image, its corresponding sample layer information, and each sample image template; then, a first certificate image and its corresponding first layer information are obtained, and the first layer information is used to indicate the first type of content in the first certificate image that needs to be displayed in the required generated layer; based on the first certificate image and the first layer information, through the layer generation model, a first layer corresponding to the first layer information is obtained to obtain a first image template, where the first image template is used to generate an image, and the first image template includes multiple first layers, and the multiple first layers are respectively used to display various types of content in the first certificate image; at least based on the first image template corresponding to the first image, a generated image is obtained, where the generated image is used as a negative sample for training the image recognition network, and the image recognition network is used to identify whether the input image is a real image.

[0028] In the above process, through a pre-trained layer generation model, based on the first document image and the first layer information indicating the first type of content in the first document image to be displayed in the required generated layer, a first layer corresponding to the first layer information is automatically generated to obtain a first image template corresponding to the first document image and including a plurality of first layers for displaying various types of content in the first document image, realizing the automation of generating an image template corresponding to an image, so as to improve the efficiency of generating an image template for a document image. Further, at least based on the first image template corresponding to the first image, a generated image is obtained to improve the efficiency and diversity of negative samples of the generated image recognition network.

[0029] The following combines specific embodiments to elaborate in detail on the method for generating a document image provided in this specification.

[0030] Figure 2 The flowchart of the method for generating a document image in an embodiment of this specification is shown. This method is executed by an electronic device, and the electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. In the process of generating a document image, as Figure 2 shown, the method includes the following steps S210 - S230:

[0031] In step S210, a first document image and its corresponding first layer information are obtained, where the first layer information is used to indicate the first type of content in the first document image to be displayed in the required generated layer.

[0032] Among them, the document in the first document image can be any type of document, and its types can include, for example, but are not limited to: ID card, driver's license, passport, and documents processed for certain needs, etc. The documents processed for certain needs include, for example, work permit and health certificate, etc. It can be understood that the first document image can be any image obtained by an image acquisition device by photographing a real document. Correspondingly, the first document image is any real document image.

[0033] The first layer information is used to indicate the first type of content in the first certificate image to be displayed in the layer to be generated. The first type of content may be one or more of various types of content in the first certificate image. In some possible examples, each type of content may include at least one of the following: the face area image in the certificate, the text content in the certificate, the certificate background, and the environmental background where the certificate is located. In some cases, the text content in the certificate may be further divided into: the text title in the certificate and the personal information of the certificate user. Among them, the text titles in the certificate include, for example, but are not limited to, relevant text titles such as "Name" and "Date of Birth" that are fixedly present in the certificate, and the personal information of the certificate user includes, for example, but is not limited to, the specific name content of the certificate user, the specific date of birth, etc.

[0034] In still some other possible examples, each type of content may further include: the image of the area where the certificate is located and the environmental background where the certificate is located, etc. Among them, the image of the certificate area may refer to the entire image of the area where the certificate is located, including the certificate background, the text content in the certificate, and the face area image in the certificate. It can be understood that the specific classification of this type of content can be set according to actual needs.

[0035] Exemplarily, the environmental background where the certificate is located may include, for example, but is not limited to: the desktop, the ground, or the hand-related background when the certificate is held.

[0036] In some implementation manners, the first layer information may be text information and / or image information, and its specific type may be set according to the structure of the layer generation model.

[0037] After the electronic device obtains the first certificate image and its corresponding first layer information, in step S220, based on the first certificate image and the first layer information, through the layer generation model, a first layer corresponding to the first layer information is obtained to obtain a first image template.

[0038] Among them, the first image template is used to generate an image. The first image template includes multiple first layers, and the multiple first layers are respectively used to display various types of content in the first certificate image. The layer generation model is pre-trained based on each sample certificate image, its corresponding sample layer information, and each sample image template.

[0039] In some implementation manners, the electronic device may obtain a pre-trained layer generation model. Then, the first certificate image and the first layer information are input into the layer generation model to process the first certificate image and the first layer information through the image generation model, and a first layer corresponding to the first layer information is obtained, where the first layer displays the first type of content in the first certificate image.

[0040] In some possible examples, there may be one or more pieces of the first layer information, and each piece of the first layer information is respectively used to indicate a certain type of content in the first document image to be displayed in a required generated layer. For example, the first layer information includes the first layer information 1, and the first layer information 1 is used to indicate, for example, the face area image in the document in the first document image to be displayed in the required generated layer. Another example is that on the basis of including the first layer information 1, the first layer information may further include the first layer information 2, and the first layer information 2 is used to indicate, for example, the text content in the document in the first document image to be displayed in the required generated layer, and so on. Correspondingly, based on the first document image, the first layer information 1, and the first layer information 2, the electronic device obtains, through the layer generation model, the layer corresponding to the first layer information 1 and the layer corresponding to the first layer information 2.

[0041] In some possible implementation manners, the layer generation model may be implemented as a network model based on the Diffusion algorithm, and it may also be implemented as other models capable of generating layers. For example, the layer generation model may be implemented as an image segmentation network and an image inpainting network. Specifically, through the image segmentation network and the image inpainting network in the layer generation model, various types of content in the document image can be segmented first, and then, for example, the document background and the background of the environment where the document is located obtained by segmentation are inpainted to obtain each layer showing various types of content.

[0042] Taking the layer generation model implemented as a network model based on the Diffusion algorithm as an example for illustration below. The network model based on the Diffusion algorithm can be understood as adding an encoding network and a decoding network for learning transparent layers on the framework of a latent diffusion model (such as the Stable Diffusion model SD model). Among them, the latent diffusion model has been trained based on a number of sample text-image pairs, and it may include a text encoding network, an image encoding network, a diffusion network, and an image decoding network, and it already has the ability to generate an image matching the input, such as including input text and a noisy image.

[0043] Exemplarily, the image encoding network and the image decoding network may be implemented as the encoder part and the decoder part of a VAE (Variational Autoencoder), or may also be implemented as other encoding and decoding networks capable of encoding and decoding images. The diffusion network is a U-Net structure, which can gradually denoise the input noisy image to generate a corresponding latent representation, so as to decode the latent representation through the image decoding network to obtain an image. Among them, the noisy image may be generated by adding Gaussian noise to the output of the image encoding network, or may be directly generated based on Gaussian noise.

[0044] In some possible examples, the network model based on the Diffuse algorithm may include the aforementioned trained diffusion model, as well as an encoding network and a decoding network for learning the transparent layer. Hereinafter, the encoding network and the decoding network for learning the transparent layer are respectively referred to as the transparent layer encoding network and the transparent layer decoding network. The structure of the transparent layer encoding network may refer to the structure of the image encoding network. The transparent layer decoding network may be a U-net structure. The encoding sub-network part and the decoding sub-network part in the transparent layer decoding network may respectively refer to the encoder part and the decoder part of the VAE.

[0045] In some cases, the transparent layer encoding network and the transparent layer decoding network in the network model based on the Diffusion algorithm may be trained first using a preset first training dataset, so that the trained network model based on the Diffusion algorithm can implement the encoding and decoding capabilities of the transparent layer. During this training process, the parameters of the diffusion model therein may be fixed, and the parameters of the transparent layer encoding network and the transparent layer decoding network therein may be trained and adjusted. The first training dataset may include a number of transparent images and their corresponding input texts. The transparent image may refer to an image in the ARGB (Alpha Red Green Blue) format, where Alpha is the transparent layer channel. The training for this can refer to the training process of the transparent layer encoding network and the transparent layer decoding network in the network model based on the Diffusion algorithm in the related art.

[0046] Through the aforementioned training process, a pre-trained network model based on the Diffusion algorithm is obtained. Then, a layer generation model can be constructed based on the network model based on the Diffusion algorithm. The construction process may include: setting corresponding fine-tuning structures for each type of content of the document image, that is, each layer required to display each type of content in the document image. For example, assume that it is necessary to layer the document image into 4 layers, namely layer 1 for displaying the text content in the document in the document image, layer 2 for displaying the face area image in the document in the document image, layer 3 for displaying the background of the document in the document image, and layer 4 for displaying the environmental background of the document in the document image, such as a layer for displaying a desktop. Corresponding fine-tuning structures need to be set for the aforementioned layers 1-4 respectively.

[0047] Exemplarily, the fine-tuning structure can be a LORA (Low-Rank Adaptation of Large Language) fine-tuning structure. Correspondingly, the fine-tuning structure corresponding to each layer can include multiple trainable low-rank matrices, where each low-rank matrix in the fine-tuning structure corresponding to each layer corresponds to each specified network layer in the aforementioned diffusion network respectively, so as to obtain a layer diffusion network corresponding to each layer, that is, each type of content, and thus obtain a layer generation model. Among them, the layer diffusion network corresponding to each layer, that is, each type of content, includes the aforementioned diffusion network and the fine-tuning structure corresponding to each layer. Each specified network layer can include the attention layer and the specified linear layer in the aforementioned diffusion network, or each specified network layer can include each network layer in the aforementioned diffusion network.

[0048] In some possible examples, as Figure 3 shown, the layer generation model can include: an image encoding network, an auxiliary encoding network, a transparent layer decoding network, and a first layer diffusion network corresponding to the first type of content. It can be understood that the layer generation model also includes layer diffusion networks corresponding to other types of content.

[0049] After that, the layer generation model can be trained based on each sample certificate image, its corresponding sample layer information, and each sample image template, so that each layer diffusion network in the layer generation model respectively learns to denoise the latent noise in the features of the type of content it corresponds to, so as to obtain a trained layer generation model; then, the trained layer generation model is used to execute the subsequent process of generating the image template of the certificate image, so as to realize the layering of each type of content in the certificate image.

[0050] It can be understood that during the process of training the layer generation model, generally only the parameters of the fine-tuning structure in the layer diffusion network corresponding to each type of content can be adjusted, while the parameters of other structures of the layer generation model are fixed. In some examples, by sharing the parameters of the attention layer in the layer diffusion network corresponding to each type of content, such as the QKV (query key value) parameters in the attention layer, the parameters of the fine-tuning structure corresponding to each type of content in the layer generation model can be adjusted to maintain the coherence between the generated layers.

[0051] The sample image template corresponding to each sample document image includes multiple sample layers, and the multiple sample layers are respectively used to display various types of content in the corresponding sample document image. The format of the sample image template can be PSD (Photoshop Document) format. Each sample layer information is used to indicate various types of content in the sample document image to be displayed in the layer to be generated. Among them, in the layer that displays the environmental background where the sample document is located in the sample document image in the sample image template, the environmental background is image-restored after the area where the sample document is located is segmented from the sample document image; and in the layer that displays the sample document background (which can also be called the document foreground relative to the environmental background where the sample document is located) in the sample document image in the sample image template, the sample document background can also be image-restored after the text content and the face area image in the sample document are segmented from the sample document image.

[0052] Correspondingly, in some possible examples, step S220 may include the following steps 11-13:

[0053] In step 11, based on the first layer information, an auxiliary feature is obtained through an auxiliary encoding network. In this step, the electronic device may input the first layer information into the auxiliary encoding network, and process the first layer information through the auxiliary encoding network to obtain the auxiliary feature.

[0054] In some possible examples, the auxiliary encoding network may include: a text encoding network. Correspondingly, the first layer information includes: a first text prompt, and the first text prompt includes the type of content to be displayed in the layer to be generated and the display position information. Among them, the text encoding network may be implemented as a CLIP (Contrastive Language-Image Pre-training) text encoder. It can be understood that the text encoding network belongs to the aforementioned diffusion model.

[0055] In some examples, if the type of content to be displayed in the layer to be generated is text, in order to ensure the accuracy of the generated text content, the first text prompt may further include the text content to be displayed in the layer to be generated. Exemplarily, if the type of content to be displayed in the layer to be generated is text and the personal information of the document user needs to be generated, correspondingly, the first text prompt may further include the specific personal information of the document user, such as but not limited to: specific name (such as xx), specific date of birth information, specific address information, etc. Correspondingly, the display position information includes but is not limited to the display position information corresponding to the specific name (such as xx), the display position information corresponding to the specific date of birth information, and the display position information corresponding to the specific address information.

[0056] In some other possible examples, the auxiliary encoding network may include: the aforementioned transparent layer encoding network. The first layer information includes: a first transparent layer image, which is used to occlude other contents in the layer to be generated except the first type of content in the first document image to be displayed. Among them, the first transparent layer image may be a mask image. The first transparent layer image may be obtained by segmenting the type content of the first document image. The value range of the pixel values in the first transparent layer image may be [0, 1], where 1 may represent opaque and 0 may represent transparent. The first transparent layer image may indicate that other contents in the layer to be generated except the first type of content in the first document image to be displayed are all transparent.

[0057] In some examples, as Figure 3 shown, the auxiliary encoding network may also include the aforementioned transparent layer encoding network and text encoding network at the same time.

[0058] In step 12, based on the first document image, image features are obtained through the image encoding network. In this step, the electronic device inputs the first document image into the image encoding network, encodes the first document image through the image encoding network, and obtains the latent representation of the first document image, that is, obtains the image features.

[0059] In step 13, based on the auxiliary features and the image features, the first layer corresponding to the first layer information is obtained through the first layer diffusion network and the transparent layer decoding network.

[0060] In some possible examples, when the auxiliary features include the latent representation of the transparent layer obtained after the aforementioned transparent layer encoding network encodes the first transparent layer image, the electronic device may add the auxiliary features to the image features, for example, add the auxiliary features and the image features pixel by pixel to obtain a mixed image feature mixed with the latent representation of the transparent layer.

[0061] After adding noise to the mixed image features to obtain first noise features, the first noise features are input into the first layer diffusion network so that the first layer diffusion network performs denoising processing on the first noise features to obtain the output features of the first layer diffusion network. Among them, the denoising process may include: the first layer diffusion network uses its own structure (including each network layer of the original diffusion network and each low-rank matrix corresponding to the specified network layer) to process the first noise features to obtain predicted noise, and then removes the predicted noise from the first noise features to obtain new first noise features; the first layer diffusion network continues to process the new first noise features to obtain new predicted noise, and then continues to remove the new predicted noise from the new first noise features, and so on, until the output condition is reached to obtain the output image features of the first layer diffusion network, and the output features correspond to the features corresponding to the first type of content in the first document image. Exemplarily, the output condition may be that the number of times of predicted noise reaches a preset number of times.

[0062] In some other possible examples, when the auxiliary features include the text latent representation obtained after the text encoding network encodes the first text prompt as described above, the electronic device may add noise to the image features to obtain second noise features; then, the second noise features and the text latent representation are input into the first layer diffusion network so that the first layer diffusion network performs denoising processing on the second noise features based on the text latent representation to obtain the output features of the first layer diffusion network. Among them, the denoising process may refer to the foregoing denoising process and will not be elaborated here. It should be noted that in this example, when the first layer diffusion network determines the predicted noise, it is determined by jointly processing the second noise features and the text latent representation.

[0063] In some other possible examples, when the auxiliary features include the text latent representation obtained after the text encoding network encodes the first text prompt as described above, and the transparent layer latent representation obtained after the transparent layer encoding network encodes the first transparent layer image as described above, the process of the foregoing example may be combined. The electronic device may add the transparent layer latent representation to the image features to obtain mixed image features. Noise is added to the mixed image features to obtain third noise features, and the third noise features and the foregoing text latent representation are input into the first layer diffusion network so that the first layer diffusion network uses its own structure (including each network layer of the original diffusion network and each low-rank matrix corresponding to the specified network layer) to perform denoising processing on the third noise features based on the text latent representation to obtain the output features of the first layer diffusion network, and the output features correspond to the features corresponding to the first type of content in the first document image.

[0064] Next, after obtaining the output features of the first layer diffusion network, input the output features of the first layer diffusion network into the aforementioned image decoding network to obtain an initial decoded image; input the initial decoded image and the aforementioned mixed image features (or the output features of the first layer diffusion network) into the transparent layer decoding network to process the initial decoded image and the aforementioned mixed image features (or the output features of the first layer diffusion network) through the transparent layer decoding network, and obtain a first layer corresponding to the first layer information, where the first layer can be a layer in ARGB format.

[0065] In this way, through the above method, based on the first document image and the corresponding layer information for indicating various types of content in the first document image that need to be displayed in the required generated layer, through the layer generation model, multiple first layers for respectively displaying various types of content in the first document image can be obtained, thereby obtaining a first image template.

[0066] Next, in step S230, based at least on the first image template corresponding to the first image, a generated image is obtained, where the generated image is used as a negative sample for training an image recognition network, and the image recognition network is used to identify whether the input image is a real image.

[0067] Among them, the first image template includes multiple layers for respectively displaying various types of content in the first document image, and each layer displays its corresponding type of content. In some examples, the area for displaying its corresponding type of content in a single layer is generally opaque, and the other areas except the area for displaying its corresponding type of content are generally transparent.

[0068] In some possible examples, in step S230, it may include: in step 21, use a specified tool to modify the content displayed in any first layer in the first image template to obtain a generated image.

[0069] In some implementation manners, the specified tool can be any image editing tool that can modify the content of the layer. Exemplarily, the format of the first image template can be PSD format. Correspondingly, the specified tool can be a tool that can edit the template in this format, such as the PS software. It can be understood that when the format of the first image template is other formats, the specified tool can also be an image editing tool that can modify the content of each layer in the corresponding other format template.

[0070] In some possible examples, the first image template may include, for example: a first layer 1 for displaying the face region image in the first document image, a first layer 2 for displaying the text content in the first document image, a first layer 3 for displaying the document background of the document in the first document image, and a first layer 4 for displaying the environmental background where the document in the first document image is located.

[0071] As described above, modifying the content displayed by any of the first layers in the first image template may involve modifying the content displayed by one or more of the first layers in the first image template. For example, the specific name content (e.g., from XX to YY), specific date of birth (e.g., from x year x month x day to y year y month y day), and / or specific address content and other text content in the personal information of the document user displayed in the aforementioned first layer 2 can be modified. For another example, the aforementioned first layer 4 can be modified. For example, the original ornament object on the desktop background displayed in the first layer 4 can be removed, or the original color of the desktop background displayed in the first layer 4 can be adjusted to another color. For another example, the first layer 1 can be modified by replacing the face region image therein with another face region image.

[0072] In some specific examples, when modifying the first layer corresponding to the text content in the document, the electronic device also needs to detect the font information of the text content displayed in the first layer, and then, based on the detected font information, adjust the font of the obtained target text content to obtain the target text content with the adjusted font; and use the target text content with the adjusted font to replace the corresponding text content to be replaced in the first layer.

[0073] After modifying the content displayed by any of the first layers in the first image template using a specified tool, a modified first image template is obtained, and then the different layers in the modified first image template are fused and superimposed. Among them, a layer fusion and superposition algorithm in related technologies such as a transparent layer fusion and superposition algorithm can be used to fuse and superimpose the different layers in the modified first image template.

[0074] It can be understood that the first image template and the second image template mentioned later may correspond to corresponding transparent layer information. For example, each layer therein can respectively correspond to its corresponding transparent layer information. The transparent layer information can be generated by the aforementioned layer generation model or set by the user according to requirements.

[0075] In some other possible examples, in step S230, the following steps 31-32 may be included:

[0076] In step 31, obtain a second image template corresponding to the second document image. The second image template includes multiple second layers, and the multiple second layers are respectively used to display various types of content in the second document image. In this step, the second image template can be generated according to the generation process of the foregoing first image template. It can be pre-stored in a specified storage space. Correspondingly, the electronic device can obtain the second image template corresponding to the second document image from the specified storage space.

[0077] In step 32, based on the first image template and the second image template, obtain a generated image. In this step, the electronic device can replace the layer for displaying the specified type of content in the first image template with the layer for displaying the specified type of content in the second image template to obtain a third image template, and then fuse and superimpose different layers in the third image template to obtain a generated image.

[0078] In the above embodiment, through a pre-trained layer generation model, based on the first document image and the first layer information indicating the first type of content in the first document image that needs to be displayed in the layer to be generated, automatically generate a first layer corresponding to the first layer information to obtain a first image template corresponding to the first document image, which includes multiple layers for displaying various types of content in the first document image. It is possible to more conveniently obtain image templates corresponding to more document images, so as to improve the efficiency of generating image templates for document images. Then, based on these image templates corresponding to the document images, more generated images with more diverse styles can be obtained. Through these generated images, the image recognition network can be better trained to better improve the accuracy of the recognition results of the image recognition network, and achieve accurate recognition of the generated images, that is, forged images.

[0079] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments, and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or consecutive order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible, or may be advantageous.

[0080] Corresponding to the above method embodiment, an embodiment of this specification provides a device 400 for generating a document image, and its schematic block diagram is as Figure 4 shown, including:

[0081] An obtaining module 410, configured to obtain a first document image and its corresponding first layer information, where the first layer information is used to indicate the first type of content in the first document image that needs to be displayed in the layer to be generated;

[0082] The first obtaining module 420 is configured to obtain a first layer corresponding to the first layer information through a layer generation model based on the first certificate image and the first layer information, so as to obtain a first image template, where the first image template is used to generate an image, the first image template includes a plurality of first layers, and the plurality of first layers are respectively used to display various types of content in the first certificate image. The layer generation model is pre-trained based on each sample certificate image, its corresponding sample layer information, and each sample image template;

[0083] The second obtaining module 430 is configured to obtain a generated image based at least on the first image template corresponding to the first image, where the generated image is used as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

[0084] In some possible examples, the second obtaining module 430 is specifically configured to use a specified tool to modify the content displayed by any first layer in the first image template to obtain the generated image.

[0085] In some possible examples, the second obtaining module 430 is specifically configured to obtain a second image template corresponding to a second certificate image, where the second image template includes a plurality of second layers, and the plurality of second layers are respectively used to display various types of content in the second certificate image; and obtain the generated image based on the first image template and the second image template.

[0086] In some possible examples, the various types of content include at least one of the following: an image of a face area in a certificate, text content in a certificate, a certificate background, and an environmental background where the certificate is located.

[0087] In some possible examples, the layer generation model includes: an image encoding network, an auxiliary encoding network, a transparent layer decoding network, and a first layer diffusion network corresponding to the first type of content;

[0088] The first obtaining module 420 is specifically configured to obtain auxiliary features through the auxiliary encoding network based on the first layer information; obtain image features through the image encoding network based on the first certificate image; and obtain the first layer corresponding to the first layer information through the first layer diffusion network and the transparent layer decoding network based on the auxiliary features and the image features.

[0089] In some possible examples, the auxiliary encoding network includes: a text encoding network, and the first layer information includes: a first text prompt, and the first text prompt includes the type of content to be displayed in the layer to be generated and the display position information.

[0090] In some possible examples, if the type of content to be displayed in the layer to be generated is text, the first text prompt further includes the text content to be displayed in the layer to be generated.

[0091] In some possible examples, the auxiliary encoding network includes: a transparent layer encoding network, and the first layer information includes: a first transparent layer image, and the first transparent layer image is used to block other content in the layer to be generated except the first type of content in the first document image to be displayed therein.

[0092] The above device embodiments correspond to the method embodiments. For specific descriptions, reference may be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference may be made to the corresponding method embodiments.

[0093] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method for generating a document image provided in this specification.

[0094] The embodiments of this specification also provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method for generating a document image provided in this specification is implemented.

[0095] The various embodiments in this specification are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the descriptions are relatively simple, and for the relevant parts, reference can be made to the partial descriptions of the method embodiments.

[0096] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0097] The specific embodiments described above further elaborate in detail the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above description is only the specific embodiments of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a certificate image, comprising: Obtaining a first certificate image and its corresponding first layer information, wherein the first layer information is used to indicate the first type of content in the first certificate image to be displayed in the layer to be generated; Based on the first certificate image and the first layer information, through a layer generation model, obtaining a first layer corresponding to the first layer information to obtain a first image template, wherein the first image template is used to generate an image, the first image template includes a plurality of first layers, and the plurality of first layers are respectively used to display various types of content in the first certificate image, and the layer generation model is pre-trained based on each sample certificate image, its corresponding sample layer information, and each sample image template; Obtaining a generated image based at least on the first image template corresponding to the first image, wherein the generated image is used as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

2. The method according to claim 1, wherein, The obtaining the generated image includes: Using a specified tool to modify the content displayed in any first layer of the first image template to obtain the generated image.

3. The method according to claim 1, wherein The obtaining the generated image includes: Obtaining a second image template corresponding to a second certificate image, wherein the second image template includes a plurality of second layers, and the plurality of second layers are respectively used to display various types of content in the second certificate image; Based on the first image template and the second image template, obtaining the generated image.

4. The method according to claim 1, wherein The various types of content include at least one of the following: the face area image in the certificate, the text content in the certificate, the certificate background, and the environmental background where the certificate is located.

5. The method according to any one of claims 1-4, wherein, The layer generation model includes: an image encoding network, an auxiliary encoding network, a transparent layer decoding network, and a first layer diffusion network corresponding to the first type of content; The obtaining the first layer corresponding to the first layer information includes: Based on the first layer information, through the auxiliary encoding network, obtaining auxiliary features; Based on the first certificate image, through the image encoding network, obtaining image features; Based on the auxiliary features and the image features, through the first layer diffusion network and the transparent layer decoding network, obtaining the first layer corresponding to the first layer information.

6. The method according to claim 5, wherein, The auxiliary encoding network includes: a text encoding network, and the first layer information includes: a first text prompt, and the first text prompt includes the type of content to be displayed in the layer to be generated and the display position information.

7. The method according to claim 6, wherein If the type of content to be displayed in the layer to be generated is text, the first text prompt further includes the text content to be displayed in the layer to be generated.

8. The method according to claim 5, wherein, The auxiliary encoding network includes: a transparent layer encoding network, and the first layer information includes: a first transparent layer image, and the first transparent layer image is used to block other content in the layer to be generated except the first type of content in the first certificate image to be displayed therein.

9. A device for generating a certificate image, comprising: An acquisition module, configured to acquire a first certificate image and its corresponding first layer information, where the first layer information is used to indicate first type content in the first certificate image to be displayed in a required generated layer; A first obtaining module, configured to obtain a first layer corresponding to the first layer information through a layer generation model based on the first certificate image and the first layer information, so as to obtain a first image template, where the first image template is used to generate an image, the first image template includes a plurality of first layers, and the plurality of first layers are respectively used to display various types of content in the first certificate image, and the layer generation model is pre-trained based on each sample certificate image, its corresponding each sample layer information, and each sample image template; A second obtaining module, configured to obtain a generated image at least based on the first image template corresponding to the first image, where the generated image is used as a negative sample for training an image recognition network, and the image recognition network is used to identify whether an input image is a real image.

10. A computing device, comprising a memory and a processor, wherein, An executable code is stored in the memory, and when the processor executes the executable code, the method described in any one of claims 1-8 is implemented.