A training method and device of a generation model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-08-11
AI Technical Summary
但是,一些攻击者可能使用图像注入攻击的方式实施非法认证,即将客户端的拍摄图像替换为其它图像,并将该其它图像上传至服务端进行认证
[0034] The generative model training method and apparatus provided in one or more embodiments of this specification aim to reduce the visibility score of the discriminant model in the output of the image generated by the generative model, which indicates the visibility of steganographic information. This reduces the visibility of steganographic information in the image generated by the trained generative model, thereby effectively resisting image injection attacks and ensuring the security of identity authentication.
Smart Images

Figure CN118840630B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification belong to the field of image processing technology, and in particular relate to a training method and apparatus for a generative model. Background Technology
[0002] In many business sectors, such as commerce and finance, user authentication is often performed by verifying images captured by the client. For example, in EKYC (Electronic Know Your Customer) scenarios, a key step is to authenticate the user's true identity based on their ID image and facial image. However, some attackers may use image injection attacks to carry out unauthorized authentication, replacing the client's captured image with another image and uploading that other image to the server for authentication.
[0003] Therefore, there is an urgent need to provide a reliable solution that can generate images that can effectively resist injection attacks to ensure the security of identity authentication. Summary of the Invention
[0004] The embodiments in this specification aim to provide a method for training a generative model, which can generate images that can effectively resist image injection attacks.
[0005] The first aspect of this specification provides a method for training a generative model, including:
[0006] Obtain the first training sample, which includes the first initial image and the first information;
[0007] The first training sample is input into the generation model for model processing to obtain the first target image carrying steganographic information;
[0008] The first target image is input into a pre-trained discrimination model to obtain a first visibility score; the first visibility score indicates the visibility of the steganographic information.
[0009] The parameters of the generative model are adjusted with the goal of reducing the first loss; wherein the first loss is positively correlated with the first visibility score.
[0010] A second aspect of this specification provides a method for verifying an image, executed by a client having a generative model trained according to the method described in the first aspect above, the method comprising:
[0011] Obtain the initial image to be verified;
[0012] The initial image is input into the generation model to obtain a target image carrying steganographic information;
[0013] Send the target image carrying steganographic information to the server.
[0014] A third aspect of this specification provides a method for verifying images, executed by a server that has a pre-trained discriminative model deployed on it. The method includes:
[0015] The client receives a target image to be verified from the client; the client is used to generate an image carrying steganographic information using a generative model trained according to the method described in the first aspect above.
[0016] Steganographic information is extracted from the target image using the discrimination model;
[0017] The target image is verified based on the steganographic information.
[0018] The fourth aspect of this specification provides a training apparatus for a generative model, comprising:
[0019] An acquisition unit is used to acquire a first training sample, which includes a first initial image and first information;
[0020] The input unit is used to input the first training sample into the generation model for model processing to obtain a first target image carrying steganographic information.
[0021] The input unit is further configured to input the first target image into a pre-trained discrimination model to obtain a first visibility score; the first visibility score indicates the visibility of the steganographic information.
[0022] An adjustment unit is used to adjust the parameters of the generative model with the goal of reducing a first loss; wherein the first loss is positively correlated with the first visibility score.
[0023] A fifth aspect of this specification provides an apparatus for verifying an image, disposed on a client, the client having deployed a generative model trained according to the method described in the first aspect above, the apparatus comprising:
[0024] The acquisition unit is used to acquire the initial image to be verified.
[0025] An input unit is used to input the initial image into the generation model to obtain a target image carrying steganographic information;
[0026] The sending unit is used to send the target image carrying steganographic information to the server.
[0027] The sixth aspect of this specification provides an apparatus for verifying images, disposed on a server, the server having a pre-trained discriminative model deployed thereon, the apparatus comprising:
[0028] A receiving unit is configured to receive a target image to be verified from a client; the client is configured to generate an image carrying steganographic information using a generative model trained according to the method described in the first aspect above.
[0029] The extraction unit is used to extract steganographic information from the target image through the discrimination model;
[0030] The verification unit is used to verify the target image based on the steganographic information.
[0031] A seventh aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in any one of the first to third aspects.
[0032] This specification provides a computing device in an eighth aspect, including a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method described in any one of the first to third aspects.
[0033] A ninth aspect of this specification provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any one of the first to third aspects.
[0034] The generative model training method and apparatus provided in one or more embodiments of this specification aim to reduce the visibility score of the discriminant model in the output of the image generated by the generative model, which indicates the visibility of steganographic information. This reduces the visibility of steganographic information in the image generated by the trained generative model, thereby effectively resisting image injection attacks and ensuring the security of identity authentication. Attached Figure Description
[0035] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a schematic diagram illustrating an implementation scenario of one of the embodiments disclosed in this specification;
[0037] Figure 2 This is a flowchart of the training method for generating the model in one embodiment of this specification;
[0038] Figure 3 This is an interactive diagram of a method for verifying images in one embodiment of this specification;
[0039] Figure 4 This is an interactive diagram of a method for verifying images in one embodiment of this specification;
[0040] Figure 5 This is a schematic diagram of a training device for generating a model in one embodiment of this specification;
[0041] Figure 6 This is a schematic diagram of an image verification device in one embodiment of this specification;
[0042] Figure 7 This is a schematic diagram of an image verification device in one embodiment of this specification. Detailed Implementation
[0043] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0044] As mentioned earlier, during user authentication, attackers may use image injection attacks to compromise the images uploaded by the client. To defend against image injection attacks, one approach is to add steganographic information to the captured image. This steganographic information needs to be low in visibility and highly robust. Low visibility ensures that the image carrying the steganographic information will not become a target for attackers, while high robustness ensures that the steganographic information remains unchanged even if the image is cropped, compressed, or subjected to common image processing operations.
[0045] To generate images that carry steganographic information with low visibility and high robustness, this specification proposes to train a generative model. The training process of this generative model is described below.
[0046] Figure 1 This is a schematic diagram illustrating an implementation scenario of one of the embodiments disclosed in this specification. Figure 1 In this process, the training of the generative and discriminative models can be performed alternately. For example, the discriminative model can be pre-trained first. Then, the parameters of the discriminative model are fixed, and the generative model is trained. Next, the parameters of the generative model are fixed, and the discriminative model is trained, and so on. It should be understood that each training iteration for either the discriminative or generative model can include multiple rounds of iteration.
[0047] Furthermore, regarding the aforementioned generative model, the corresponding training objective is to reduce the visibility score of the discriminative model's output indicating the visibility of steganographic information in the images generated by the generative model. Conversely, for the discriminative model, the corresponding training objective is to increase the visibility score. Therefore, this scheme trains the generative model using adversarial training, ensuring that the generative model generates images with low steganographic visibility while simultaneously ensuring that the discriminative model can recognize the steganographic information in the images generated by the generative model.
[0048] The pre-training process of the above discrimination model will be explained below.
[0049] In one embodiment, image f1 carrying explicit information can be first acquired as a positive sample and labeled as a positive example, and image f2 carrying steganographic information can be acquired as a negative sample and labeled as a negative example. Then, based on the visibility score and positive label output by the discriminant model for the positive sample, loss L1 is calculated, and based on the visibility score and negative label output by the discriminant model for the negative sample, loss L2 is calculated. Finally, based on the combined loss L1 and L2, the parameters of the discriminant model are adjusted to obtain a pre-trained discriminant model. Specifically, assuming the positive label is set to 1 and the negative label is set to 0, during training, the parameters of the discriminant model are adjusted so that the score output by the discriminant model for image f1 in the positive sample is closer to 1, and the score output for image f2 in the negative sample is closer to 0.
[0050] In a more specific embodiment, the steganographic information in image f2 is added to image f2 using a digital watermarking algorithm. This digital watermarking algorithm may include, but is not limited to, any of the following: Least Significant Bit (LSB) algorithm, frequency domain watermarking algorithm (such as DCT, DFT), wavelet transform-based algorithm, etc.
[0051] Of course, in practice, the negative samples mentioned above can also include image f3 that does not carry any information. Therefore, the loss L3 can be calculated based on the visibility score and negative example label output by the discriminant model for image f3, and the parameters of the discriminant model can be adjusted based on the combined loss of losses L1, L2 and L3.
[0052] In one example, the mean squared error (MSE) loss function or cross-entropy loss function can be used to determine the aforementioned losses L1, L2, or L3 based on the positive and negative examples' labels and the predicted visibility scores. Then, based on the combined loss, backpropagation can be used to calculate the update gradients corresponding to the discriminant model's parameters, and the discriminant model's parameters can be adjusted based on these update gradients.
[0053] Furthermore, in practical applications, the pre-training of the discriminative model involves multiple iterations, and the pre-training steps described above are only the steps involved in the t-th iteration. It can be understood that by repeatedly executing the pre-training steps, multiple rounds of parameter adjustment of the discriminative model can be achieved, and the discriminative model after the final round of parameter adjustment can then be used as the pre-trained discriminative model.
[0054] After pre-training the discriminative model, the parameters of the discriminative model can be fixed, and then the generative model can be trained. The training process of the generative model is explained below.
[0055] Figure 2 This is a flowchart illustrating a training method for generating a model in one embodiment of this specification. This method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities. Figure 2 As shown, the method may include the following steps:
[0056] Step S202: Obtain training samples, including initial images and first information.
[0057] In one embodiment, the initial image described above may be captured by the client during the execution of the target service. The target service here could be, for example, an identity authentication service.
[0058] When the target business is identity authentication, the initial image mentioned above is a user's facial image or ID card image.
[0059] The first piece of information mentioned above can be, for example, an image, text, a symbol, or a number. The following explanation will use text as the first piece of information.
[0060] Step S204: Input the training samples into the generation model for model processing to obtain the target image f1 carrying steganographic information.
[0061] In one embodiment, the generative model includes encoder t1, encoder t2, and decoder. Encoder t1 is the encoder corresponding to the image modality, which can be implemented as a convolutional neural network, such as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), FCN (Fully Convolutional Networks), etc. It can also be implemented as other neural networks, such as MLP (Multi-Layer Perceptron Network), etc. Encoder t2 is the encoder corresponding to the text modality, also known as a text encoder, and can be implemented as a Transformer encoder, meaning it encodes the first information based on an attention mechanism.
[0062] Specifically, the initial image can be processed using encoder t1 to obtain an image representation, and the first information can be processed using encoder t2 to obtain a text representation. Then, in the decoder, a target image f1 carrying steganographic information is generated based on the image representation and the text representation.
[0063] Of course, in practice, the above-mentioned generative model can also be implemented based on other networks or models that can output images based on images, and this specification does not limit this.
[0064] Step S206: Input the target image f1 into the pre-trained discrimination model to obtain the visibility score s1, which indicates the visibility of the steganographic information.
[0065] The discrimination model here can be trained through the pre-training process described above, or it can be trained through other methods for training binary classification models. This specification does not limit this method.
[0066] Of course, in practice, in addition to the visibility score s1 mentioned above, the pre-trained discriminant model can also output predicted steganalytic information.
[0067] Step S208: Adjust the parameters of the generative model with the goal of reducing the loss L1, which is positively correlated with the visibility score s1.
[0068] It should be understood that the training objective of the generative model is to identify the steganalytic information in the target image f1 that the model cannot recognize. Therefore, for the generative model, it is desirable to minimize the visibility score s1. Thus, the loss L1 can be set to be positively correlated with the visibility score s1. In this way, the direction of decreasing the loss L1 is to decrease the visibility score s1.
[0069] In one example, the loss L1 can be calculated according to the following formula (1):
[0070]
[0071] As shown in formula (1), θ G The parameters for generating the model, z i To generate the model input data, namely the initial image and initial information, G(z) i ,θ G ) represents the output of the generative model, i.e., the target image f1, D(G(z) i ,θ GLet θ be the visibility score s1 output by the pre-trained discriminative model for the target image f1. As shown in equation (1), the visibility score s1 output by the discriminative model is positively correlated with the loss L1. That is, the smaller the visibility score s1, the smaller the loss L1 of the generative model. Therefore, θ can be adjusted using various optimization algorithms, such as gradient descent. G The smaller the loss L1, the smaller the visibility score s1, thus optimizing the generative model.
[0072] Furthermore, for the steganalysis added to the target image f1, it is generally desirable that it be highly robust, meaning that even after degradation processing such as compression, cropping, adding noise, or resizing, the discriminative model should still be able to recognize the steganalysis. It should be understood that the goal of the discriminative model recognizing the steganalysis is to maximize the visibility score s2 output by the discriminative model for the downgraded image f1'. However, since the discriminative and generative models are trained adversarially, the training objective of the generative model is to minimize the visibility score s2, i.e., the loss L1 is set to be positively correlated with the visibility score s2 output by the discriminative model for the downgraded image f1'. Thus, the direction of decreasing the loss L1 is to decrease the visibility score s2.
[0073] In another example, the loss L1 can be calculated according to the following formula (2):
[0074]
[0075] The definition of the first part of Formula 2 can be found in Formula 1 above, D(f i (D(G(z i ,θ G ))) represents the visibility score s2 output by the pre-trained discriminative model for the downgraded image f1'. As can be seen from equation (2), the visibility score s2 output by the discriminative model is positively correlated with the loss L1. That is, the smaller the visibility score s2, the smaller the loss L1 of the generative model. Therefore, θ can be adjusted using various optimization algorithms, such as gradient descent. G The smaller the loss L1, the smaller the visibility score s2, thus optimizing the generative model.
[0076] It should be understood that in practice, training a generative model involves multiple iterations, meaning the steps of adjusting the generative model's parameters are repeated many times. After repeatedly performing these adjustments, the generative model's parameters can be fixed, and then a discriminative model can be trained. The training process for the discriminative model is explained below.
[0077] Specifically, the training samples mentioned above can be input into a generative model with fixed parameters to obtain a target image f2 carrying steganographic information. The target image f2 can then be input into a discriminative model to obtain a visibility score s3. The parameters of the discriminative model are adjusted with the goal of reducing the loss L2, where the loss L2 is inversely correlated with the visibility score s3.
[0078] It should be understood that the training objective of the discriminative model is to identify as much steganalytic information as possible in the target image f2. Therefore, for the discriminative model, it is desirable to maximize the visibility score s3. Thus, the loss L2 can be set to be inversely correlated with the visibility score s3. In this way, decreasing the loss L2 leads to increasing the visibility score s3.
[0079] Furthermore, the input to the discriminative model can also include the aforementioned initial image, thereby allowing L2 to be set to be inversely correlated with the visibility score s4 output by the discriminative model for the initial image.
[0080] In one example, the loss L2 can be calculated according to the following formula (3):
[0081]
[0082] As shown in formula (3), θ D To determine the parameters of the model, Positive samples include the target image f2. Let m represent the negative sample, i.e., the initial image, where the sum of i and j is m. To determine the visibility score s3 output by the model for positive samples,
[0083] To determine the visibility score s4 output by the model for negative samples, it can be seen from formula (3) that the loss L2 is inversely correlated with the visibility score s3, meaning that the larger the visibility score s3, the smaller the loss L2; conversely, the loss L2 is positively correlated with the visibility score s4, meaning that the smaller the visibility score s4, the smaller the loss L2. θ can be adjusted using, for example, gradient descent. D This reduces the L2 loss and makes the discrimination model more accurate.
[0084] Furthermore, as mentioned earlier, it is desirable for the discrimination model to output a visibility score as high as possible for the downgraded image. Therefore, L2 can also be set to be inversely correlated with the visibility score output by the discrimination model for the downgraded image. Thus, the positive samples in Equation 3 above can also include the downgraded image f2' of the target image f2.
[0085] Finally, in practice, the output of the aforementioned discriminative model can also include predicted steganalytic information. Therefore, the training objective of the discriminative model also includes aiming for the predicted steganalytic information to be as accurate as possible (i.e., as close as possible to the first information). Thus, the loss L2 can also be set to be positively correlated with the difference loss determined based on the comparison between the first information and the predicted steganalytic information. In this way, decreasing the loss L2 corresponds to reducing the aforementioned difference loss.
[0086] In one example, the difference loss is determined based on the vector distance between the first information and the predicted steganalytic information, and the difference loss is positively correlated with the vector distance, which can be, for example, Euclidean distance or cosine distance.
[0087] In summary, as can be seen from the definitions of L1 and L2 losses above, the training objectives of the generative and discriminative models are adversarial. In other words, this approach trains the generative model by borrowing the idea of adversarial learning.
[0088] It should be understood that in practice, training the discriminative model also involves multiple iterations, meaning that the steps of adjusting the discriminative model parameters are repeated many times. After repeatedly performing the steps of adjusting the discriminative model parameters, the parameters of the discriminative model can be fixed, and then the generative model can be trained. This process of alternating between training the generative and discriminative models continues until the training termination condition is met (e.g., the parameters of the generative and discriminative models converge).
[0089] In summary, the generative model training method provided in this specification's embodiments trains the generative model by drawing on the concept of adversarial learning. Furthermore, this scheme trains the generative model with the goal of reducing the visibility score of the discriminative model's output indicating the visibility of steganographic information in images generated by the generative model, resulting in images with low steganographic information visibility after training. Additionally, this scheme also trains the generative model with the goal of reducing the visibility score of the discriminative model's output indicating the degradation of images generated by the generative model, resulting in images with high steganographic robustness after training. In conclusion, the generative model trained by this scheme can generate images with low steganographic information visibility and high robustness, thereby effectively resisting image injection attacks and ensuring the security of identity authentication.
[0090] It should be noted that the generative model and discriminative model trained in the embodiments of this specification can be applied to the image verification step in the identity authentication process (i.e., verifying whether the image has been subjected to an injection attack). The image verification process based on these two models will be described below.
[0091] Figure 3 This is an interactive diagram of a method for verifying images in one embodiment of this specification, such as... Figure 3As shown, the method may include the following steps:
[0092] Step S302: The client obtains the initial image to be verified.
[0093] Among them, the client is deployed with the above-mentioned Figure 2 The generative model trained using the method shown is an example.
[0094] In one embodiment, the initial image described above may be captured by the client during the execution of the target service. The target service here could be, for example, an identity authentication service.
[0095] When the target business is identity authentication, the initial image mentioned above is a user's facial image or ID card image.
[0096] Step S304. The client inputs the initial image into the above generation model to obtain the target image carrying steganographic information.
[0097] In one embodiment, the client can also input the first information into the generative model, thereby generating a target image that carries steganographic information as the model output.
[0098] Of course, in practice, this initial information can also be built into the generative model after it has been trained, so that the client only needs to input the initial image into the generative model.
[0099] In step S306, the client sends the target image carrying steganographic information to the server.
[0100] In practice, the client can also send its client identifier to the server.
[0101] In this solution, the server deploys a discriminative model that is trained together with the aforementioned generative model.
[0102] In step S308, the server extracts steganographic information from the target image using a discrimination model.
[0103] In step S310, the server verifies the target image based on the extracted steganographic information.
[0104] Specifically, the server can retrieve the verification information corresponding to the client from the storage unit based on the client identifier received from the client, and match the verification information with the steganographic information extracted from the target image. If the match is consistent, it can be determined that the initial image was generated by the client, thereby verifying whether the initial image has been subjected to an injection attack.
[0105] In summary, the image verification method provided in this specification involves the client generating an image to be verified using a generative model, and the server extracting steganographic information from the image to be verified using a discriminative model, and verifying whether the image has been subjected to an injection attack based on the steganographic information.
[0106] Figure 4 This is an interactive diagram of a method for verifying images in one embodiment of this specification, such as... Figure 4 As shown, the method may include the following steps:
[0107] In step S402, the client inputs the user's face image and ID card image into the generation model to obtain a face image carrying steganographic information and an ID card image carrying steganographic information.
[0108] In step S404, the client sends the face image carrying steganographic information and the ID card image carrying steganographic information to the server.
[0109] In step S406, the server extracts steganographic information from the face image and the ID card image respectively using a discrimination model, and matches the two extracted steganographic information. If they match, it is determined that both the face image and the ID card image were generated by the client. If they do not match, it is determined that neither the face image nor the ID card image was generated by the client.
[0110] This enables the verification of whether facial images and ID card images have been subjected to injection attacks.
[0111] Corresponding to the training method of the generative model described above, one embodiment of this specification also provides a training apparatus for the generative model, such as... Figure 5 As shown, the device may include:
[0112] The acquisition unit 502 is used to acquire the first training sample, which includes the first initial image and the first information.
[0113] The input unit 504 is used to input the first training sample into the generation model for model processing to obtain the first target image carrying steganographic information.
[0114] The input unit 504 is also used to input the first target image into the pre-trained discrimination model to obtain a first visibility score, which indicates the visibility of the steganographic information.
[0115] Adjustment unit 506 is used to adjust the parameters of the generative model with the goal of reducing the first loss, wherein the first loss is positively correlated with the first visibility score.
[0116] In one embodiment, the above-described generative model includes: a first encoder, a second encoder, and a decoder;
[0117] Input unit 504 is specifically used for:
[0118] A first initial image is processed using a first encoder to obtain a first image representation. A second encoder is used to process the first information to obtain a first text representation. A decoder is then used to generate a first target image based on the first image representation and the first text representation.
[0119] In one embodiment, the input to the discriminative model further includes a first downgraded image obtained by downgrading the first target image, wherein the first loss is also positively correlated with the visibility score output by the discriminative model for the first downgraded image.
[0120] The aforementioned degradation processing includes at least one of the following: compression, cropping, adding noise, and resizing.
[0121] In one embodiment, the input unit 504 is further configured to input the first training sample into the generative model with fixed parameters to obtain a second target image carrying steganographic information.
[0122] The input unit 504 is also used to input the second target image into the discrimination model to obtain a second visibility score;
[0123] The adjustment unit 506 is also used to adjust the parameters of the discrimination model with the goal of reducing the second loss, wherein the second loss is inversely correlated with the second visibility score.
[0124] In one embodiment, the input to the discriminative model further includes a second downgraded image obtained by downgrading the second target image, and the second loss is also inversely correlated with the visibility score output by the discriminative model for the second downgraded image.
[0125] In one embodiment, the input to the discriminative model further includes the aforementioned first initial image, and the second loss is also positively correlated with the visibility score output by the discriminative model for the first initial image.
[0126] In one embodiment, the output of the discriminative model further includes predicted steganalysis information, and the second loss is positively correlated with the difference loss determined based on the comparison between the first information and the predicted steganalysis information.
[0127] In one embodiment, the device further includes:
[0128] The acquisition unit 508 is used to acquire positive samples and negative samples. The positive samples include a first image carrying explicit information, and the negative samples include a second image carrying steganographic information.
[0129] Training unit 510 is used to train the discrimination model based on positive and negative samples.
[0130] The steganographic information in the second image is added to the second image using a digital watermarking algorithm.
[0131] In a more specific embodiment, the negative sample mentioned above also includes a third image that does not carry any information.
[0132] The functions of each functional unit of the apparatus in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the apparatus provided in one embodiment of this specification will not be repeated here.
[0133] This specification provides a training apparatus for a generative model in one embodiment, wherein the trained generative model is capable of generating images that can effectively resist image injection attacks.
[0134] Corresponding to the image verification method described above, one embodiment of this specification also provides an image verification apparatus, configured on a client, the client being deployed with... Figure 2 The generative model trained using the method shown is an example. Figure 6 As shown, the device may include:
[0135] Acquisition unit 602 is used to acquire the initial image to be verified;
[0136] The input unit 604 is used to input the initial image into the generation model to obtain the target image carrying steganographic information.
[0137] The sending unit 606 is used to send the target image carrying steganographic information to the server.
[0138] The functions of each functional unit of the apparatus in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the apparatus provided in one embodiment of this specification will not be repeated here.
[0139] Corresponding to the image verification method described above, one embodiment of this specification also provides an image verification apparatus, located on a server, which deploys a discriminative model trained together with the generative model. For example... Figure 7 As shown, the device may include:
[0140] Receiving unit 702 is configured to receive a target image to be verified from a client, the client being configured to utilize... Figure 2 The generative model trained by the method shown generates images carrying steganographic information.
[0141] Extraction unit 704 is used to extract steganalytic information from the target image through a discrimination model.
[0142] The verification unit 706 is used to verify the target image based on steganographic information.
[0143] The functions of each functional unit of the apparatus in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the apparatus provided in one embodiment of this specification will not be repeated here.
[0144] This specification provides an image verification apparatus according to one embodiment, which can verify whether an image has been subjected to an injection attack.
[0145] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform a combination Figure 2 The method described.
[0146] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements a combination... Figure 2 The method described.
[0147] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the medium or device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0148] The steps of the methods or algorithms described in conjunction with the disclosure in this specification can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, external hard disk, CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a server. Of course, the processor and storage medium can also exist as discrete components in the server.
[0149] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog are the most commonly used. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0150] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0151] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0152] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.
[0153] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0154] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0155] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0156] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0157] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0158] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0159] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0160] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0161] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0162] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0163] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A method for training a generative model, comprising: Obtain the first training sample, which includes the first initial image and the first information; The first training sample is input into the generative model for model processing to obtain a first target image carrying steganographic information, which is used to resist image injection attacks during the identity authentication process. The first target image is input into a pre-trained discrimination model to obtain a first visibility score; The first visibility score indicates the degree of visibility of the steganographic information; The parameters of the generative model are adjusted with the goal of reducing the first loss; wherein the first loss is positively correlated with the first visibility score. After adjusting the parameters of the generative model, the parameters of the generative model are fixed, and the discriminative model is trained, with the discriminative model acting as an adversarial target to the training objective of the generative model.
2. The method according to claim 1, wherein, The generative model includes: a first encoder, a second encoder, and a decoder; The model processing includes: The first initial image is processed using the first encoder to obtain a first image representation; the first information is processed using the second encoder to obtain a first text representation; and the first target image is generated using the decoder based on the first image representation and the first text representation.
3. The method according to claim 1, wherein, The input to the discrimination model also includes a first downgraded image obtained by downgrading the first target image, and the first loss is positively correlated with the visibility score output by the discrimination model for the first downgraded image.
4. The method according to claim 3, wherein, The degradation process includes at least one of the following: compression, cropping, adding noise, and resizing.
5. The method according to claim 1, wherein, Training the discriminative model includes: By inputting the first training sample into the generative model with fixed parameters, a second target image carrying steganographic information is obtained. The second target image is input into the discrimination model to obtain the second visibility score; The parameters of the discrimination model are adjusted with the goal of reducing the second loss; wherein the second loss is inversely correlated with the second visibility score.
6. The method according to claim 5, wherein, The input to the discrimination model also includes a second downgraded image obtained by downgrading the second target image, and the second loss is also inversely correlated with the visibility score output by the discrimination model for the second downgraded image.
7. The method according to claim 5, wherein, The input to the discriminant model also includes the first initial image; the second loss is also positively correlated with the visibility score output by the discriminant model for the first initial image.
8. The method according to claim 5, wherein, The output of the discriminant model also includes predicted steganalytic information; the second loss is also positively correlated with the difference loss determined based on the comparison between the first information and the predicted steganalytic information.
9. The method according to claim 1, wherein, The discriminant model is pre-trained through the following steps: Obtain positive and negative samples; the positive samples include a first image carrying explicit information, and the negative samples include a second image carrying steganographic information; The discrimination model is trained based on the positive and negative samples.
10. The method according to claim 9, wherein, The steganographic information in the second image is added to the second image using a digital watermarking algorithm.
11. The method according to claim 9, wherein, The negative samples also include a third image that does not carry any information.
12. A method for verifying an image, executed by a client, the client being deployed with... The generative model trained according to the method of claim 1, wherein the method comprises: Obtain the initial image to be verified; The initial image is input into the generation model to obtain a target image carrying steganographic information; Send the target image carrying steganographic information to the server.
13. A method for verifying an image, executed by a server, the server being deployed with... The discriminant model trained according to the method described in claim 5, wherein the method comprises: The client receives a target image to be verified from the client; the client is used to generate an image carrying steganographic information using a generative model trained according to the method of claim 1. Steganographic information is extracted from the target image using the discrimination model; The target image is verified based on the steganographic information.
14. The method according to claim 13, wherein, The target image includes a face image and an identification document image; the verification of the target image includes: The steganographic information extracted from the face image is matched with the steganographic information extracted from the ID card image, and based on the matching result, it is determined whether the face image and the ID card image were generated by the client.
15. A training device for a generative model, comprising: An acquisition unit is used to acquire a first training sample, which includes a first initial image and first information; The input unit is used to input the first training sample into the generation model for model processing to obtain a first target image carrying steganographic information, which is used to resist image injection attacks during the identity authentication process. The input unit is further configured to input the first target image into a pre-trained discrimination model to obtain a first visibility score; the first visibility score indicates the visibility of the steganographic information. An adjustment unit is used to adjust the parameters of the generative model with the goal of reducing a first loss; wherein the first loss is positively correlated with the first visibility score; After adjusting the parameters of the generative model, the parameters of the generative model are fixed, and the discriminative model is trained, with the discriminative model acting as an adversarial target to the training objective of the generative model.
16. An image verification device, configured on a client, the client having deployed... The apparatus for training the generative model according to claim 1 includes: The acquisition unit is used to acquire the initial image to be verified. An input unit is used to input the initial image into the generation model to obtain a target image carrying steganographic information; The sending unit is used to send the target image carrying steganographic information to the server.
17. An image verification device, disposed on a server, the server being equipped with The discriminant model trained according to the method of claim 5, the apparatus comprises: A receiving unit is configured to receive a target image to be verified from a client; the client is configured to generate an image carrying steganographic information using a generative model trained according to the method described in claim 1. The extraction unit is used to extract steganographic information from the target image through the discrimination model; The verification unit is used to verify the target image based on the steganographic information.
18. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-14.
Citation Information
Patent Citations
Digital watermark training method and device, equipment and storage medium
CN114445256A
Image steganography method based on generative adversarial network
CN115086674A
Data processing method and device, electronic equipment and storage medium
CN117495646A