Model training data construction method and device and storage medium
The construction of training data sets through automated modification methods solves the problem of low efficiency in building tampered image data sets in the existing technology, and efficient and accurate image detection model training is achieved, and a variety of tampering methods can be identified, which improves detection accuracy.
Patent Information
- Application Number
- CN202510963811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The method of constructing tampered image data sets in the prior art is inefficient, requires a lot of human resources, and the image detection model is insufficient in detecting tampered credentials.
The training data set is constructed using automated modification methods, including text replacement, image mirroring, information erasing, cross-graph modification and noise addition, etc., and complex tampered images are generated through the diffusion model to train the image detection model.
It improves the efficiency of the training data set construction, enhances the accuracy of the image detection model for tampering credentials, and can identify complex tampering images generated by traditional and diffusion models.
Smart Images

Figure CN120472267A_ABST
Abstract
Description
Technical Field
[0001] This application specification relates to the field of artificial intelligence technology, and in particular to a method, device, and storage medium for constructing model training data. Background Art
[0002] With the advancement of digitalization, users are required to provide credentials to verify the authenticity of business data in various business scenarios, including but not limited to finance, e-commerce, and online education. In these scenarios, verifying the authenticity of credentials is a major challenge in automated review. Credentials are often in the form of images.
[0003] With the continuous development of image processing technology and artificial intelligence (AI), attackers have also begun to use image processing technology and AI to tamper with credentials and create false credentials that are difficult to distinguish between true and false. This has brought great risks and pressure to the business data review process.
[0004] To verify the authenticity of credentials, we can currently train image detection models by manually constructing datasets containing tampered images. The trained image detection models are then used to perform image detection on credentials. However, this existing method of manually constructing datasets containing tampered images is inefficient and requires a significant amount of human resources.
[0005] Based on this, this specification provides a method for constructing model training data. Summary of the Invention
[0006] The present application provides a method, device, storage medium and electronic device for constructing model training data to at least partially solve the above-mentioned problems existing in the prior art.
[0007] This application specification adopts the following technical solutions: This application specification provides a method for constructing model training data, the method comprising: Acquire an initial image data set, where the initial image data set includes a plurality of initial images; For each initial image in the initial image data set, determining a region to be modified in the initial image; For each area to be modified, modify the area to be modified according to a preset automated modification method to obtain a target image, wherein the preset automated modification method includes at least one method selected from the group consisting of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification using a diffusion model; A training data set for training an image detection model is constructed based on each target image and an initial image corresponding to each target image.
[0008] This application specification provides a device for constructing model training data, the device comprising: An initial image data set acquisition module is used to acquire an initial image data set, wherein the initial image data set includes a plurality of initial images; a module for determining an area to be modified, configured to determine, for each initial image in the initial image data set, an area to be modified in the initial image; a modification module, configured to modify each area to be modified according to a preset automated modification method to obtain a target image, wherein the preset automated modification method includes at least one of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification using a diffusion model; The training data set construction module is used to construct a training data set for training the image detection model based on each target image and the initial image corresponding to each target image.
[0009] The present application specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for constructing model training data.
[0010] The present application specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for constructing model training data when executing the computer program.
[0011] At least one of the above technical solutions adopted in this application specification can achieve the following beneficial effects: In the method for constructing model training data provided in the specification of the present application, an initial image data set is first obtained, and the initial image data set includes several initial images. For each initial image in the initial image data set, the area to be modified in the initial image is obtained; for each area to be modified, the area to be modified is modified according to a preset automatic modification method to obtain a target image, and a training data set for training an image detection model is constructed based on each target image and the initial image corresponding to each target image. The preset automatic modification method includes at least one method of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification through a diffusion model. When modifying the area to be modified, a fully automatic modification method is adopted, without the need for manual participation, and the efficiency of constructing training data is high. And because the training data set contains a variety of methods for automatically modifying images, then, when the image detection model trained according to the training data set is used to detect images, the accuracy of the detection result of image tampering obtained is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present specification and constitute a part of the present specification. The illustrative embodiments and descriptions of the present specification are used to explain the present specification and do not constitute an improper limitation on the present specification. In the drawings: Figure 1 A flowchart of a method for constructing model training data provided in this application specification; Figure 2 An image detection flow chart provided in this application specification; Figure 3 A schematic diagram of a device for constructing model training data provided in this application specification; Figure 4 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0013] To make the purpose, technical solutions, and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] The execution entities of this application specification are various computing devices that can execute the method for constructing model training data provided in this application specification, such as a single server, a server cluster, etc. For the sake of convenience, the server is used as the execution entity for explanation.
[0015] The following describes in detail the technical solutions provided by each embodiment of this application specification in conjunction with the accompanying drawings.
[0016] Figure 1 This is a flowchart of a method for constructing model training data provided in this application specification, specifically including steps S100~S106.
[0017] S100: Acquire an initial image dataset, where the initial image dataset includes a number of initial images.
[0018] It should be noted that the initial image dataset includes at least one of a general image dataset, a tampered image dataset and a business image dataset. The general image dataset includes initial images corresponding to at least two business scenarios and the categories of objects included in each initial image. The tampered image dataset includes tampered images, initial images corresponding to the tampered images and the annotated areas of the initial images. The tampered images are obtained by modifying the annotated areas of the initial images; the business image dataset includes initial images corresponding to one business scenario and the categories of objects included in each initial image.
[0019] General image datasets may include the COCO dataset and the Object365 model. Business scenarios may include financial scenarios, e-commerce scenarios, and the like, and this application specification does not limit this. The initial image corresponding to the business scenario refers to a voucher image that appears in the business scenario. For example, in a financial scenario, the initial image may be an invoice image.
[0020] It is understood that the general image dataset is highly generalizable and includes initial images from a variety of business scenarios. An image detection model trained with this general image dataset can identify object categories for voucher images in a variety of business scenarios. The tampered image dataset can be obtained by modifying the initial images in the general image dataset, specifically by modifying the data in the annotated regions of the initial images to obtain tampered images. The tampered image dataset can also be any existing dataset that includes tampered and initial images, and this specification is not limited thereto. Training the image detection model with this tampered image dataset enables the image detection model to detect voucher images in various business scenarios and obtain a result indicating whether the voucher images have been tampered with. In this specification, multiple business image datasets may be included, each of which may include initial images corresponding to a business scenario and the categories of objects included in each initial image. Training the image detection model with a business image dataset provides strong business-specificity, thereby improving the accuracy of detection results for voucher image data for the business scenario corresponding to that business image dataset. For each business image dataset, the business image dataset may be obtained by classifying the general image dataset according to the business scenario.
[0021] S102: For each initial image in the initial image data set, determine a region to be modified in the initial image.
[0022] Specifically, the server can use optical character recognition (OCR) technology to determine the area where the target data is located in the initial image, and determine the area where the target data is located as the area to be modified. It can also be determined through an image recognition model. Specifically, the server can input the initial image into the image recognition model, obtain the area where the target data is located as output by the image recognition model, and determine the area where the target data is located as the area to be modified. The server can also determine the area to be modified in the initial image through other methods, which are not limited in this application specification. Each initial image may include at least one area to be modified.
[0023] When determining the area to be modified in the initial image, it is more efficient to choose OCR technology or artificial intelligence model that does not require human intervention, thereby improving the efficiency of the training data set for building the model and reducing human resources.
[0024] It should be noted that the target data includes a target object and / or target text. The target data may be preset so that it can be subsequently modified. This allows an image detection model trained using a training dataset including the modified target data to detect whether the target data in the image to be detected has been tampered with, thereby improving the pertinence of the image detection model and thereby increasing the accuracy of the detection results obtained by the image detection model.
[0025] It's understandable that target data may differ for different business scenarios. For a financial scenario, if the initial image is an invoice, the target data may be the target text within the invoice, which is part of the pre-defined target data. For a pet registration scenario, the target data may be the pet in the captured image, such as a cat or dog.
[0026] If the OCR technology is used to determine the area to be modified in the invoice image, the text in the invoice image can be recognized first, and the area where the recognized text is located can be determined as the area to be modified.
[0027] Similar to the tampered image dataset, the initial image in the business image dataset may also include the business tampered image, the initial image corresponding to the business tampered image, and the annotated area of the initial image. Therefore, when determining the area to be modified in the initial image, the annotated area of the initial image in the business image dataset may also be determined as the area to be modified in the initial image. In other words, if the initial image includes the annotated area, the annotated area is determined as the area to be modified.
[0028] In other words, even if the initial image dataset already includes a tampered image, the original image corresponding to the tampered image, and the annotated area of the original image, OCR technology, image recognition models, and other methods can be used to determine the target data area of the initial image in the initial image dataset, and the target data area can be determined as the area of the initial image to be modified. In this case, the area of the initial image to be modified includes not only the original annotated area, but also the identified target data area.
[0029] By using an OCR or image recognition model, the model automatically detects the areas to be modified within the initial image, particularly where text or other key information resides. This allows the model to accurately locate the areas to be modified, providing a foundation for subsequent modifications. Furthermore, by pre-setting target data, the model can automatically adjust the modified areas based on specific business scenarios, effectively avoiding the limitations of traditional methods.
[0030] S104: For each area to be modified, modify the area to be modified according to a preset automated modification method to obtain a target image, wherein the preset automated modification method includes at least one method selected from text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification through a diffusion model.
[0031] In this application, when the preset automated modification method is text replacement, the text in the to-be-modified region is replaced with text in other to-be-modified regions other than the to-be-modified region, where the other to-be-modified regions and the to-be-modified region belong to the same initial image. Specifically, a pair of target texts in the to-be-modified regions of similar size can be randomly selected in the same initial image and the positions of the target texts can be swapped.
[0032] For example, the first row and the second row in the initial image A are swapped, and the sizes of the areas to be modified corresponding to the first row and the second row are similar.
[0033] When the preset automatic modification method is image mirroring, the object in the area to be modified is flipped and / or mirrored, wherein the flipping can be done in any direction.
[0034] When the preset automated modification method is information erasure, the target data in the area to be modified will be erased, and the background area in the area to be modified will be retained. Specifically, the target data in the area to be modified can be erased based on OpenCV technology. Through the erasure method based on OpenCV technology, certain content in the area to be modified can be erased, such as text or image information, leaving only the background color, so that a certain key information disappears to achieve the purpose of modification. This application specification does not limit the amount of erased information. It is understandable that part of the target data in the area to be modified can be erased, or all the target data in the area to be modified can be erased.
[0035] When the preset automated modification method is cross-image modification, the image to which the to-be-modified region belongs is determined as the first image. A target replacement region is determined in images other than the first image, and the target replacement region is fused into the to-be-modified region using the Poisson fusion method. Specifically, the server may perform random scaling, rotation, and flipping operations on the target replacement region, and then perform image fusion using the Poisson fusion method.
[0036] Bosong Fusion effectively blends the copied target replacement area seamlessly with the original image, making the modification difficult to detect. This method further enhances the concealment and authenticity of the modified target image, thereby increasing the difficulty of image detection. However, by fully simulating this advanced modification technique during the training phase, the image detection model can learn to identify highly concealed modification traces, effectively overcoming the detection difficulties caused by subtle artifacts.
[0037] When the preset automated modification method is noise addition, the variational autoencoder extracts the image features of the area to be modified. Gaussian noise is added to these features to obtain the noisy image features. These noisy image features are then input into the corresponding decoder of the variational autoencoder to obtain the decoder output image. By adding noise to the image features, the image features are disrupted by the noise and cannot be displayed after decoding, thus achieving the purpose of image modification.
[0038] When the preset automated modification method is modification through a diffusion model, the server can first obtain a text prompt word. The text prompt word can be preset, specifically the modification type of the area to be modified, and the modification type includes information erasure and information replacement. When obtaining the text prompt word, one can be randomly obtained from the preset prompt words as the text prompt word for injection into the diffusion model. Afterwards, the server can input the text prompt word into the text encoder in the diffusion model to obtain the text features output by the text encoder. The image features of the area to be modified can also be extracted through an image encoder, and the image features and text features are determined as generation conditions. The generation conditions are injected into a pre-trained diffusion model so that the diffusion model denoises the preset standard noise image, and the denoising is used to achieve redrawing or erasing the target information in the area to be modified. Then, when the modification type is information erasure, the denoising is used to achieve erasure of the target information in the area to be modified. When the modification type is information replacement, the denoising is used to achieve redrawing of the area to be modified.
[0039] The modification methods provided in this application specification are all automated modifications and do not require human intervention, which greatly reduces the workload of technical personnel, saves labor costs, and improves the efficiency of constructing training data sets.
[0040] It should be noted that the aforementioned multiple preset automated modification methods can be used in conjunction with each other. By integrating multiple modification methods, the diversity and complexity of the training data are ensured. These modification methods can not only simulate traditional image tampering, but also cover complex forged images generated by diffusion models, thereby training image detection models with greater generalization capabilities. Through diverse tampering methods, the model can learn the details and signs of various image forgeries, including subtle visual artifacts and pixel-level differences, improving the detection capabilities of the image detection model and obtaining more accurate detection results.
[0041] In particular, using a diffusion model to modify the target image and generate the target image can better simulate complex tampering scenarios, increase the realism of the image, and improve the model's sensitivity to subtle tampering. Compared to traditional methods based on manual image modification, automated modification methods can generate more diverse and realistic training data, thus avoiding the limitations of dataset coverage.
[0042] S106: Constructing a training data set for training an image detection model according to each target image and the initial image corresponding to each target image.
[0043] In this application, the target image and the initial image corresponding to the target image can be used as training data for the image detection model. In this case, multiple training data can constitute the training data set. Of course, the area to be modified in the initial image can also be marked, and the marked area to be modified can also be used as training data to prompt the image detection model to modify the target image area when training the image detection model, thereby accelerating model training.
[0044] based on Figure 1 The method for constructing model training data shown in the figure first obtains an initial image dataset, which includes several initial images. For each initial image in the initial image dataset, the area to be modified in the initial image is obtained; for each area to be modified, the area to be modified is modified according to a preset automatic modification method to obtain a target image, and a training dataset for training an image detection model is constructed based on each target image and the initial image corresponding to each target image. The preset automatic modification method includes at least one of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification through a diffusion model. When modifying the area to be modified, a fully automatic modification method is adopted, which does not require manual participation, and the efficiency of constructing training data is high. And because the training dataset contains a variety of automatic image modification methods, the image detection model trained according to the training dataset is used to detect images, and the accuracy of the image tampering detection result is high.
[0045] Before executing step S104, the server may also determine whether to modify the area to be modified in the initial image based on the type of target data in the initial image. In other words, the server may choose to perform an automated modification operation on specific target data, and the specific target data may correspond to a business scenario.
[0046] Specifically, when the business scenario is a financial one, the server may determine whether to modify the to-be-modified area in the initial image based on whether the target data region contains text after determining the target data. If the target data region contains only the target object but no text, then the business image dataset may contain an error, possibly including a non-financial voucher image. In this case, the server may choose not to modify the to-be-modified area in the initial image.
[0047] After executing step S106, a training data set for training the image detection model is obtained. Based on this, this application specification also provides a method for training an image detection model. The execution subject of the image detection model training can be any electronic device that can train the model, such as a server, a personal computer, etc. For ease of explanation, this application specification uses a server as the execution subject.
[0048] Obtain a training dataset. The training dataset can be obtained through steps S100-S106. The training dataset includes several target images and an initial image corresponding to each target image. The initial image corresponding to each target image belongs to at least one of a general image dataset, a tampered image dataset, and a business image dataset. It will be understood that the target image can be obtained by executing steps S102-S104 for each initial image in the general image dataset, the tampered image dataset, and the business image dataset. Therefore, the target image in the training dataset may come from at least one of the general image dataset, the tampered image dataset, and the business image dataset.
[0049] First, an image detection model is pre-trained using an initial image corresponding to the target image included in a general image dataset to obtain an initial image detection model. Specifically, the initial image is input into the image detection model, and the image detection model outputs a predicted object category in the initial image. A classification loss is determined based on the object category in the initial image included in the general image dataset and the predicted object category. The image detection model is then pre-trained based on the classification loss to obtain the initial image detection model. At this point, the image detection model is capable of recognizing target data.
[0050] Afterwards, the server uses the initial image and the target image included in the tampered image dataset that correspond to the target image to obtain a general image detection model. Specifically, any one of the initial image and the target image included in the tampered image dataset that correspond to the target image is used as an input image, and the input image is input into the initial image detection model. The initial image detection model detects the input image to obtain a first detection result. Based on the first detection result and the category of the input image, a first detection loss is determined. Based on the first detection loss, the initial image detection model is trained to obtain a general image detection model. The category of the input image includes any one of the target image and the initial image, and the first detection result includes any one of whether the input image has been tampered with or not tampered with.
[0051] It should be noted that the initial image used to train the initial image detection model may be the same as or different from the initial image used to train the image detection model.
[0052] In order to further improve the detection accuracy of the general image detection model for credential images corresponding to a certain business scenario, the server can also use the initial image and target image corresponding to the target image included in the business image dataset to fine-tune the general image detection model to obtain a business-specific image detection model. Specifically, the server can input any of the initial image and target image corresponding to the target image included in the business image dataset into the general image detection model. The general image detection model detects the input image to obtain a second detection result. Based on the second detection result and the category of the input image, a second detection loss is determined. Based on the second detection loss, the general image detection model is trained to obtain a business-specific image detection model. The category of the input image includes any one of the target image and the initial image, and the second detection result includes any one of whether the input image has been tampered with or not tampered with.
[0053] It should be noted that the initial image and target image corresponding to the target image used to train the general image detection model may or may not be the same as the initial image and target image corresponding to the target image used to train the initial image detection model. Because the business image dataset includes initial images corresponding to a business scenario and the categories of objects included in each initial image, the model trained based on this business image dataset is more targeted and has stronger detection capabilities for the business scenario corresponding to this business image dataset. In other words, the detection results obtained for images of this business scenario are more accurate.
[0054] After the image detection model is trained, the image detection business can be performed through the image detection model. Figure 2 This is an image detection flow chart provided in this application specification, such as Figure 2shown.
[0055] Figure 2 The image detection model in can be any one of a general image detection model and a business-specific image detection model.
[0056] Users can communicate with the server through their terminal, requesting the server to provide them with business services. Business services can be services corresponding to any business scenario, such as tax information registration, pet information registration, merchant store information authentication, and so on. The user first sends a business request to the server through the terminal, including a request for a business service provided by the server. The server receives the business request from the terminal and returns a credential acquisition request to the terminal. The terminal receives the credential acquisition request and displays the credential acquisition interface to the user. The user operates the credential acquisition node, such as clicking the Add Credential button in the credential acquisition interface to upload a credential image. In response to the user's credential upload operation, the terminal sends the uploaded credential image to the server. The server receives the credential image, identifies it as the image to be detected, and inputs it into the image detection model. The image detection model detects the image to be detected and obtains a detection result, which can indicate whether the image to be detected has been tampered with or not. The image detection model sends this detection result to the server, which then determines whether to provide the business service to the terminal based on the detection result. Specifically, when the detection result shows that the image to be detected has not been tampered with, the terminal is provided with business services. When the detection result shows that the image to be detected has been tampered with, the terminal may be denied services and a prompt message such as "credential image unqualified" may be sent to the terminal to prompt the user to upload a real credential image. The aforementioned image detection model may be a general image detection model or a business-specific image detection model. If it is a business-specific image detection model, after receiving the credential image, the server may first determine the business scenario to which the credential image belongs, and then determine the corresponding business-specific image detection model based on the business scenario. The business scenario to which the credential image belongs can be determined by the business service required by the user.
[0057] This application specification also provides a method for training a diffusion model, which is still described with the server as the execution entity.
[0058] The diffusion model in this application specification includes a text encoder, which is used to encode input text to obtain text features.
[0059] When training the diffusion model, the server first obtains a training image set, which includes sample training images, sample tampered images, sample mask images and sample prompt text. The sample prompt text includes a preset modification type, which is used to guide the diffusion model to generate tampered images that meet the modification type.
[0060] For example, type 1: text modification, which changes text A to text B; type 2: information erasure: which erases text or objects and retains background information.
[0061] The sample tampered image is obtained by performing a modification operation corresponding to a preset modification type on the sample training image using the sample prompt text by the tampering model. The tampering model is any model capable of modifying the sample training image. When using the tampering model to modify the sample training image, the preset modification type can be used as prompt information. The prompt information and the sample training image are input into the tampering model, and the sample tampered image is output by the tampering model. In addition, the sample mask image is obtained by processing the to-be-modified area in the sample training image into a mask.
[0062] Next, the sample training image is denoised to obtain a noisy sample training image; the sample mask image and the noisy sample training image are respectively input into an image encoder to obtain sample image features of the sample training image and sample image features of the sample mask image, respectively output by the image encoder; and the sample prompt text is input into a text encoder in the diffusion model to obtain sample text features output by the text encoder. It should be noted that the modification type included in the sample prompt text corresponding to each sample training image is consistent with the preset modification type included in the prompt information, wherein the prompt information is used to obtain the sample tampered image corresponding to the sample training image.
[0063] Next, the sample text features and each sample image feature are used as generation conditions. These generation conditions are then injected into the diffusion model, causing it to denoise the preset standard noise image to produce the target tampered image. Finally, the diffusion model is trained based on the target tampered image and the sample tampered images. Specifically, the difference between the target tampered image and the sample tampered images is determined, and the diffusion model is trained with the goal of reducing this difference.
[0064] It should be noted that the diffusion model also includes an image erasing module, an image editing module, an adaptation network and a denoising network. Among them, the image erasing module can be a Flux Fill model for realizing the information erasing function, the image editing module can be a Step1X model for realizing the information replacement function, and the denoising network can be a U-Net. The image editing module includes a character positioning module, and the character positioning model can be a fine-tuned Qwen2.5VL model. It can be understood that, unlike the ordinary diffusion model, the diffusion model in this application specification includes a large model. Therefore, the diffusion model in this application specification can also be referred to as a diffusion large model.
[0065] Furthermore, the adaptation network includes a high-rank adaptation network and a low-rank adaptation network. The rank of the high-rank adaptation network is greater than that of the low-rank adaptation network. The text encoder is connected to the image editing module, which is connected to the denoising network. The high-rank adaptation network is embedded in the deep network of the denoising network, and the low-rank adaptation network is embedded in the shallow network of the denoising network. The image erasure module is connected to the denoising network.
[0066] A deep network is one with a preset number of layers, while a shallow network is one with a preset number of layers. For example, if the diffusion model has 100 layers and the preset range is 40-60, then a deep network is one with 40-60 layers, while a shallow network is one with 1-39 layers and 61-100 layers.
[0067] By assigning an adaptation network with different rank values to each denoising network, the stability of the global modification area is guaranteed, while the diffusion model can generate diverse details.
[0068] When the modification type in the sample prompt text is information modification, the server can train Flux Fill to optimize the erasure function, adopt multi-Lora fine-tuning technology with a hierarchical rank allocation mechanism, and adjust the U-Net structure of the diffusion model to adapt to the business scenario tasks. Specifically, different LoRa rank values and LoRa models are designed for different U-Net modules, such as DownBlock, CrossAttn, and UpBlock. The deep network uses a high-rank LoRa model (e.g., r=64) to capture the global semantics of the area to be modified and ensure the stability of the global modification area. The shallow layer uses a low-rank LoRa model (e.g., r=16) to focus on local details and preserve the diversity of model-generated details.
[0069] When the modification type in the sample prompt text is information modification, the server can train the Step1X model to optimize the editing function. In addition to using the multi-Lora fine-tuning technology of the hierarchical rank allocation mechanism mentioned above to enhance the diffusion model's ability to recognize characters and text, the fine-tuned Qwen2.5VL model is also embedded in Step1X to ensure the accuracy of character positioning. The text encoder in the diffusion model is also trained to ensure the accuracy of the prompt information and achieve the expected editing effect. Finally, the new Step1X can more accurately modify specific locations to the expected data, such as numbers, characters, etc. based on the prompt word.
[0070] The method for constructing model training data provided in this application specification can achieve the following effects: 1. Improving the model's generalization capabilities: By modifying the initial image through various modification methods, the image detection model's generalization capabilities are significantly improved. The image detection model can not only identify traditional forgeries, but also handle highly realistic target images from artificial intelligence generated content (AIGC) generated by the diffusion model. This allows the image detection model to maintain high accuracy across different types of manipulated images.
[0071] 2. Reduced computational complexity: Traditional tampering detection methods are computationally complex and resource-intensive. This solution, by utilizing automated modification methods, can reduce the number of manually designed tampering types. Furthermore, the automated image modification process improves training efficiency, enabling image detection model training to converge to higher performance more quickly.
[0072] 3. Enhanced detection accuracy and robustness: Automated modification methods can fully cover a variety of tampering scenarios during model training, capturing subtle traces of various image tampering. This makes the trained image detection model more robust, enabling it to more accurately detect images modified by various tampering methods in real-world scenarios.
[0073] 4. Greater practicality and flexibility: The method can be trained on a variety of datasets and automatically generates regions to be modified, offering high flexibility. Whether processing publicly available tampered datasets or custom datasets for specific business scenarios, the image detection model can automatically adapt and optimize, making the method more applicable across a wide range of application scenarios.
[0074] The above is a method for constructing model training data provided in one or more embodiments of this application specification. Based on the same idea, this application specification also provides a corresponding device for constructing model training data, such as Figure 3 As shown, the device includes: An initial image data set acquisition module 300 is used to acquire an initial image data set, wherein the initial image data set includes a plurality of initial images; A to-be-modified region determining module 302 is configured to determine, for each initial image in the initial image data set, a to-be-modified region in the initial image; A modification module 304 is configured to modify each region to be modified according to a preset automated modification method to obtain a target image, wherein the preset automated modification method includes at least one of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification using a diffusion model; The training data set construction module 306 is used to construct a training data set for training the image detection model according to each target image and the initial image corresponding to each target image.
[0075] Optionally, the initial image dataset includes at least one of a general image dataset, a tampered image dataset and a business image dataset, the general image dataset includes initial images corresponding to at least two business scenarios and the category of objects included in each initial image, the tampered image dataset includes a tampered image, an initial image corresponding to the tampered image and an annotated area of the initial image, and the tampered image is obtained by modifying the annotated area of the initial image; the business image dataset includes an initial image corresponding to a business scenario and the category of objects included in each initial image.
[0076] Optionally, the to-be-modified region determining module 302 is specifically configured to determine the region where the target data is located in the initial image by using optical character recognition technology; or Inputting the initial image into an image recognition model to obtain an area where target data output by the image recognition model is located, wherein the target data includes a target object and / or target text; The area where the target data is located is determined as the area to be modified.
[0077] Optionally, the to-be-modified region determining module 302 is specifically configured to determine the marked region as the to-be-modified region if the initial image includes a marked region.
[0078] Optionally, the modification module 304 is specifically configured to replace the text in the to-be-modified region with text in other to-be-modified regions except the to-be-modified region when the preset automatic modification method is text replacement; When the preset automatic modification method is image mirroring, performing flipping and / or mirroring operations on the object in the area to be modified; When the preset automatic modification method is information erasure, the target data in the area to be modified will be erased, and the background area in the area to be modified will be retained; When the preset automatic modification method is cross-image modification, the image to which the area to be modified belongs is determined as a first image, a target replacement area is determined in images other than the first image, and the target replacement area is fused into the area to be modified using a Poisson fusion method; When the preset automatic modification method is noise addition, extracting image features of the area to be modified through a variational autoencoder, adding Gaussian noise to the image features to obtain noisy image features, inputting the noisy image features into a decoder corresponding to the variational autoencoder, and obtaining an image output by the decoder; When the preset automatic modification method is modification through a diffusion model, a text prompt word is obtained, the text prompt word is input into a text encoder in the diffusion model, and text features output by the text encoder are obtained. The image features of the area to be modified are extracted through an image encoder, and the image features and the text features are determined as generation conditions. The generation conditions are injected into a pre-trained diffusion model so that the diffusion model denoises the preset standard noise image.
[0079] Optionally, the training dataset includes a plurality of target images and an initial image corresponding to each target image, and the initial image corresponding to one target image belongs to at least one of a general image dataset, a tampered image dataset, and a business image dataset; The apparatus further includes: an image detection model training module 308, configured to pre-train an image detection model using an initial image corresponding to a target image included in the general image dataset to obtain an initial image detection model; Using the initial image corresponding to the target image and the target image included in the tampered image dataset, the initial image detection model is trained to obtain a general image detection model; The general image detection model is fine-tuned and trained using the initial image corresponding to the target image and the target image included in the business image dataset to obtain a business-specific image detection model.
[0080] Optionally, the diffusion model includes a text encoder; The apparatus further includes: a diffusion model training module 310, configured to obtain a training image set, wherein the training image set includes a sample tampered image, a sample mask image, and a sample prompt text, wherein the sample tampered image is obtained by performing a modification operation corresponding to a preset modification type on the sample training image by the tampering large model based on the sample prompt text, and the sample mask image is obtained by processing a to-be-modified region in the sample training image into a mask; Adding noise to the sample training image to obtain the noisy sample training image; Inputting the sample mask image and the sample training image after noise addition into the image encoder respectively, obtaining sample image features of the sample training image and the sample mask image respectively output by the image encoder; and inputting the sample prompt text into the text encoder in the diffusion model, obtaining sample text features output by the text encoder; Determining the sample text features and each sample image feature as generation conditions, and injecting the generation conditions into the diffusion model so that the diffusion model denoises the preset standard noise image to obtain a target tampered image; The diffusion model is trained according to the target tampered image and the sample tampered image.
[0081] Optionally, the diffusion model also includes an image erasure module, an image editing module, an adaptation network and a denoising network; the image editing module includes a character positioning module, the adaptation network includes a high-rank adaptation network and a low-rank adaptation network, and the rank value of the high-rank adaptation network is greater than the rank value of the low-rank adaptation network; the text encoder is connected to the image editing module, the image editing module is connected to the denoising network, the high-rank adaptation network is embedded in the deep network in the denoising network, the low-rank adaptation network is embedded in the shallow network in the denoising network, and the image erasure module is connected to the denoising network.
[0082] This application also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 The method for constructing the provided model training data.
[0083] This application also provides Figure 4 The structural diagram of the electronic device shown in FIG. Figure 4 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The method for constructing the model training data. Of course, in addition to software implementation, this application specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0084] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0085] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0086] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0087] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0088] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0090] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0092] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0093] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0094] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0096] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0098] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0099] The foregoing is merely an example of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations may be made to the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for constructing model training data, the method comprising: Acquire an initial image data set, where the initial image data set includes a plurality of initial images; For each initial image in the initial image data set, determining a region to be modified in the initial image; For each area to be modified, modify the area to be modified according to a preset automated modification method to obtain a target image, wherein the preset automated modification method includes at least one method selected from the group consisting of text replacement, image mirroring, information erasure, cross-image modification, noise addition, and modification using a diffusion model; A training data set for training an image detection model is constructed based on each target image and an initial image corresponding to each target image.
2. The method according to claim 1, wherein the initial image dataset includes at least one of a general image dataset, a tampered image dataset, and a business image dataset, the general image dataset including initial images corresponding to at least two business scenarios and the category of objects included in each initial image, the tampered image dataset including a tampered image, an initial image corresponding to the tampered image, and an annotated area of the initial image, the tampered image being obtained by modifying the annotated area of the initial image; and the business image dataset including an initial image corresponding to one business scenario and the category of objects included in each initial image.
3. The method of claim 1, wherein determining the area to be modified in the initial image comprises: Using optical character recognition technology, determine the area where the target data is located in the initial image; or Inputting the initial image into an image recognition model to obtain an area where target data output by the image recognition model is located, wherein the target data includes a target object and / or target text; The area where the target data is located is determined as the area to be modified.
4. The method of claim 2, wherein determining the area to be modified in the initial image comprises: If the initial image includes a marked area, the marked area is determined as the area to be modified.
5. The method according to claim 1, wherein for each area to be modified, modifying the area to be modified according to a preset automated modification method, specifically comprises: When the preset automatic modification method is text replacement, the text in the area to be modified is replaced with text in other areas to be modified except the area to be modified; When the preset automatic modification method is image mirroring, performing flipping and / or mirroring operations on the object in the area to be modified; When the preset automatic modification method is information erasure, the target data in the area to be modified will be erased, and the background area in the area to be modified will be retained; When the preset automatic modification method is cross-image modification, the image to which the area to be modified belongs is determined as a first image, a target replacement area is determined in images other than the first image, and the target replacement area is fused into the area to be modified using a Poisson fusion method; When the preset automatic modification method is noise addition, extracting image features of the area to be modified through a variational autoencoder, adding Gaussian noise to the image features to obtain noisy image features, inputting the noisy image features into a decoder corresponding to the variational autoencoder, and obtaining an image output by the decoder; When the preset automatic modification method is modification through a diffusion model, a text prompt word is obtained, the text prompt word is input into a text encoder in the diffusion model, and text features output by the text encoder are obtained. The image features of the area to be modified are extracted through an image encoder, and the image features and the text features are determined as generation conditions. The generation conditions are injected into a pre-trained diffusion model so that the diffusion model denoises the preset standard noise image.
6. The method of claim 2, wherein the training dataset includes a plurality of target images and an initial image corresponding to each target image, wherein the initial image corresponding to each target image belongs to at least one of a general image dataset, a tampered image dataset, and a business image dataset; the method further comprising: Pre-training the image detection model using the initial image corresponding to the target image included in the general image dataset to obtain an initial image detection model; Using the initial image corresponding to the target image and the target image included in the tampered image dataset, the initial image detection model is trained to obtain a general image detection model; The general image detection model is fine-tuned and trained using the initial image corresponding to the target image and the target image included in the business image dataset to obtain a business-specific image detection model.
7. The method of claim 1, wherein the diffusion model comprises a text encoder; Training the diffusion model specifically includes: Obtaining a training image set, the training image set including a sample training image, a sample tampered image, a sample mask image, and a sample prompt text, wherein the sample tampered image is obtained by the tampering large model performing a modification operation corresponding to a preset modification type on the sample training image based on the sample prompt text, and the sample mask image is obtained by processing a to-be-modified area in the sample training image into a mask; Adding noise to the sample training image to obtain the noisy sample training image; Inputting the sample mask image and the sample training image after noise addition into an image encoder respectively, obtaining sample image features of the sample training image and sample image features of the sample mask image respectively output by the image encoder; and inputting the sample prompt text into a text encoder in the diffusion model, obtaining sample text features output by the text encoder; Determining the sample text features and each sample image feature as generation conditions, and injecting the generation conditions into the diffusion model so that the diffusion model denoises the preset standard noise image to obtain a target tampered image; The diffusion model is trained according to the target tampered image and the sample tampered image.
8. According to the method as claimed in claim 7, the diffusion model also includes an image erasure module, an image editing module, an adaptation network and a denoising network; the image editing module includes a character positioning module, the adaptation network includes a high-rank adaptation network and a low-rank adaptation network, and the rank value of the high-rank adaptation network is greater than the rank value of the low-rank adaptation network; the text encoder is connected to the image editing module, the image editing module is connected to the denoising network, the deep network in the denoising network is embedded with the high-rank adaptation network, the shallow network in the denoising network is embedded with the low-rank adaptation network, and the image erasure module is connected to the denoising network.
9. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
False identity card back photo identification method based on feature extraction
CN115294437A
Image tampering detection method and system for generative model image editing
CN117690005A
Picture tampering detection model construction method and picture tampering detection method
CN118644767A
ACDSee image tampering positioning method and device based on large model, terminal and medium
CN119723308A
Cited By
Image retrieval method, device and equipment
CN120929633A
An image retrieval method, apparatus and device
CN120929633B
Image editing method and device, equipment and storage medium
CN121033227A
Image editing method, device, apparatus and storage medium
CN121033227B