Data set construction method and system, electronic equipment and storage medium

By dividing, classifying and extracting real images and synthetic images, a high-quality data set is constructed, which solves the problem of low accuracy in identifying synthetic images by existing detection models, and significantly improves detection accuracy.

CN120047958AActive Publication Date: 2025-05-27ZHONGDIAN DATA IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510519280.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing detection models have low accuracy when identifying synthetic images generated by the diffusion model and lack effective dataset construction methods to improve detection accuracy.

Method used

A dataset construction method is proposed, by dividing the real images into different sets and generating synthetic images based on these sets, the synthetic images are further classified and extracted to construct the target dataset. The method includes dividing the real image into first and second real image sets, generating a synthetic image, dividing the synthetic image into a set with defects and without defects, and extracting images equally in these sets, and ultimately constructing a target data set based on the real image and the synthetic image.

Benefits of technology

By constructing a high-quality dataset containing all features of the synthetic image, the detection model's recognition accuracy of the synthetic image is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047958A_ABST
    Figure CN120047958A_ABST
Patent Text Reader

Abstract

The invention discloses a data set construction method and system, electronic equipment and a storage medium, and relates to the technical field of model training, and the disclosed data set construction method comprises the steps: dividing each real image into a first real image set and a second real image set, and obtaining each composite image based on each real image in the first real image set; dividing each composite image into a first composite image set and a second composite image set, and classifying the composite images in the first composite image set; equivalently extracting images of each category from the first composite image set, and extracting images from the second composite image set to obtain a third composite image set; and constructing a target data set based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target data set. The invention aims to solve the technical problem of how to improve the detection and recognition accuracy of a detection model on a synthetic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model training, and particularly to a method and system for constructing a dataset, an electronic device, and a storage medium. Background Art

[0002] Images, as a key form of information carrier, play an important role in the information transmission process. Currently, there are various ways to obtain images. According to the generation characteristics of images, they can generally be classified into two categories: real images and synthetic images. Among them, real images are images that have not been tampered with by humans; synthetic images are images generated by using specific models and algorithms.

[0003] Obviously, since synthetic images are constructed through artificial intervention and model building, compared with real images, there are risks and potential hazards of being misused or even maliciously abused. In view of this, it is crucial to determine synthetic images from real images. At present, detection models are widely used in the field of synthetic image detection and are one of the main detection tools. However, with the technological evolution, synthetic images generated by diffusion models exhibit new characteristics. Synthetic images generated by diffusion models often no longer carry traditional recognition features such as artifacts, resulting in a low detection accuracy of detection models for synthetic images.

[0004] Therefore, how to improve the detection accuracy of detection models for synthetic images is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] The main purpose of this application is to provide a method and system for constructing a dataset, an electronic device, and a storage medium, aiming to solve the technical problem of how to improve the detection accuracy of detection models for synthetic images.

[0006] To achieve the above object, this application proposes a method for constructing a dataset, the method including: Dividing each real image into a first real image set and a second real image set, and obtaining each synthetic image based on each real image in the first real image set; Dividing each synthetic image into a first synthetic image set and a second synthetic image set, and classifying the synthetic images in the first synthetic image set, where the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; Equally extracting images of each category in the first synthetic image set, and extracting images in the second synthetic image set to obtain a third synthetic image set; Construct a target dataset based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target dataset.

[0007] In one embodiment, before the step of dividing each real image into a first real image set and a second real image set, the method further includes: Obtain the original real images in each data source, and filter the original real images through a preset image size, image clarity range, aspect ratio range, and redundant information to obtain each intermediate real image; Deduplicate each of the intermediate real images, and use each of the intermediate real images obtained after deduplication as each of the real images.

[0008] In one embodiment, before the step of deduplicating each of the intermediate real images and using each of the intermediate real images obtained after deduplication as each of the real images, the method further includes: Calculate the similarity of each text-image pair composed of each of the intermediate real images and the text corresponding to each of the intermediate real images, obtain the similarity corresponding to each text-image pair, and screen out each target real image from each of the intermediate real images based on each similarity; The step of deduplicating each of the intermediate real images and using each of the intermediate real images obtained after deduplication as each of the real images includes: Deduplicate each of the target real images, and use each of the target real images obtained after deduplication as each of the real images.

[0009] In one embodiment, before the step of dividing each real image into a first real image set and a second real image set, the method further includes: Obtain the text corresponding to the original real images in each data source, and filter the text corresponding to the original real images through a preset text size threshold to obtain each intermediate text; Deduplicate each of the intermediate texts, and use the original real images corresponding to each of the intermediate texts obtained after deduplication as each of the real images.

[0010] In one embodiment, before the step of dividing each real image into a first real image set and a second real image set, the method further includes: Filter the text corresponding to each real image through a preset sensitive word detector, and filter each real image through a preset NSFW image detector to obtain a preprocessed real image set; The step of dividing each real image into a first real image set and a second real image set includes: Dividing the preprocessed real image set into a first real image set and a second real image set.

[0011] In one embodiment, the step of equally extracting images of each category from the first synthetic image set and extracting images from the second synthetic image set to obtain a third synthetic image set includes: Respectively extracting a first preset number of images from each category in the first synthetic image set to obtain a fourth synthetic image set; Determine the total number of defect-free images according to the number of images in the fourth synthetic image set and a preset coefficient value, and extract images from the second synthetic image set according to the total number of defect-free images to obtain a fifth synthetic image set; Combine the fourth synthetic image set and the fifth synthetic image set to obtain a third synthetic image set.

[0012] In one embodiment, the step of obtaining each synthetic image based on each real image in the first real image set includes: Input each real image in the first real image set into an image-to-image model to obtain each synthetic image; Or, Input the text corresponding to each real image in the first real image set into a text-to-image model to obtain each synthetic image; Or, Input each real image in the first real image set into an image-to-text model to obtain the material text corresponding to each real image in the first real image set; Perform text enhancement on each material text, and input each text-enhanced material text into a text-to-image model to obtain the synthetic image corresponding to each real image in the first real image set.

[0013] In addition, to achieve the above object, the present application also proposes a dataset construction system, and the dataset construction system includes: A first image processing module, configured to divide each real image into a first real image set and a second real image set, and obtain each synthetic image based on each real image in the first real image set; A second image processing module, configured to divide each synthetic image into a first synthetic image set and a second synthetic image set, and classify the synthetic images in the first synthetic image set, where the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; A third image processing module, configured to equally extract images of each category from the first synthetic image set and extract images from the second synthetic image set to obtain a third synthetic image set; A dataset integration module, configured to construct a target dataset based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target dataset.

[0014] In addition, to achieve the above object, the present application further provides an electronic device, where the electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the dataset construction method as described above.

[0015] In addition, to achieve the above object, the present application further provides a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the dataset construction method as described above.

[0016] In the embodiment of the present application, by dividing each real image into a first real image set and a second real image set, and obtaining synthetic images based on the real images in the first real image set, real images and synthetic images for constructing a dataset can be obtained; then dividing the synthetic images into a first synthetic image set and a second synthetic image set, and classifying the synthetic images in the first synthetic image set, the features of defective images of each category can be easily counted through the method of defective image classification; then equally extracting images of each category from the first synthetic image set and extracting images from the second synthetic image set to obtain a third synthetic image set can further subdivide the defective images, so as to facilitate the extraction of synthetic images with different defective features, and also, by extracting images from the second synthetic image set, the third synthetic image set can include both images without defects and defective images with various defective features, so that the features of the synthetic images can be comprehensively summarized through the third synthetic image set; finally, constructing a target dataset based on the real images in the second real image set and the synthetic images in the third synthetic image set, and training a detection model based on the target dataset, a high-quality target dataset can be constructed through the third synthetic image set including all features of the synthetic images and the second real image set images including real image features, and then a model with high detection accuracy for synthetic images can be trained based on the high-quality target dataset. It can be seen that the present application achieves the object of improving the detection accuracy of the detection model for synthetic images. Description of the Drawings

[0017] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 The flowchart provided for the first embodiment of the method for constructing a data set of this application; Figure 2 The flowchart provided for another embodiment of the method for constructing a data set of this application; Figure 3 The flowchart processing diagram for a specific embodiment of the method for constructing a data set of this application; Figure 4 The module structure diagram of the data set construction system for the embodiments of this application; Figure 5 The structural diagram of the electronic device involved in the method for constructing a data set in the embodiments of this application.

[0020] The realization of the purpose, functional features, and advantages of this application will be further described in combination with the embodiments with reference to the accompanying drawings. Specific Embodiments

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0022] To better understand the technical solutions of this application, the following will be described in detail in combination with the specification drawings and specific embodiments.

[0023] As a key carrier form of information, images play an important role in the information transmission process. Currently, there are various ways to obtain images. According to the generation characteristics of images, they can generally be classified into two categories: real images and synthetic images. Among them, real images are images that have not been tampered with by anyone; while synthetic images are images constructed and generated by means of specific models and algorithms.

[0024] Obviously, since synthetic images are artificially intervened and model - constructed, compared with real images, there are risks and potential hazards of being misused or even maliciously abused. In view of this, it is crucial to determine synthetic images from real images. At present, detection models are widely used in the field of synthetic image detection and are one of the main detection tools. However, with the technological evolution, synthetic images generated by diffusion models exhibit new characteristics. Synthetic images generated by diffusion models often no longer carry traditional recognition features such as artifacts, resulting in a low detection accuracy of detection models for synthetic images.

[0025] Therefore, how to improve the detection accuracy of detection models for detecting and recognizing synthetic images is a technical problem that those skilled in the art still need to solve.

[0026] To solve the above - mentioned problems, in the embodiments of the present application, each real image is divided into a first real - image set and a second real - image set, and synthetic images are obtained based on the real images in the first real - image set, so that real images and synthetic images for constructing a data set can be obtained. Then, the synthetic images are divided into a first synthetic - image set and a second synthetic - image set, and the synthetic images in the first synthetic - image set are classified. By classifying defective images, it is convenient to statistically analyze the characteristics of defective images of each category. Then, images of each category are equally sampled from the first synthetic - image set, and images are sampled from the second synthetic - image set to obtain a third synthetic - image set, which can further subdivide defective images, so as to facilitate the extraction of synthetic images with different defect characteristics. Moreover, by sampling images from the second synthetic - image set, the third synthetic - image set can include both images without defects and defective images with various defect characteristics, so that the characteristics of synthetic images can be comprehensively summarized through the third synthetic - image set. Finally, based on the real images in the second real - image set and the synthetic images in the third synthetic - image set, a target data set is constructed to train a detection model based on the target data set. A high - quality target data set can be constructed through the third synthetic - image set that contains all the characteristics of synthetic images and the real - image set images in the second real - image set that contain real - image characteristics. Furthermore, a model with high detection accuracy for synthetic images can be trained based on the high - quality target data set. It can be seen that the present application achieves the purpose of improving the detection accuracy of the detection model for synthetic images.

[0027] It should be noted that the execution subject of the data - set construction method of the present application can be an electronic device with data - processing, network - communication, and program - running functions, such as a tablet computer, a personal computer, a server, etc.

[0028] Based on this, the embodiments of the present application provide a data - set construction method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the data - set construction method of the present application.

[0029] In this embodiment, the dataset construction method includes steps S10 to S40: Step S10: Divide each real image into a first real image set and a second real image set, and obtain each synthetic image based on each real image in the first real image set; It should be noted that a synthetic image refers to an image created through computer generation algorithms (such as generative adversarial networks, diffusion models, etc.). A synthetic image is not directly captured from the real world, but is generated by an AI (Artificial Intelligence) model according to input data or parameters. A real image refers to an image directly captured from the real world through a camera or other imaging device. A real image represents an actual scene, object, or event in nature. Based on this, in this embodiment, the first real image set includes real images used to guide the generation of synthetic images, and the second real image set refers to a part of the real images directly used for training the detection model. The detection model is a model used to identify synthetic images, and the type of the model is not limited in this application.

[0030] Exemplarily, each real image can be randomly split into a first real image set and a second real image set, and then models such as DALL-E (an artificial intelligence image generation technology, a deep learning-based image generation model) can be used to perform operations such as image-to-image, image-to-text, and text-to-image processing on the real images in the first real image set, so as to obtain synthetic images.

[0031] It should be noted that the above-mentioned random allocation method can ensure that there is no systematic correlation between the first real image set and the second real image set, and further ensure that there is also no systematic correlation between the synthetic images generated based on the first real image set and the second real image set. Thus, the dataset for model training can be enriched based on the synthetic images and real images without systematic correlation.

[0032] Step S20: Divide each synthetic image into a first synthetic image set and a second synthetic image set, and classify the synthetic images in the first synthetic image set, where the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; Exemplarily, for the defects of synthetic images, the synthetic images can be divided into six different defect categories: physical defects, geometric defects, human body part defects, image distortion, text representation problems, and semantic scene problems. Specifically, the definitions of each category are as follows: Physical defects: Include all elements in the image that are contradictory to or violate physical laws. Specifically, physical defects mainly include optical anomalies and gravitational anomalies. Among them, optical anomalies include specular reflection paradox (the reflection angle on the metal surface or water surface deviates from the Fresnel's law. For example, the deviation of the water surface reflection from the actual position of the object exceeds 15°), shadow contradiction (the direction of the object's shadow is inconsistent with the position of the light source. For example, in a multi-light source scenario, the shadow directions conflict with each other), and light transmission anomaly (the light refraction path of the transparent material is not correctly represented. For example, the shape and position of the object seen through the glass are significantly distorted). Gravitational anomalies include anti-gravity suspension (the object lacks physical support. For example, a teacup floats 10 cm above the table), hydrodynamic error (the water flow exhibits non-Newtonian fluid characteristics. For example, a stationary water curtain on a vertical wall), and cloth movement paradox (the fluttering direction of the clothing is opposite to the direction of the wind force. For example, the wind blows to the left but the clothing flutters to the right).

[0033] Geometric defects: There are morphological errors in the synthetic image that violate the principles of projective geometry, manifested as non-uniformity in the shape, perspective relationship, or spatial structure of the object. Specifically, geometric defects mainly include morphological distortion, abnormal spatial relationship, and symmetry break. Among them, morphological distortion includes topological error (the abnormal number of furniture legs. For example, a chair with five legs) and surface curvature break (the sudden change of the surface curvature of a cylinder. For example, the curvature change rate exceeds 0.25 per square millimeter). Abnormal spatial relationship includes perspective contradiction (violation of the law of near-big-and-far-small. For example, a pedestrian in the distance is larger than a car nearby) and depth layering error (the foreground object is unreasonably occluded by the background object. For example, a person stands in front, but his shadow is projected on the person behind). Symmetry break includes mirror asymmetry (for example, the difference in pupil diameter between the left and right sides of the face exceeds 15%) and repeated pattern break (for example, the tile texture is misaligned by more than 5 pixels at the seam).

[0034] Human body part defects: Human body part defects refer to the unnatural manifestations in the anatomical structure or biomechanical characteristics of the human body that are easily detectable by the human visual system. Specifically, human body part defects mainly include local anomalies, global anomalies, and material anomalies. Among them, local anomalies include hand distortion (abnormal number of fingers. For example, a palm with six fingers, and the joint rotation angle exceeds the limit. For example, the wrist joint bends more than 120°), facial contradiction (the color difference of bilateral pupils. For example, the LAB color difference ΔE > 8, and the tooth arrangement violates the dental arch curve), global anomaly (disproportionate limb ratio. For example, the ratio of the forearm length to the upper arm exceeds the normal range of 1:1.2 - 1.5), and kinematic paradox (the bending direction of the knee joint is contradictory to the force direction. For example, the calf swings forward during running). Material anomalies include skin color mutation (for example, the LAB color difference ΔE > 5 between adjacent skin areas) and material penetration (for example, the non-physical fusion of the earring and the earlobe tissue).

[0035] Image distortion: Image distortion refers to signal processing defects introduced during the generation process in an image, resulting in impaired visual information integrity. Specifically, image distortion mainly includes frequency-domain anomalies, spatial-domain anomalies, and style contradictions. Among them, frequency-domain anomalies include checkerboard artifacts (e.g., periodic square noise appears in the high-frequency region, with an amplitude exceeding 3 dB) and color banding effects (e.g., discrete color levels appear in a gradually changing sky, with a hue mutation exceeding 5°). Spatial-domain anomalies include local blurring (e.g., the 50% value of the modulation transfer function at the edge of a key object is lower than 0.3 cycles per pixel) and detail annihilation (e.g., the structural similarity index in the texture area is lower than 0.65 compared to the original image). Style contradictions include brushstroke mutations (e.g., the edge of a vector graphic feature appears in an oil painting-style image) and lighting style conflicts (e.g., a realistic scene contains cartoon-rendered shadows).

[0036] Text representation problems: Text representation problems refer to morphological or semantic errors in the text elements in an image. Specifically, text representation problems mainly include morphological anomalies, semantic anomalies, and spatial anomalies. Among them, morphological anomalies include character distortion (e.g., the right arc of the letter "B" is missing, forming a combination of "13") and font mutation (e.g., a mixture of Song typeface and boldface appears within the same word). Semantic anomalies include meaningless combinations (e.g., the sign shows garbled characters such as "@#GmbH_2023") and context contradictions (e.g., a Chinese-style building sign shows the English word "PIZZERIA"). Spatial anomalies include perspective distortion (e.g., the text on the wall does not follow the planar projection transformation) and curved surface adaptation failure (e.g., the text on the surface of a cylinder is not deformed according to the curvature).

[0037] Semantic scene problems: Semantic scene problems refer to the situation where the combination of elements in an image conforms to physical / geometric laws, but the overall semantic logic violates common sense. Specifically, semantic scene problems mainly include spatio-temporal contradictions, ecological contradictions, and functional contradictions. Among them, spatio-temporal contradictions include season conflicts (e.g., a person in a snow scene is wearing a short-sleeved T-shirt) and day-night dislocation (e.g., the midday sun appears against a starry sky background). Ecological contradictions include abnormal species distribution (e.g., a group of penguins appears in a desert scene) and cultural mismatch (e.g., a person wearing Tang Dynasty clothing uses a smartphone). Functional contradictions include misuse of equipment (e.g., a microscope is used as a drinking cup) and behavioral paradoxes (e.g., a diver lights a lighter at the bottom of the sea).

[0038] In a feasible implementation manner, the defective images can be classified by means of a model or manual annotation, and the present application does not limit this.

[0039] Step S30, equally extract images of each category from the first synthetic image set, and extract images from the second synthetic image set to obtain a third synthetic image set; It should be noted that the third synthetic image set is a set of images extracted from the first synthetic image set and images extracted from the second synthetic image set. In addition, the above equal extraction means that for each category, the same number of images are extracted. For example, when there are 10 defect categories in the first synthetic image set and the preset number of extracted images is 100, the equal extraction is to extract 100 images from each of the 10 defect categories, so that 10 * 100 = 1000 images can be extracted from the first synthetic image set. In addition, to facilitate the detection model to learn the features of synthetic images without defects, images can also be extracted from the second synthetic image set, where the number of images extracted from the second synthetic image set is not limited in this embodiment.

[0040] It can be understood that different categories in the first synthetic image set have different defect features, while the second synthetic image set does not have the defects of synthetic images but has the general features of synthetic images. Therefore, in order to enable the detection model to comprehensively learn the features corresponding to different defect types and the general features of synthetic images, the third synthetic image set for training the detection model can be obtained by extracting images from each defect category and extracting images from the second synthetic image set.

[0041] In addition, in a feasible implementation manner, for the images extracted from the first synthetic image set and the images extracted from the second synthetic image set, data augmentation processing can also be performed, and then combined to obtain the third synthetic image set. Among them, data augmentation can be performed by rotation and flipping, so that the third synthetic image set can be expanded on the basis of retaining the image features.

[0042] In addition, in a feasible implementation manner, when there is a need for an application scenario, the extraction ratio of synthetic images of some defect categories in the first synthetic data set can also be increased specifically, so as to specifically enhance the recognition ability of the detection model for certain synthetic defect situations. And by reasonably controlling the extraction ratio of images of various categories in the first synthetic data set, it is possible to avoid the detection model from overfitting to a certain frequently occurring synthetic defect situation.

[0043] Step S40, based on the real images in the second real image set and the synthetic images in the third synthetic image set, construct a target data set to train a detection model based on the target data set.

[0044] It can be understood that after obtaining the third set of synthetic images containing the features of the synthetic images and the real images containing the features of the real images, the real images in the second set of real images and the third set of synthetic images can be combined to obtain the target data set. Furthermore, the detection model can be trained based on the target data set, so that the detection model can fully learn the difference in features between the synthetic images and the real images based on the target data set, thereby improving the accuracy of synthetic image detection.

[0045] In this embodiment, by dividing each real image into a first set of real images and a second set of real images, and obtaining synthetic images based on the real images in the first set of real images, real images and synthetic images for constructing a data set can be obtained; then the synthetic images are divided into a first set of synthetic images and a second set of synthetic images, and the synthetic images in the first set of synthetic images are classified. By classifying the defective images, it is convenient to count the features of the defective images of each category; then an equal number of images of each category are randomly selected from the first set of synthetic images, and images are randomly selected from the second set of synthetic images to obtain the third set of synthetic images, which can further subdivide the defective images, so as to facilitate the extraction of synthetic images with different defective features. Moreover, by randomly selecting images from the second set of synthetic images, the third set of synthetic images can simultaneously include images without defects and defective images with various defective features, so that the features of the synthetic images can be comprehensively summarized through the third set of synthetic images; finally, based on the real images in the second set of real images and the synthetic images in the third set of synthetic images, a target data set is constructed to train the detection model based on the target data set. A high-quality target data set can be constructed through the third set of synthetic images containing all the features of the synthetic images and the images in the second set of real images containing the features of the real images. Furthermore, a model with high accuracy in detecting synthetic images can be trained based on the high-quality target data set. It can be seen that the present application achieves the purpose of improving the accuracy of the detection model in detecting synthetic images.

[0046] Furthermore, based on the first embodiment of the data set construction method of the present application, a second embodiment of the data set construction method of the present application is proposed.

[0047] In this embodiment, before the above step S10, the method further includes: Step S100, obtaining the original real images in each data source, and filtering the original real images through a preset image size, image clarity range, aspect ratio range, and redundant information to obtain each intermediate real image; It can be understood that the features of the original real images from a single data source often have limitations. Therefore, by obtaining the original real images from multiple data sources, it is convenient for the detection model to learn the features of different original real images. Among them, the original real image refers to the real image obtained from the data source without any processing.

[0048] It can be understood that low-quality images and images with an unbalanced aspect ratio will affect the training effect of the detection model. Therefore, the original real images can be filtered by the image size, the range of image clarity, the range of aspect ratio, and the redundant information of the images, and the remaining original real images after filtering are used as intermediate real images.

[0049] In addition, the QR codes, emojis, and region segmentation symbols in the original real images are invalid data in the process of synthetic image detection. Therefore, in a feasible implementation manner, the invalid data can also be filtered based on the data characteristics of each data source.

[0050] Step S200, deduplicate each of the intermediate real images, and use each of the intermediate real images obtained after deduplication as each of the real images.

[0051] It can be understood that there may be duplicate images in different data sources. Therefore, the convolutional neural network can be trained through contrast learning technology to detect duplicate images, and then the duplicate data can be removed from the intermediate real images.

[0052] In this embodiment, the present application filters the images by setting various filtering conditions and removes duplicate images through data deduplication, thereby improving the quality of the images and laying a foundation for improving the accuracy of the detection model.

[0053] In a feasible implementation manner, before the above step S200, the method further includes: Step S300, calculate the similarity of the text-image pairs formed by each of the intermediate real images and the text corresponding to each of the intermediate real images, obtain the similarity corresponding to each of the text-image pairs, and screen out each target real image from each of the intermediate real images based on each of the similarities; It can be understood that in addition to images, the data source may also include preset text corresponding to the images. For example, use a contrastive language-image pre-training model to calculate the similarity of the text-image pairs formed by the intermediate real images and the corresponding text. Thus, the similarity between the intermediate real image and the text corresponding to it can be obtained based on the pre-trained contrastive language-image pre-training model. Then, the obtained similarities can be compared with a preset similarity threshold respectively, and the intermediate real images with similarities greater than or equal to the similarity threshold are used as target real images.

[0054] Exemplarily, if the similarity between the intermediate real image A and the text corresponding to the intermediate real image A is greater than the similarity threshold, the intermediate real image A can be used as the target real image.

[0055] Based on this, the above step S200 includes: Step S2001, removing duplicates from each of the target real images, and using each of the target real images obtained after duplicate removal as each of the real images.

[0056] In this embodiment, the present application can retain images with higher quality by comparing the language-image pre-training model to obtain the similarity and obtaining real images based on the similarity, thereby improving the quality of the data set and further increasing the probability that the detection model learns correct features.

[0057] In a feasible implementation manner, before the above step S10, the method further includes: Step S400, obtaining the text corresponding to the original real images in each data source, and filtering the text corresponding to the original real images through a preset text size threshold to obtain each intermediate text; Step S500, removing duplicates from each of the intermediate texts, and using the original real images corresponding to each of the intermediate texts obtained after duplicate removal as each of the real images.

[0058] It can be understood that in addition to factors such as image size and image clarity affecting the quality of real images and thus the accuracy of the detection model for detecting synthetic images, the text corresponding to the real images also affects the accuracy of the detection model for detecting synthetic images. Therefore, the text corresponding to the original real images can also be filtered through a preset text size threshold, where the text size refers to the storage space size occupied by the text.

[0059] Specifically, the text with a text size lower than the preset value can be used as the intermediate text, and after removing duplicates from the intermediate text, the original real images corresponding to the intermediate text after duplicate removal can be used as real images.

[0060] Thus, the present application can further improve the efficiency of the detection model for detecting synthetic images by processing the text corresponding to the images.

[0061] In a feasible implementation manner, before the above step S10, the method further includes: Step S600, filtering the text corresponding to each real image through a preset sensitive word detector, and filtering each of the real images through a preset NSFW image detector to obtain a preprocessed set of real images; It is understandable that there may also be NSFW images (Not Safe For Work, images containing pornographic, violent or other inappropriate content) in the real images, and the text corresponding to the real images may include personal data. Therefore, sensitive data such as ID numbers, mobile phone numbers, email addresses, IP addresses, etc. in the text can be identified and removed through a preset sensitive word detector, and NSFW images can be filtered out through an NSFW image detector.

[0062] Based on this, step S10 above further includes: Step S101, dividing the preprocessed real image set into a first real image set and a second real image set.

[0063] In this embodiment, the technical problem of privacy leakage can be avoided through a sensitive word detector and an NSFW image detector.

[0064] Furthermore, based on the first embodiment and / or the second embodiment of the dataset construction method of the present application, a third embodiment of the dataset construction method of the present application is proposed.

[0065] Please refer to Figure 2 , in this embodiment, step S30 above includes: Step S301, in each category of the first synthetic image set, respectively extract a first preset number of images to obtain a fourth synthetic image set; Step S302, determine the total number of defect-free images according to the number of images in the fourth synthetic image set and a preset coefficient value, and extract images from the second synthetic image set according to the total number of defect-free images to obtain a fifth synthetic image set; It should be noted that the coefficient value is used to adjust the ratio of the number of defective images to the number of defect-free images. The total number of defect-free images can be the product of the number of images in the fourth synthetic image set and the coefficient value.

[0066] Step S303, merge the fourth synthetic image set and the fifth synthetic image set to obtain a third synthetic image set.

[0067] In this embodiment, by equally extracting the synthetic images corresponding to different defect types and extracting defect-free synthetic images according to a preset ratio, the present application can ensure that the detection model can learn the features of different defects and the features of defect-free synthetic images, thereby improving the accuracy of the detection model in identifying synthetic images.

[0068] Furthermore, based on the first embodiment and / or the second embodiment of the dataset construction method of the present application, a fourth embodiment of the dataset construction method of the present application is proposed.

[0069] In this embodiment, step S10 includes: Step S102: Classify each of the real images, and respectively extract a second preset number of images from each category to obtain a second set of real images; It should be noted that the basis for classifying real images can be the image scene or the type of image object, and this application does not limit this.

[0070] In a feasible manner, the categories may include: people, animals, food, nature and environment, architecture and city landscapes, urban infrastructure, transportation vehicles, daily life items, health and medicine, art and design, social activities, equipment and tools.

[0071] To facilitate the detection model to learn the real image features of each category, a second preset number of images can be respectively extracted from each category, and the extracted images are used as the second set of real images.

[0072] Step S103: Combine the real images that do not belong to the second set of real images to obtain a first set of real images.

[0073] In this embodiment, by first classifying and then dividing the second set of real images and the first set of real images, this application can enrich the features in the dataset, and thus facilitate the detection model to learn more comprehensive real image features.

[0074] In a feasible implementation manner, step S10 further includes: Step S104: Input each of the real images in the first set of real images into an image-to-image model to obtain respective synthesized images; In this embodiment, synthesized images can be obtained through an image-to-image model. Among them, the image-to-image model can be an extended model or a generative adversarial network model, and this application does not limit this.

[0075] Step S105: Input the text corresponding to each of the real images in the first set of real images into a text-to-image model to obtain respective synthesized images; In this embodiment, in addition to obtaining synthesized images based on the image-to-image model, synthesized images can also be obtained based on the text corresponding to the real images through the text-to-image model. In this application, the specific type of the text-to-image model is not limited.

[0076] Step S106: Input each of the real images in the first set of real images into an image-to-text model to obtain the material text corresponding to each of the real images in the first set of real images; Step S107: Perform text enhancement on each of the material texts, and input each of the material texts after text enhancement into the text-to-image model respectively to obtain the synthetic images corresponding to each of the real images in the first real image set.

[0077] It should be noted that when directly obtaining synthetic images based on real images through the image-to-image model, due to the randomness inherent in the image-to-image model, it will be impossible to control the generation details of the synthetic images. Therefore, this application proposes a solution that combines the image-to-text model, text enhancement, and the text-to-image model.

[0078] Specifically, after obtaining the material texts through the image-to-text model, the industry knowledge can be incorporated into the material texts through text enhancement technology, so that the process of generating synthetic images based on the material texts is controllable, thereby improving the quality of the synthetic images, and further enabling the detection model to detect synthetic images in a specific industry.

[0079] Furthermore, based on each of the above embodiments of the data set construction method of this application, a specific embodiment of the data set construction method of this application is proposed.

[0080] Please refer to Figure 3 , Figure 3 which is the flow chart of a specific embodiment of the data set construction method of this application. In Figure 3 it, the steps of the data set construction method in this application are as follows: 1. Data processing: Obtain image data from SA1B (a large-scale image segmentation data set containing 11 million diverse, high-resolution, privacy-protected images, and more than 1.1 billion high-quality segmentation masks), Wanjuan1.0 (the first open-source version of the anjuan multimodal corpus, including three parts: text set, image-text data set, and video data set), COCO (a large-scale object detection, segmentation, and caption data set), Laion5B (a large-scale multimodal data set containing 5.85 billion high-quality image-text pairs), and DOCCI (containing approximately 15,000 high-resolution images with detailed human annotation descriptions). For example, 50,000 images can be extracted from SA1B, Wanjuan1.0, COCO, and Laion5B respectively, and 15,000 image-text data can be extracted from DOCCI to obtain 215,000 pieces of data. Then, perform data processing according to several steps including data filtering, data deduplication, security detection, and image-text pair processing.

[0081] (1) Data filtering: Filter out low-quality images that are too small or too large (such as images smaller than 150 pixels and images larger than 50,000 pixels), and images with extremely unbalanced aspect ratios (such as images with a ratio exceeding 4:1). In addition, explore the data characteristics of multiple types of data sources and remove invalid data such as QR codes, emojis, and region-segmented images.

[0082] (2) Data deduplication: Use contrastive learning to train a convolutional neural network to solve the problem of image duplicate detection and remove duplicate data from the dataset.

[0083] (3) Security detection: Apply the NSFW image detector and sensitive word detection to all data in the dataset. If an image is an NSFW image, delete the image. In addition, to reduce the risk of personal data leakage, sensitive data such as ID numbers, mobile phone numbers, email addresses, and IP addresses in the data can also be removed.

[0084] (4) Image-text pair data processing: To ensure the usability of image-text pair data, the content of the image-text pairs can be filtered. For example, the cosine similarity between the image and text encodings can be calculated using the ViT-B / 32 CLIP model (a variant of the CLIP model, where "B" represents the Base model and "32" refers to the patch size of 32 used by the model), and then the image-text pairs with low cosine similarity can be deleted (such as removing all image-text pairs with a similarity lower than 0.28).

[0085] After 215,000 pieces of data from SA1B, Wanjuan1.0, COCO, Laion5B, and DOCCI have gone through several steps including data filtering, deduplication, security detection, and image-text pair data processing, 128,500 pieces of compliant data are obtained.

[0086] 2. Dataset construction: (1) Based on the 128,500 pieces of data obtained after data processing, 40,000 pieces of data are extracted as the real data for training the detection model (i.e., obtaining the second real image set). The real data includes categories such as people, animals, food, nature and environment, architecture and urban landscapes, urban infrastructure, transportation vehicles, daily life items, health and medicine, art and design, social activities, equipment and tools, etc. The remaining 88,500 pieces of data are used as materials for generating synthetic images (i.e., obtaining the first real image set).

[0087] (2) Select four types of models, namely ProGAN (a variant of the generative adversarial network), SD3 (an open-source text-to-image model), Kolors (a large text-to-image model), and DALLE3 (a version of the DALL-E series of models). Randomly and evenly divide 88,500 pieces of data into four groups as the source data for the four types of models. Among them, the text required for the text-to-image models is generated by an image-to-text model, such as the CPM2.5 model (an autoregressive language model based on the Transformer architecture), or it can also be obtained by text augmentation using the text corresponding to real images. One image or one piece of text generates one synthetic image. In this way, 88,500 synthetic images are constructed (i.e., Figure 3 the synthetic data in

[0088] (3) Perform annotation processing on the 88,500 synthetic images according to the classification of defective images (i.e., Figure 3 the annotation of the synthetic data in

[0089] (4) Construct the target dataset from 40,000 real images and the 40,000 synthetic images extracted in a 1:1 ratio. Then divide the target dataset into a fine-tuning dataset and an evaluation dataset. The ratio of the fine-tuning dataset to the evaluation dataset is 8:2, that is, 64,000 images (32,000 real images and 32,000 synthetic images) are used to construct the fine-tuning dataset, and 16,000 images (8,000 real images and 8,000 synthetic images) are used to construct the evaluation dataset.

[0090] 3. Model fine-tuning and evaluation: Based on the constructed dataset, perform binary classification fine-tuning on the CLIP ViT-L / 14 model to determine whether it is a synthetic image, and adapt the fine-tuned model to the classification task of distinguishing synthetic images from real images.

[0091] (1) Model fine-tuning: For the binary classification task of distinguishing synthetic images from real images, the output of the penultimate layer of the model is fine-tuned to identify synthetic images. 32,000 real images and 32,000 synthetic images are used for fine-tuning. The corresponding feature vectors {𝐫1,…,𝐫N} and {𝐟1,…,𝐟N} extracted at the output of the penultimate layer are collected, where 𝐫i = CLIP*(Ri) and 𝐟i = CLIP*(Fi); The Stochastic Gradient Descent (SGD) optimizer is used to update the parameters. The cross-entropy loss function is used to calculate the difference between the model prediction and the actual data by minimizing the distribution difference of the feature vectors in the high-dimensional space. The learning rate is set to 4e-6 and the weight decay is set to 1e-3, and the model is trained for 3 - 8 epochs (training cycles).

[0092] (2) Model evaluation: The performance of the model is evaluated on an independent evaluation dataset. The fine-tuned CLIP ViT-L / 14 model is applied to extract the feature vectors of the images to be evaluated. By calculating the distance from the image features to the known real image features in the feature space, the nearest neighbor method is used for scoring, and a confidence estimate (in the range of 0 - 1) of whether the picture is a synthetic image is given. Scoring criteria: Output range: 0 - 1; The higher the score (1), the more likely it is a synthetic image, and the lower the score (0), the more likely it is a real image; Given a threshold of 0.5, if the score is greater than 0.5, it is considered a synthetic image.

[0093] The evaluation results of each synthetic image detection method on the evaluation dataset are compared as shown in Table 1 below: Table 1

[0094] The specific meanings of each index in Table 1 are as follows: Accuracy refers to the proportion of the number of samples correctly predicted by the model to the total number of samples. The calculation formula for accuracy Accuracy is: Accuracy = (TP + TN) / (TP + TN + FP + FN); Among them, TP (True Positive) is the number of synthetic images correctly identified, TN (True Negative) is the number of real images correctly identified, FP (False Positive) is the number of real images misidentified as synthetic images, and FN (False Negative) is the number of synthetic images misidentified as real images.

[0095] The Average Precision (AP) is obtained by calculating the area under the Precision-Recall (P-R) curve. It combines Precision and Recall, providing a comprehensive evaluation metric. The calculation steps are as follows: Calculate the prediction score for each sample; Sort all samples in descending order according to the prediction scores; Gradually increase the threshold and calculate the Precision and Recall at different thresholds; Plot the Precision-Recall curve; Calculate the area under the curve: AP = Σ(R_n - R_(n-1)) * P_n; Where: R_n is the Recall at the nth threshold, R_(n-1) is the Recall at the (n - 1)th threshold, and P_n is the Precision at the nth threshold.

[0096] The true accuracy refers to the accuracy of real images, that is, the prediction accuracy of the detection model in identifying real images. This metric focuses on the ability of the detection model to distinguish real images in practical applications, and its calculation formula is the same as that of the detection accuracy.

[0097] The forged accuracy refers to the accuracy of synthetic images, that is, the prediction accuracy of the detection model in identifying synthetic images. This metric measures the ability of the detection model to distinguish synthetic images, and its calculation formula is also the same as that of the detection accuracy.

[0098] As shown in the above table, the fine-tuned detection model of this application has an accuracy of 72.45%, an average precision of 69.45%, a true accuracy of 99.58%, and a forged accuracy of 70.54% on the evaluation dataset, showing a significant improvement compared to the previous traditional methods. The methods of CNNSpot, FreDect, Fusing, GramNet, LGrad, and UnivFD perform poorly on the evaluation dataset, with an overall accuracy less than 50%, lower than random; the average precision is between 32% and 45%, indicating that the overall discriminative ability of traditional methods is very weak; the true accuracy is generally very high, with most greater than 90%, indicating that the detection model is very good at identifying real images and severely biased towards judging images as real images; while the forged accuracy is extremely low, between 0 and 0.15, indicating that traditional methods can hardly identify the synthetic images in the evaluation test set. In addition, the fine-tuned CLIP ViT-L / 14 model can better distinguish between synthetic images and real images, can adapt to images generated by different generative models, and since the detection model proposed in this application uses a fixed feature extractor, it does not require targeted training, and the evaluation of this application is stable and not affected by the features of a single generative model, showing excellent performance in cross-model tests. It can be seen that this application improves the accuracy and generalization ability of synthetic image detection and expands the application scenarios of synthetic image detection models.

[0099] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the dataset construction method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0100] This application also provides a dataset construction system. Please refer to Figure 4 , and the dataset construction system includes: The first image processing module 10 is used to divide each real image into a first real image set and a second real image set, and obtain each synthetic image based on each real image in the first real image set; The second image processing module 20 is used to divide each synthetic image into a first synthetic image set and a second synthetic image set, and classify the synthetic images in the first synthetic image set, where the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; The third image processing module 30 is used to equally extract images of each category from the first synthetic image set and extract images from the second synthetic image set to obtain a third synthetic image set; The dataset integration module 40 is used to construct a target dataset based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target dataset.

[0101] In one embodiment, the dataset construction system further includes: The first filtering module is used to obtain the original real images in each data source, and filter the original real images through a preset image size, image clarity range, aspect ratio range, and redundant information to obtain each intermediate real image; The first deduplication module is used to deduplicate each of the intermediate real images, and use each of the intermediate real images obtained after deduplication as each of the real images.

[0102] In one embodiment, the dataset construction system further includes: The similarity screening module is used to calculate the similarity of each text-image pair composed of each of the intermediate real images and the text corresponding to each of the intermediate real images, obtain the similarity corresponding to each text-image pair, and screen out each target real image from each of the intermediate real images based on each similarity; Based on this, the above-mentioned first deduplication module is further used to: Deduplicate each of the target real images, and use each of the target real images obtained after deduplication as each of the real images.

[0103] In one embodiment, the dataset construction system further includes: The second filtering module is used to obtain the text corresponding to the original real images in each data source, and filter the text corresponding to the original real images through a preset text size threshold to obtain each intermediate text; The second deduplication module is used to deduplicate each of the intermediate texts, and use the original real images corresponding to each of the intermediate texts obtained after deduplication as each of the real images.

[0104] In one embodiment, the dataset construction system further includes: The third filtering module is used to filter the text corresponding to each real image through a preset sensitive word detector, and filter each of the real images through a preset NSFW image detector to obtain a preprocessed real image set; Based on this, the above-mentioned first image processing module 10 is further used to: Divide the preprocessed real image set into a first real image set and a second real image set.

[0105] In one embodiment, the third image processing module 30 is further used to: In each category of the first synthetic image set, extract a first preset number of images to obtain a fourth synthetic image set; Determine the total number of defect-free images according to the number of images in the fourth synthetic image set and a preset coefficient value, and extract images from the second synthetic image set according to the total number of defect-free images to obtain a fifth synthetic image set; Combine the fourth synthetic image set and the fifth synthetic image set to obtain a third synthetic image set.

[0106] In one embodiment, the first image processing module 10 is further configured to: Classify each of the real images, and extract a second preset number of images from each category to obtain a second real image set; Combine the real images that do not belong to the second real image set to obtain a first real image set.

[0107] In one embodiment, the first image processing module 10 is further configured to: Input each of the real images in the first real image set into an image-to-image model to obtain respective synthetic images; Or, Input the text corresponding to each of the real images in the first real image set into a text-to-image model to obtain respective synthetic images; Or, Input each of the real images in the first real image set into an image-to-text model to obtain the material text corresponding to each of the real images in the first real image set; Perform text enhancement on each of the material texts, and input each of the text-enhanced material texts into a text-to-image model to obtain the synthetic image corresponding to each of the real images in the first real image set.

[0108] The dataset construction system provided by the present application adopts the dataset construction method in the above embodiment, and can solve the technical problem of how to improve the accuracy of the detection model in detecting and recognizing synthetic images. Compared with the prior art, the beneficial effects of the dataset construction system provided by the present application are the same as those of the dataset construction method provided by the above embodiment, and other technical features in the dataset construction system are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0109] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the dataset construction method in the first embodiment above.

[0110] Refer to the following Figure 5 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0111] As Figure 5 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. Instead, more or fewer systems may be implemented or had.

[0112] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0113] The electronic device provided by the present application, adopting the dataset construction method in the above-mentioned embodiment, can solve the technical problem of how to improve the accuracy of the detection model in detecting and recognizing synthetic images. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the dataset construction method provided by the above-mentioned embodiment, and other technical features in this electronic device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0114] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0115] As mentioned above, only the specific implementation manners of the present application are described, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0116] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the dataset construction method in the above-mentioned embodiment.

[0117] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0118] The above computer-readable storage medium may be included in an electronic device; or may exist separately without being assembled into the electronic device.

[0119] The above computer-readable storage medium carries one or more programs, which when executed by the electronic device, cause the electronic device to: divide each real image into a first real image set and a second real image set, and obtain a synthetic image based on the real images in the first real image set; divide the synthetic image into a first synthetic image set and a second synthetic image set, and classify the synthetic images in the first synthetic image set; equally extract images of each category from the first synthetic image set, and extract images from the second synthetic image set to obtain a third synthetic image set; construct a target data set based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target data set.

[0120] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider using the Internet).

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0122] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0123] The readable storage medium provided in the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned dataset construction method, and can solve the technical problem of how to improve the accuracy of the detection model in detecting and recognizing synthetic images. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as those of the dataset construction method provided in the above embodiments, and will not be elaborated here.

[0124] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made using the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for constructing a data set, characterized in that: The data set construction method comprises: Dividing each real image into a first real image set and a second real image set, and obtaining each synthetic image based on each real image in the first real image set; Dividing each of the synthetic images into a first synthetic image set and a second synthetic image set, and classifying the synthetic images in the first synthetic image set, wherein the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; Extracting an equal number of images of each category from the first synthetic image set, and extracting images from the second synthetic image set, to obtain a third synthetic image set; Based on the real images in the second real image set and the synthetic images in the third synthetic image set, a target dataset is constructed to train a detection model based on the target dataset.

2. The method for constructing a data set according to claim 1, wherein: Before the step of dividing each real image into a first real image set and a second real image set, the method further includes: Acquire original real images from various data sources, and filter the original real images by using preset image size, image definition range, aspect ratio range, and redundant information to obtain various intermediate real images; Deduplication is performed on each of the intermediate real images, and each of the intermediate real images obtained after deduplication is used as each of the real images.

3. The method for constructing a data set according to claim 2, wherein: Before the step of removing duplicates from each of the intermediate real images and using each of the intermediate real images obtained after the removal of duplicates as each of the real images, the method further includes: Calculating the similarity of the image-text pairs consisting of the intermediate real images and the texts corresponding to the intermediate real images to obtain the similarities corresponding to the image-text pairs, and selecting the target real images from the intermediate real images based on the similarities; The step of removing duplicates from each of the intermediate real images and using each of the intermediate real images obtained after the removal of duplicates as each of the real images comprises: Deduplication is performed on each of the target real images, and each of the target real images obtained after deduplication is used as each of the real images.

4. The method for constructing a data set according to claim 1, wherein: Before the step of dividing each real image into a first real image set and a second real image set, the method further includes: Acquire the text corresponding to the original real image in each data source, and filter the text corresponding to the original real image by a preset text size threshold to obtain each intermediate text; Deduplication is performed on each of the intermediate texts, and the original real images corresponding to each of the intermediate texts obtained after deduplication are used as each real image.

5. The method for constructing a data set according to claim 1, wherein: Before the step of dividing each real image into a first real image set and a second real image set, the method further includes: The text corresponding to each real image is filtered by a preset sensitive word detector, and each real image is filtered by a preset NSFW image detector to obtain a preprocessed real image set; The step of dividing each real image into a first real image set and a second real image set comprises: The preprocessed real image set is divided into a first real image set and a second real image set.

6. The method for constructing a data set according to claim 1, wherein: The step of extracting an equal amount of images of each category from the first synthetic image set and extracting images from the second synthetic image set to obtain a third synthetic image set comprises: Extracting a first preset number of images from each category in the first synthetic image set to obtain a fourth synthetic image set; Determining the total number of defect-free images according to the number of images in the fourth synthetic image set and a preset coefficient value, and extracting images from the second synthetic image set according to the total number of defect-free images to obtain a fifth synthetic image set; The fourth composite image set and the fifth composite image set are combined to obtain a third composite image set.

7. The method for constructing a data set according to claim 1, wherein: The step of obtaining each synthetic image based on each of the real images in the first real image set comprises: Inputting each of the real images in the first real image set into the graph-generated graph model to obtain each synthetic image; or, Inputting the text corresponding to each of the real images in the first real image set into the text-generated graph model to obtain each synthetic image; or, Inputting each of the real images in the first real image set into the image-to-text model to obtain the material text corresponding to each of the real images in the first real image set; Text enhancement is performed on each of the material texts, and each of the material texts after text enhancement is input into the text-generated graph model to obtain a synthetic image corresponding to each of the real images in the first real image set.

8. A data set construction system, characterized in that: The data set construction system comprises: A first image processing module, used for dividing each real image into a first real image set and a second real image set, and obtaining each synthetic image based on each real image in the first real image set; a second image processing module, configured to divide each of the synthetic images into a first synthetic image set and a second synthetic image set, and classify the synthetic images in the first synthetic image set, wherein the first synthetic image set includes synthetic images with defects, and the second synthetic image set includes synthetic images without defects; A third image processing module, used to extract images of each category in equal quantities from the first synthetic image set, and to extract images from the second synthetic image set, so as to obtain a third synthetic image set; A data set integration module is used to construct a target data set based on the real images in the second real image set and the synthetic images in the third synthetic image set, so as to train a detection model based on the target data set.

9. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data set construction method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data set construction method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Face forgery detection method based on multi-feature fusion network

    CN118015714A

  • Artificial intelligence generated image detection method and device, storage medium and electronic equipment

    CN119741396A

  • Method and apparatus for training text-graph model, device, and storage medium

    WO2025035924A1