Image detection and model training method and device, equipment, medium and product
By aligning and extracting high-frequency information of the detected images, combined with deep learning models, the problem of poor detection of fake images is solved, and higher detection accuracy and visualization effects are achieved.
Patent Information
- Application Number
- CN202510526643.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-29
AI Technical Summary
With the development of AI technology, the fidelity of forged images has increased, and traditional face image detection methods are not effective, making it difficult to effectively distinguish between real and forged images.
By aligning the image to be detected, high-frequency information is extracted and fused with the aligned image, image detection is performed using a deep learning model, and a heat map is generated to identify the forged part.
It improves the accuracy and visualization of image detection, can effectively identify and prevent the propagation of forged images, and improves the security and accuracy of image detection.
Smart Images

Figure CN120564012A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, deep learning and other technical fields, and in particular to an image detection and model training method, device, equipment, medium and product. Background Art
[0002] With the development of artificial intelligence (AI) technology, images forged using AI technology are becoming more and more realistic, so it is necessary to improve image detection effects. Summary of the Invention
[0003] The present disclosure provides an image detection and model training method, apparatus, device, medium and product.
[0004] According to one aspect of the present disclosure, an image detection method is provided, comprising: aligning an image to be detected to obtain an aligned image; extracting high-frequency information in the aligned image to obtain a high-frequency image; fusing the aligned image and the high-frequency image to obtain a fused image; and obtaining an image detection result based on the fused image.
[0005] According to another aspect of the present disclosure, a method for training an image detection model is provided, comprising: aligning sample images to obtain an aligned image; extracting high-frequency information from the aligned image to obtain a high-frequency image; fusing the aligned image with the high-frequency image to obtain a fused image; processing the fused image using an image detection model to obtain a prediction result; constructing a loss function based on the prediction result and a true value result corresponding to the sample image, and adjusting model parameters of the image detection model based on the loss function.
[0006] According to another aspect of the present disclosure, an image detection device is provided, including: an alignment module for aligning an image to be detected to obtain an aligned image; an extraction module for extracting high-frequency information in the aligned image to obtain a high-frequency image; a fusion module for fusing the aligned image and the high-frequency image to obtain a fused image; and an acquisition module for acquiring an image detection result based on the fused image.
[0007] According to another aspect of the present disclosure, an image detection model training device is provided, including: an alignment module for aligning sample images to obtain an aligned image; an extraction module for extracting high-frequency information in the aligned image to obtain a high-frequency image; a fusion module for fusing the aligned image and the high-frequency image to obtain a fused image; an acquisition module for processing the fused image using an image detection model to obtain a prediction result; and an adjustment module for constructing a loss function based on the prediction result and the true value result corresponding to the sample image, and adjusting the model parameters of the image detection model based on the loss function.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods described in any one of the above aspects.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods according to any one of the above aspects.
[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of the above aspects.
[0011] According to the technical solution disclosed in the present invention, the image detection effect can be improved.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0014] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0015] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram of an alignment process provided according to an embodiment of the present disclosure;
[0017] Figure 4is a schematic diagram according to a third embodiment of the present disclosure;
[0018] Figure 5 This is an overall architecture diagram of the model training phase provided according to an embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0020] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0021] Figure 8 Schematic diagram of an electronic device used to implement the image detection method or image detection model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] Taking face image forgery as an example, it is necessary to detect whether a face image is a forged image.
[0024] In related technologies, face authentication is usually performed based on traditional machine learning technology or deep learning technology. However, with the development of AI technology, AI-forged face images are becoming more and more realistic, and traditional face image detection methods have the problem of poor effectiveness.
[0025] In order to improve the image detection effect, the present disclosure provides the following embodiments.
[0026] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides an image detection method, the method comprising:
[0027] 101. Align the images to be detected to obtain aligned images.
[0028] 102. Extract high-frequency information from the aligned images to obtain a high-frequency image.
[0029] 103. Fuse the aligned image and the high-frequency image to obtain a fused image.
[0030] 104. Obtain an image detection result based on the fused image.
[0031] Among them, alignment refers to adjusting the key points in the image to be detected to the standard position and posture.
[0032] Specifically, taking the image to be detected as a face image as an example, the key points of the facial features in the face image (such as eyes, nose, mouth, etc.) can be detected, and the key points of the facial features can be adjusted to standard positions and postures through mathematical transformations (such as affine transformations). This can reduce the impact of facial expressions and postures on the detection results.
[0033] Aligned images refer to images after alignment processing.
[0034] Image information can be divided into high-frequency information and low-frequency information. Low-frequency information mainly corresponds to the parts of the image where the pixel values change slowly, such as the background and large objects. High-frequency information mainly corresponds to the parts of the image where the pixel values change dramatically, such as the edges, details, textures, and noise of the image.
[0035] A high-frequency image refers to an image composed of high-frequency information in the aligned image.
[0036] Specifically, a spatial rich model (SRM) may be used to process the aligned images to obtain high-frequency images.
[0037] The fused image refers to the image obtained by fusing the high-frequency image and the aligned image.
[0038] For example, the high-frequency image and the aligned image may be added to obtain a fused image.
[0039] After obtaining the fused image, the fused image can be input into a pre-trained image detection model, and the output is the image detection result, which is used to characterize whether the image to be detected is a real image or a forged image.
[0040] The image detection result may specifically include a real type score and a forged type score. If the real type score is higher than the forged type score, the image to be detected is determined to be a real image. If the real type score is lower than the forged type score, the image to be detected is determined to be a forged image.
[0041] In this embodiment, by aligning the image to be detected to obtain an aligned image, and performing subsequent processing based on the aligned image, the influence of the bending deformation of the image to be detected on the detection result can be avoided, thereby improving the accuracy of the image detection result; by extracting the high-frequency information of the aligned image to obtain a high-frequency image, and performing subsequent processing based on the high-frequency image, the depth features of the aligned image can be effectively captured, thereby further improving the accuracy of the image detection result; by fusing the aligned image and the high-frequency image to obtain a fused image, thereby further improving the accuracy and effect of image detection.
[0042] Figure 2is a schematic diagram according to the second embodiment of the present disclosure, which provides an image detection method, such as Figure 2 As shown, the method includes:
[0043] 201. Align the images to be detected to obtain aligned images.
[0044] 202. Input the aligned image into the SRM, and extract high-frequency information from the aligned image through the SRM to output a high-frequency image.
[0045] 203. Add the aligned image and the high-frequency image to obtain a fused image.
[0046] 204. Input the fused image into a pre-trained image detection model to output an image detection result.
[0047] In step 201, key points in the image to be detected may be extracted; and based on the key points, an affine transformation may be performed on the image to be detected to obtain the aligned image.
[0048] Figure 3 is a schematic diagram of alignment processing provided according to an embodiment of the present disclosure.
[0049] like Figure 3 As shown in FIG, taking the face image to be detected as an example, a face key point detection algorithm can be used to detect key points (represented by circles) in the face image, such as eyes, nose, mouth, etc., and then an aligned image can be obtained through affine transformation.
[0050] The transformation formula can be as follows:
[0051]
[0052] Among them, (x', y') is the pixel coordinate after transformation; (x, y) is the pixel coordinate before transformation; M is the affine transformation matrix, which can be obtained based on the key point information (such as position and posture) obtained by detection and its corresponding standard information.
[0053] In this embodiment, the aligned image can be obtained simply and efficiently through affine transformation. When the image to be detected is a face image, the influence of facial expression and posture can be avoided, thereby improving the accuracy of the image detection result.
[0054] In step 202, SRM is a steganalysis technique for spatially coded images. SRM uses multiple sub-models to extract more types of features. A sub-model refers to extracting corresponding features after the image is filtered specifically. Since neighborhood correlation can be represented by the prediction error between local pixels, filtering here generally refers to the operation of outputting this prediction error. This type of error is generally called residuals.
[0055] SRM mainly includes the following steps: calculating residuals, quantization and truncation, and statistical co-occurrence matrix.
[0056] Compute the residuals:
[0057]
[0058] Among them, X ij It is the pixel at row i and column j in the aligned image X;
[0059] R=(R ij ) is the residual image corresponding to X, R ij It's X ij The corresponding residual value;
[0060] It's X ij Neighborhood pixels of
[0061] Is the prediction function set to use To predict cX ij ;
[0062] c is the residual order, which is generally the number of neighborhood pixels.
[0063] Quantization and truncation:
[0064]
[0065] Among them, q is the set quantization step size;
[0066] trunc T () is the truncation function, defined as:
[0067]
[0068] T is the set cutoff threshold;
[0069] round() indicates rounding operation.
[0070] Through quantization and truncation, we can analyze areas with strong correlation and small residuals to improve targeting; we can also reduce the dimension of the final extracted features and reduce the amount of calculation.
[0071] Statistical co-occurrence matrix: The SRM feature is ultimately expressed as the fourth-order joint distribution probability of the above quantized and truncated residuals under each sub-model, that is, the four consecutive sample points d = (d1, d2, d3, d4) ∈ T4 = {-T, ..., T} in the horizontal or vertical direction of the residual 4 The joint distribution probability of .
[0072] Taking the horizontal direction as an example, the vertical direction is similar, and the joint distribution probability in the horizontal direction is:
[0073]
[0074] Where Z is the total number of all occurrences, which is a normalized function such that
[0075]
[0076] The co-occurrence matrix obtained above is used as the feature map of each channel, and the feature maps of each channel constitute a high-frequency image.
[0077] In this embodiment, high-frequency images are obtained through SRM, so that the generated multi-channel feature map is rich in spatial information, providing high-quality input for the subsequent deep network and improving the detection accuracy.
[0078] In step 203, after obtaining the high-frequency image and the aligned image, the two may be added together to obtain a fused image.
[0079] In this embodiment, the fusion is performed by addition, so that a fused image can be obtained simply and efficiently.
[0080] In step 204, the model structure of the image detection model can be selected as needed, for example, a ConvNeXt model can be selected. The ConvNeXt model is an improved residual network (ResNet) model with stronger feature expression capabilities and high-frequency information sensitivity. It can specifically include an input layer, multiple intermediate layers and an output layer.
[0081] After the fused image is input into the image detection model, the image detection model can output the forged type score and the real type score, and then use the type with the higher score as the image detection result.
[0082] In this embodiment, the image detection results are obtained through the image detection model, and the excellent performance of the deep learning model can be utilized to improve the image detection effect.
[0083] In some embodiments, the method may further generate a heat map, wherein the heat map is used to characterize the forged portion in the image to be detected, wherein a highlighted portion indicates the presence of forgery.
[0084] Specifically, GradCAM can be used to generate a heat map based on the features of the output layer of the image detection model. GradCAM, short for Gradient-weighted Class Activation Mapping, is a technology used to visualize deep learning models. It generates a heat map to show the areas that the model focuses on when making decisions, thereby providing a visual explanation of the model's decision-making process.
[0085] In this embodiment, by generating a heat map, the forged part of the image to be detected can be more intuitively known, thereby improving the visualization level.
[0086] The above image detection method can be applied to a variety of scenarios, such as:
[0087] Application process for social media platforms:
[0088] When users upload photos or videos to the platform, the system automatically invokes a face forgery detection algorithm for real-time detection. Using face alignment and SRM feature extraction, the ConvNeXt model determines whether the image or video is forged. If forgery is detected, the system automatically flags or refuses to publish it, protecting users from false information. This improves the authenticity and security of platform content and reduces the spread of false information.
[0089] Application process of e-government and public services:
[0090] When applying for various government services (such as electronic ID applications and accessing public services), users' uploaded identity photos will be reviewed by a detection algorithm. The system automatically verifies the authenticity of the photos, preventing the use of forged documents and ensuring the legitimacy and security of public services. This improves the security of government services and prevents identity fraud.
[0091] The application process of the online education and examination system: During exam monitoring, the online education platform collects real-time student facial images for forgery detection. Through face alignment and feature extraction, the ConvNeXt model determines whether there is forgery. The system records and notifies the invigilator to ensure exam fairness. This ensures the fairness of online exams and prevents cheating.
[0092] The application process of digital content creation and review: Before publishing, digital content creation platforms conduct forgery detection on uploaded images or videos. Detection algorithms screen for potential forgeries to ensure the authenticity of platform content. Review feedback is provided to help content creators improve the quality of their work. This helps maintain the platform's content ecosystem and prevent the proliferation of fake content.
[0093] Figure 4is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides an image detection model training method, the method comprising:
[0094] 401. Align the sample images to obtain an aligned image.
[0095] 402. Extract high-frequency information from the aligned images to obtain a high-frequency image.
[0096] 403. Fuse the aligned image and the high-frequency image to obtain a fused image.
[0097] 404. Use an image detection model to process the fused image to obtain a prediction result.
[0098] 405. Construct a loss function based on the prediction result and the true value result corresponding to the sample image, and adjust the model parameters of the image detection model based on the loss function.
[0099] In this embodiment, by aligning the images to be detected to obtain aligned images, and performing subsequent processing based on the aligned images, the influence of bending deformation of the images to be detected on the detection results can be avoided, thereby improving the accuracy of the image detection model; by extracting high-frequency information of the aligned images to obtain high-frequency images, and performing subsequent processing based on the high-frequency images, the depth features of the aligned images can be effectively captured, thereby further improving the accuracy of the image detection model; by fusing the aligned images and the high-frequency images to obtain the fused images, thereby further improving the accuracy and effect of the image detection model.
[0100] The sample images refer to images used as training samples and may come from an existing dataset.
[0101] Furthermore, the sample images may include sample images of various forgery types.
[0102] For example, the sample images include the following types of forged images:
[0103] Face Swapping: Replace one face with another to generate a fused image.
[0104] Lip-sync Alignment image: Adjust the lip movements of the characters in the video to match the audio content.
[0105] Expression Synthesis Image: Generates facial images with specific expressions.
[0106] Full-face Generation: Generates full face images using a Generative Adversarial Network (GAN).
[0107] In this embodiment, by acquiring sample images of various forgery types, the generalization ability and robustness of the image detection model can be improved.
[0108] After obtaining the sample image, it can also be labeled to indicate whether it is a forged image, such as using 0 to indicate that the sample image is real and 1 to indicate that the sample image is forged.
[0109] Figure 5 It is an overall architecture diagram of the model training stage provided according to an embodiment of the present disclosure.
[0110] like Figure 5 As shown in the figure, after obtaining the sample image, similar to the reasoning process, after alignment, extraction of high-frequency information, and fusion processing, a fused image is obtained, which is input into the image detection model, and the output is the prediction result.
[0111] The prediction results can specifically include scores for real types and scores for fake types. The scores can specifically be probabilities. Afterwards, a loss function is constructed based on the labeled true value results (0 or 1) and the prediction results. The model parameters of the image detection model are adjusted using the loss function until the end conditions are reached (such as reaching a preset number of iterations or model convergence), and the final image detection model is obtained for the inference process.
[0112] like Figure 5 As shown in the figure, when adjusting model parameters based on the loss function, the exponential moving average (EMA) method can be used to adjust the parameters. EMA is a weighted moving average algorithm.
[0113] In this embodiment, the accuracy and generalization ability of image forgery detection can be improved by training using the EMA method.
[0114] In some embodiments, aligning the sample images to obtain an aligned image includes:
[0115] Extracting key points from the sample image;
[0116] Based on the key points, an affine transformation is performed on the sample image to obtain the aligned image.
[0117] In this embodiment, the aligned image can be obtained simply and efficiently through affine transformation. When the image to be detected is a face image, the influence of facial expression and posture can be avoided, thereby improving the accuracy of the image detection model.
[0118] In some embodiments, extracting high-frequency information from the aligned images to obtain a high-frequency image includes:
[0119] The aligned image is input into the SRM, and the high-frequency information in the aligned image is extracted by the SRM to output the high-frequency image.
[0120] In this embodiment, high-frequency images are obtained through SRM, so that the generated multi-channel feature map is rich in spatial information, providing high-quality input for the subsequent deep network and improving the accuracy of the model.
[0121] In some embodiments, fusing the aligned image and the high-frequency image to obtain a fused image includes:
[0122] The aligned image and the high-frequency image are added to obtain the fused image.
[0123] In this embodiment, the fusion is performed by addition, so that a fused image can be obtained simply and efficiently.
[0124] Figure 6 Schematic diagram of the fourth embodiment of the present disclosure. This embodiment provides an image detection device, the device 600 including: an alignment module 601 , an extraction module 602 , a fusion module 603 and an acquisition module 604 .
[0125] The alignment module 601 is used to align the image to be detected to obtain an aligned image;
[0126] The extraction module 602 is used to extract high-frequency information from the aligned image to obtain a high-frequency image;
[0127] The fusion module 603 is used to fuse the aligned image and the high-frequency image to obtain a fused image;
[0128] The acquisition module 604 is used to acquire image detection results based on the fused image.
[0129] In this embodiment, by aligning the image to be detected to obtain an aligned image, and performing subsequent processing based on the aligned image, the influence of the bending deformation of the image to be detected on the detection result can be avoided, thereby improving the accuracy of the image detection result; by extracting the high-frequency information of the aligned image to obtain a high-frequency image, and performing subsequent processing based on the high-frequency image, the depth features of the aligned image can be effectively captured, thereby further improving the accuracy of the image detection result; by fusing the aligned image and the high-frequency image to obtain a fused image, thereby further improving the accuracy and effect of image detection.
[0130] In some embodiments, the alignment module 601 is further configured to:
[0131] Extracting key points from the image to be detected;
[0132] Based on the key points, an affine transformation is performed on the image to be detected to obtain the aligned image.
[0133] In this embodiment, the aligned image can be obtained simply and efficiently through affine transformation. When the image to be detected is a face image, the influence of facial expression and posture can be avoided, thereby improving the accuracy of the image detection result.
[0134] In some embodiments, the extraction module 602 is further configured to:
[0135] The aligned image is input into the SRM, and the high-frequency information in the aligned image is extracted by the SRM to output the high-frequency image.
[0136] In this embodiment, high-frequency images are obtained through SRM, so that the generated multi-channel feature map is rich in spatial information, providing high-quality input for the subsequent deep network and improving the detection accuracy.
[0137] In some embodiments, the fusion module 603 is further configured to:
[0138] The aligned image and the high-frequency image are added to obtain the fused image.
[0139] In this embodiment, the fusion is performed by addition, so that a fused image can be obtained simply and efficiently.
[0140] In some embodiments, the acquisition module 604 is further configured to:
[0141] The fused image is input into a pre-trained image detection model to output the image detection result.
[0142] In this embodiment, the image detection results are obtained through the image detection model, and the excellent performance of the deep learning model can be utilized to improve the image detection effect.
[0143] In some embodiments, the apparatus 600 further includes:
[0144] The generating module is used to generate a heat map, wherein the heat map is used to characterize the forged part in the image to be detected.
[0145] In this embodiment, by generating a heat map, the forged part of the image to be detected can be more intuitively known, thereby improving the visualization level.
[0146] Figure 7 Schematic diagram of the fifth embodiment of the present disclosure. This embodiment provides an image detection model training device, the device 700 including: an alignment module 701 , an extraction module 702 , a fusion module 703 , an acquisition module 704 and an adjustment module 705 .
[0147] The alignment module 701 is used to align the sample images to obtain an aligned image;
[0148] The extraction module 702 is used to extract high-frequency information from the aligned image to obtain a high-frequency image;
[0149] The fusion module 703 is used to fuse the aligned image and the high-frequency image to obtain a fused image;
[0150] The acquisition module 704 is used to process the fused image using an image detection model to obtain a prediction result;
[0151] The adjustment module 705 is used to construct a loss function based on the prediction result and the true value result corresponding to the sample image, and adjust the model parameters of the image detection model based on the loss function.
[0152] In this embodiment, by aligning the images to be detected to obtain aligned images, and performing subsequent processing based on the aligned images, the influence of bending deformation of the images to be detected on the detection results can be avoided, thereby improving the accuracy of the image detection model; by extracting high-frequency information of the aligned images to obtain high-frequency images, and performing subsequent processing based on the high-frequency images, the depth features of the aligned images can be effectively captured, thereby further improving the accuracy of the image detection model; by fusing the aligned images and the high-frequency images to obtain the fused images, thereby further improving the accuracy and effect of the image detection model.
[0153] In some embodiments, the sample image includes:
[0154] Sample images of various forgery types.
[0155] In this embodiment, by acquiring sample images of various forgery types, the generalization ability and robustness of the image detection model can be improved.
[0156] In some embodiments, the adjustment module 705 is further configured to:
[0157] Based on the loss function, the model parameters of the image detection model are adjusted using the EMA method.
[0158] In this embodiment, the accuracy and generalization ability of image forgery detection can be improved by training using the EMA method.
[0159] In some embodiments, the alignment module 701 is further configured to:
[0160] Extracting key points from the image to be detected;
[0161] Based on the key points, an affine transformation is performed on the image to be detected to obtain the aligned image.
[0162] In this embodiment, the aligned image can be obtained simply and efficiently through affine transformation. When the image to be detected is a face image, the influence of facial expression and posture can be avoided, thereby improving the accuracy of the image detection model.
[0163] In some embodiments, the extraction module 702 is further configured to:
[0164] The aligned image is input into the SRM, and the high-frequency information in the aligned image is extracted by the SRM to output the high-frequency image.
[0165] In this embodiment, high-frequency images are obtained through SRM, so that the generated multi-channel feature map is rich in spatial information, providing high-quality input for the subsequent deep network and improving the accuracy of the model.
[0166] In some embodiments, the fusion module 703 is further configured to:
[0167] The aligned image and the high-frequency image are added to obtain the fused image.
[0168] In this embodiment, the fusion is performed by addition, so that a fused image can be obtained simply and efficiently.
[0169] It can be understood that in the embodiments of the present disclosure, the same or similar contents in different embodiments can be referenced to each other.
[0170] It can be understood that the terms “first”, “second”, etc. in the embodiments of the present disclosure are only used for distinction and do not indicate the degree of importance, time sequence, etc.
[0171] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0172] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0173] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. Electronic device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 800 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0174] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0175] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0176] The computing unit 801 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as image detection methods or image detection model training methods. For example, in some embodiments, the image detection method or image detection model training method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image detection method or image detection model training method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the image detection method or the image detection model training method in any other appropriate manner (for example, by means of firmware).
[0177] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0178] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0179] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0181] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0182] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0183] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0184] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An image detection method, comprising: Align the image to be detected to obtain an aligned image; extracting high-frequency information from the aligned images to obtain a high-frequency image; fusing the aligned image and the high-frequency image to obtain a fused image; Based on the fused image, an image detection result is obtained.
2. The method according to claim 1, wherein The aligning the images to be detected to obtain aligned images includes: Extracting key points from the image to be detected; Based on the key points, an affine transformation is performed on the image to be detected to obtain the aligned image.
3. The method according to claim 1, wherein The extracting high-frequency information from the aligned images to obtain a high-frequency image includes: The aligned image is input into the SRM, and the high-frequency information in the aligned image is extracted by the SRM to output the high-frequency image.
4. The method according to claim 1, wherein The fusing the aligned image and the high-frequency image to obtain a fused image includes: The aligned image and the high-frequency image are added to obtain the fused image.
5. The method according to claim 1, wherein The obtaining of an image detection result based on the fused image includes: The fused image is input into a pre-trained image detection model to output the image detection result.
6. The method according to claim 1, further comprising: A heat map is generated, where the heat map is used to characterize the forged portion in the image to be detected.
7. A method for training an image detection model, comprising: aligning the sample images to obtain an aligned image; extracting high-frequency information from the aligned images to obtain a high-frequency image; fusing the aligned image and the high-frequency image to obtain a fused image; Using an image detection model to process the fused image to obtain a prediction result; A loss function is constructed based on the prediction result and the true value result corresponding to the sample image, and model parameters of the image detection model are adjusted based on the loss function.
8. The method according to claim 7, wherein: The sample images include: Sample images of various forgery types.
9. The method according to claim 7, wherein: The adjusting the model parameters of the image detection model based on the loss function includes: Based on the loss function, the model parameters of the image detection model are adjusted using the EMA method.
10. The method according to any one of claims 7 to 9, wherein: The step of aligning the sample images to obtain an aligned image includes: Extracting key points from the sample image; Based on the key points, an affine transformation is performed on the sample image to obtain the aligned image.
11. The method according to any one of claims 7 to 9, wherein: The extracting high-frequency information from the aligned images to obtain a high-frequency image includes: The aligned image is input into the SRM, and the high-frequency information in the aligned image is extracted by the SRM to output the high-frequency image.
12. The method according to any one of claims 7 to 9, wherein: The fusing the aligned image and the high-frequency image to obtain a fused image includes: The aligned image and the high-frequency image are added to obtain the fused image.
13. An image detection device, comprising: An alignment module is used to align the image to be detected to obtain an aligned image; an extraction module, configured to extract high-frequency information from the aligned image to obtain a high-frequency image; a fusion module, configured to fuse the aligned image and the high-frequency image to obtain a fused image; The acquisition module is used to acquire the image detection result based on the fused image.
14. A method for training an image detection model, comprising: an alignment module, configured to align sample images to obtain an aligned image; an extraction module, configured to extract high-frequency information from the aligned image to obtain a high-frequency image; a fusion module, configured to fuse the aligned image and the high-frequency image to obtain a fused image; An acquisition module, configured to process the fused image using an image detection model to obtain a prediction result; An adjustment module is used to construct a loss function based on the prediction result and the true value result corresponding to the sample image, and adjust the model parameters of the image detection model based on the loss function.
15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Face image detection method and device, model training method and device and storage medium
CN114913565A
Deep counterfeit video technology traceability method based on image frequency domain information
CN115188039A
Face deep false detection method based on multi-modal feature fusion
CN115880749A
Image detection method, electronic equipment and storage medium
CN116228644A