Image detection method, model training method, device and electronic device

By employing parallel live and synthetic image detection modules within image detection networks, the method addresses the efficiency decline in large-scale image detection, enhancing accuracy and speed in identifying real human faces.

CN115984941BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310071648.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2025-07-15
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

As the detectable range of the image detection network expands, the image detection efficiency decreases in the prior art and the complexity of the image detection process increases.

Method used

Using a live body detection module and a synthetic image detection module set in parallel, the output feature sequence of the image to be detected is obtained through the object detection module, and whether the image is a real face image based on the target feature sequence is determined, thereby simplifying the detection process.

Benefits of technology

This improves image detection efficiency, reduces the complexity of the detection process, and improves the accuracy of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984941B_ABST
    Figure CN115984941B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image detection method, a model training method, an apparatus, and an electronic device, which relate to the field of artificial intelligence technologies, particularly to the fields of deep learning, image processing, and computer vision technologies, and can be applied to scenarios such as face image detection. The specific implementation solution is as follows: input the image to be detected into an image detection network, where the image detection network includes at least one image detection model; sequentially in each image detection model, determine the target detection module corresponding to the image to be detected from a live detection module and a synthetic image detection module; through the target detection module, obtain the output feature sequence of the image to be detected; based on the target feature sequence, obtain the detection result of the image to be detected, where the target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image. Using the present disclosure can improve the image detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular to the fields of deep learning, image processing, and computer vision technologies, and can be applied to scenarios such as face image detection. Background Art

[0002] With the development of artificial intelligence technologies, the detectable range of image detection networks is continuously expanding. For example, the detectable objects can already include attack images and synthetic images. However, with the continuous expansion of the detectable range, the image detection efficiency decreases accordingly. Summary of the Invention

[0003] The present disclosure provides an image detection method, a model training method, an apparatus, and an electronic device.

[0004] According to a first aspect of the present disclosure, there is provided an image detection method, including:

[0005] Inputting an image to be detected into an image detection network, where when the image detection network includes at least one image detection model and when the image detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model;

[0006] Successively in each image detection model, determining a target detection module corresponding to the image to be detected from a live detection module and a synthetic image detection module;

[0007] Obtaining an output feature sequence of the image to be detected through the target detection module;

[0008] Based on the target feature sequence, obtaining a detection result of the image to be detected, where the target feature sequence is an output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0009] According to a second aspect of the present disclosure, there is provided a model training method, including:

[0010] Obtaining a training data set, where the training data set includes a plurality of image samples, and a module selection label and an image property label corresponding to each image sample, the module selection label is used to indicate the target detection module corresponding to the image sample, and the image property label is used to indicate whether the image sample is a real face image;

[0011] The initial detection network is trained using a training dataset to obtain an image detection network. When the initial detection network includes at least one image detection model and when the initial detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input to the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel.

[0012] According to a third aspect of the present disclosure, there is provided an image detection device, including:

[0013] An image input unit for inputting an image to be detected into the image detection network. When the image detection network includes at least one image detection model and when the image detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input to the next image detection model;

[0014] A module selection unit for sequentially determining, in each image detection model, a target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module;

[0015] A feature output unit for obtaining an output feature sequence of the image to be detected through the target detection module;

[0016] An image detection unit for obtaining a detection result of the image to be detected based on the target feature sequence. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0017] According to a fourth aspect of the present disclosure, there is provided a model training device, including:

[0018] A data acquisition unit for acquiring a training dataset. The training dataset includes a plurality of image samples, as well as a module selection label and an image property label corresponding to each image sample. The module selection label is used to indicate the target detection module corresponding to the image sample, and the image property label is used to indicate whether the image sample is a real face image;

[0019] A network training unit for training the initial detection network using the training dataset to obtain an image detection network. When the initial detection network includes at least one image detection model and when the initial detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input to the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel.

[0020] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0021] At least one processor;

[0022] A memory communicatively connected to the at least one processor;

[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in the first aspect.

[0024] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided in the first aspect.

[0025] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method provided in the first aspect when executed by a processor.

[0026] Adopting the present disclosure can improve the efficiency of image detection.

[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0029] Figure 1 is a schematic diagram of an image detection process in the prior art;

[0030] Figure 2 is a schematic diagram of the process of an image detection method provided by an embodiment of the present disclosure;

[0031] Figure 3 is a schematic diagram of the working process of an image detection network provided by an embodiment of the present disclosure;

[0032] Figure 4 is a schematic diagram of the working process of a current detection model provided by an embodiment of the present disclosure;

[0033] Figure 5 is a schematic diagram of the complete process of an image detection method provided by an embodiment of the present disclosure;

[0034] Figure 6 is a schematic diagram of the application scenario of an image detection method provided by an embodiment of the present disclosure;

[0035] Figure 7 is a schematic diagram of the process of a model training method provided by an embodiment of the present disclosure;

[0036] Figure 8 Schematic structural block diagram of an image detection device provided by an embodiment of the present disclosure;

[0037] Figure 9 Schematic structural block diagram of a model training device provided by an embodiment of the present disclosure;

[0038] Figure 10 Schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0040] As described in the background art, with the development of artificial intelligence technology, the detectable range of image detection networks is continuously expanding. For example, detectable objects can already include attack images and synthetic images. Please refer to Figure 1 , in the existing image detection network 100, the live detection module 101 and the synthetic image detection module 102 of each image detection model are usually set serially. It is necessary to first detect the image to be detected through the live detection module 101 to obtain the live detection result, and then detect the image to be detected through the synthetic image detection module 102 to obtain the synthetic image detection result. Finally, combining the live detection result and the synthetic image detection result, it is determined whether the image to be detected is a real face image. For example, when the live detection result is a real face image and the synthetic image detection result is also a real face image, it is determined that the image to be detected is a real face image; otherwise, it is determined that the image to be detected is not a real face image.

[0041] Obviously, in the prior art, it is necessary to detect the image to be detected through the live detection module 101 and the synthetic image detection module 102 in sequence, which increases the complexity of the image detection process and the image detection efficiency is relatively low.

[0042] Based on the above background, an embodiment of the present disclosure provides an image detection method, and this image detection method can be applied to an electronic device. The following will refer to Figure 2 the following process schematic diagram to describe an image detection method provided by an embodiment of the present disclosure. It should be noted that although the logical order is shown in the process schematic diagram, in some cases, the steps shown or described can also be executed in other orders.

[0043] Step S201: Input the image to be detected into the image detection network. When the image detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model.

[0044] Step S202: In each image detection model in sequence, determine the target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module.

[0045] Step S203: Obtain the output feature sequence of the image to be detected through the target detection module.

[0046] Step S204: Based on the target feature sequence, obtain the detection result of the image to be detected. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0047] Among them, the image to be detected can be a real face image, or an attack face image such as a video image, a printed photo, a hole-punched photo, a 3D mask, etc., or a synthetic face image generated by means of software editing, artificial intelligence editing, etc.

[0048] Please combine Figure 3 , in the embodiments of the present disclosure, the image detection network 300 may include at least one image detection model 301. Each image detection model 301 includes a live detection module 3011 and a synthetic image detection module 3012 arranged in parallel. When the image detection network 300 includes at least two image detection models 301, the output of the previous image detection model 301 among two adjacent image detection models 301 is used as the input of the next image detection model 301. That is, when the image detection network 300 includes at least two image detection models 301, the at least two image detection models 301 are arranged in series. In a specific example, the image detection network 300 may include 12 image detection models 301.

[0049] Based on the above description, it can be understood that in the embodiments of the present disclosure, when the current detection model is the first image detection model in the image detection network, the input feature sequence corresponding to the current detection model may include at least two image feature blocks obtained by splitting the image to be detected, or may include at least two image feature blocks obtained by splitting the image to be detected, and a learnable embedding vector (Class Token) - category feature block; when the current detection model is a non-first image detection model in the image detection network, the input feature sequence corresponding to the current detection model is the output feature sequence obtained by the previous image detection model of the current detection model.

[0050] After inputting the image to be detected into the image detection network, in each image detection model in sequence, a target detection module corresponding to the image to be detected can be determined from the live detection module and the synthesized image detection module, and an output feature sequence of the image to be detected can be obtained through the target detection module. Specifically, in the embodiments of the present disclosure, the current detection model can determine a target detection module corresponding to the image to be detected from the live detection module and the synthesized image detection module according to the input feature sequence corresponding to the current detection model, and through the target detection module, obtain an output feature sequence of the image to be detected according to the input feature sequence corresponding to the current detection model.

[0051] Finally, based on the target feature sequence, a detection result of the image to be detected is obtained. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image. Specifically, in the embodiments of the present disclosure, when it is determined from the last image detection model that the target detection module corresponding to the image to be detected is the live detection module, the detection result is used to indicate that the image to be detected is a real face image or an attacked face image; when it is determined from the last image detection model that the target detection module corresponding to the image to be detected is the synthesized image detection module, the detection result is used to indicate that the image to be detected is a real face image or a synthesized face image.

[0052] By using the image detection method provided in the embodiments of the present disclosure, in each image detection model, a target detection module corresponding to the image to be detected can be determined from the live detection module and the synthesized image detection module, and an output feature sequence of the image to be detected can be obtained through the target detection module. Then, based on the target feature sequence, a detection result of the image to be detected is obtained to indicate whether the image to be detected is a real face image. That is, in the embodiments of the present disclosure, an output feature sequence of the image to be detected can be obtained only through one of the live detection module and the synthesized image detection module, and then based on the target feature sequence, a detection result of the image to be detected is obtained, rather than detecting the image to be detected through the live detection module and the synthesized image detection module in sequence as in the prior art. Therefore, the complexity of the image detection process can be reduced, thereby improving the image detection efficiency.

[0053] In addition, it should be noted that in the embodiments of the present disclosure, the image detection network can be trained through the following steps:

[0054] Obtain a training data set. The training data set includes a plurality of image samples, as well as a module selection label and an image property label corresponding to each image sample. The module selection label is used to indicate the target detection module corresponding to the image sample, and the image property label is used to indicate whether the image sample is a real face image;

[0055] The initial detection network is trained with a training dataset to obtain an image detection network. When the initial detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module that are arranged in parallel.

[0056] Among them, the training dataset includes a plurality of positive samples. The positive samples belong to real face images. Their module selection labels are used to indicate that the target detection module corresponding to the image sample is a live detection module or a synthetic image detection module, and their image property labels are used to indicate that the image sample is a real face image. For example, the training dataset includes 10,000 positive samples. Among them, the module selection labels of 5,000 positive samples are used to indicate that the target detection module corresponding to the image sample is a live detection module, and their image property labels are used to indicate that the image sample is a real face image. For the other 5,000 positive samples, the module selection labels are used to indicate that the target detection module corresponding to the image sample is a synthetic image detection module, and their image property labels are used to indicate that the image sample is a real face image.

[0057] The training dataset also includes a plurality of first negative-positive samples. The first negative-positive samples belong to attack face images. Their module selection labels are used to indicate that the target detection module corresponding to the image sample is a live detection module, and their image property labels are used to indicate that the image sample is an attack face image.

[0058] The training dataset also includes a plurality of second negative-positive samples. The second negative-positive samples belong to synthetic face images. Their module selection labels are used to indicate that the target detection module corresponding to the image sample is a synthetic image detection module, and their image property labels are used to indicate that the image sample is a synthetic face image.

[0059] In some optional embodiments, "sequentially in each image detection model, determining the target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module" includes the following steps:

[0060] For the current detection model in at least one image detection model, obtain the input feature sequence corresponding to the current detection model;

[0061] Through the current detection model, based on the input feature sequence, determine the detection category of the image to be detected;

[0062] According to the detection category, determine the target detection module from the live detection module and the synthetic image detection module included in the current detection model.

[0063] In a specific example, when the current detection model is the first image detection model in the image detection network, the input feature sequence corresponding to the current detection model includes at least two image feature blocks obtained by slicing the image to be detected. Based on this, in the embodiments of the present disclosure, "determining the detection category of the image to be detected through the current detection model based on the input feature sequence" may include: performing feature extraction on the input feature sequence to obtain an intermediate feature sequence; calculating the mean of the corresponding feature points of each feature block in the intermediate feature sequence to obtain a third target feature block; performing binary classification on the third target feature block to obtain a third classification result; performing normalization processing on the third classification result to determine the detection category of the image to be detected, and the detection category is the attack face image detection category or the synthetic face image detection category.

[0064] In another specific example, when the current detection model is the first image detection model in the image detection network, the input feature sequence corresponding to the current detection model includes at least two image feature blocks obtained by slicing the image to be detected, and a learnable Class Token - category feature block. Based on this, in the embodiments of the present disclosure, "determining the detection category of the image to be detected through the current detection model based on the input feature sequence" may include: performing feature extraction on the input feature sequence to obtain an intermediate feature sequence; determining a first target feature block corresponding to the category feature block from the intermediate feature sequence; performing binary classification on the first target feature block to obtain a first classification result; performing normalization processing on the first classification result to determine the detection category of the image to be detected, and the detection category is the attack face image detection category or the synthetic face image detection category.

[0065] When the detection category is the attack face image detection category, determine that the target detection module corresponding to the image to be detected is the live detection module; when the detection category is the synthetic face image detection category, determine that the target detection module corresponding to the image to be detected is the synthetic image detection module.

[0066] Through the above steps included in "successively in each image detection model, determining the target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module", in the embodiments of the present disclosure, for the current detection model in at least one image detection model, the input feature sequence corresponding to the current detection model can be obtained, and then through the current detection model, based on the input feature sequence, the detection category of the image to be detected can be determined, so as to determine the target detection module from the live detection module and the synthetic image detection module included in the current detection model according to the detection category, without introducing other feature sequences for determining the detection category. On the one hand, the accuracy of the detection category can be improved based on the high correlation between the input feature sequence and the detection category, thereby improving the accuracy of the detection result. On the other hand, the complexity of the image detection process can be further reduced, thereby improving the image detection efficiency.

[0067] In some alternative embodiments, to further improve the efficiency of image detection, when the current detection model is the first image detection model in the image detection network, the input feature sequence corresponding to the current detection model includes at least two image feature blocks obtained by splitting the image to be detected, and a learnable Class Token, that is, "obtaining the input feature sequence corresponding to the current detection model" may include the following steps:

[0068] When the current detection model is the first image detection model in the image detection network, split the image to be detected to obtain at least two image feature blocks, and form the input feature sequence corresponding to the current detection model through the at least two image feature blocks and the class feature block;

[0069] When the current detection model is not the first image detection model in the image detection network, obtain the output feature sequence obtained by the previous image detection model of the current detection model as the input feature sequence corresponding to the current detection model.

[0070] In the embodiments of the present disclosure, first, the face key points of the image to be detected can be aligned, then the face part of the image can be intercepted, and the size of the face part of the image can be adjusted, for example, adjusted to a size of 224*224 pixels, and then the face part of the image can be split to obtain at least two image feature blocks, for example, 5, 9, 16, or even more image feature blocks, and the embodiments of the present disclosure do not make specific limitations on this. Thereafter, the at least two image feature blocks and the class feature block are used to form the input feature sequence corresponding to the current detection model.

[0071] Through the above steps included in "obtaining the input feature sequence corresponding to the current detection model", in the embodiments of the present disclosure, when "determining the detection category of the image to be detected based on the input feature sequence through the current detection model", the input feature sequence can be feature-extracted to obtain an intermediate feature sequence; the first target feature block corresponding to the class feature block is determined from the intermediate feature sequence; the first target feature block is subjected to binary classification to obtain a first classification result; the first classification result is normalized to determine the detection category of the image to be detected, instead of calculating the mean value of the feature points corresponding to each feature block in the intermediate feature sequence to obtain a third target feature block; the third target feature block is subjected to binary classification to obtain a third classification result; the third classification result is normalized to determine the detection category of the image to be detected, which can not only reduce the computing power consumption when determining the detection category, but also reduce the calculation consumption time, thereby further improving the efficiency of image detection.

[0072] As described above, when "obtaining the input feature sequence corresponding to the current detection model" includes the above steps, "determining the detection category of the image to be detected based on the input feature sequence" may include the following steps:

[0073] Perform feature extraction on the input feature sequence to obtain an intermediate feature sequence;

[0074] Determine a first target feature block corresponding to the category feature block from the intermediate feature sequence;

[0075] Perform binary classification on the first target feature block to obtain a first classification result;

[0076] Perform normalization on the first classification result to determine the detection category of the image to be detected.

[0077] In a specific example, the above steps included in "determining the detection category of the image to be detected based on the input feature sequence" can be implemented based on the Transformer model. For example, an adaptive mixture of experts module is introduced into the Transformer model.

[0078] Please combine with Figure 4 In the embodiments of the present disclosure, first, the input feature sequence can be subjected to "layer-normalization" processing, that is, Norm processing, and then feature extraction is performed on the input feature sequence after Norm processing to obtain a first sequence to be processed. For example, the input feature sequence after Norm processing can be subjected to feature extraction through a multi-head attention module to obtain a first sequence to be processed. Thereafter, the first sequence to be processed can be subjected to residual connection with the input feature sequence before Norm processing to obtain a second sequence to be processed, and then Norm processing is performed on the second sequence to be processed, and the second sequence to be processed after Norm processing is used as the intermediate feature sequence. Finally, in the adaptive mixture of experts module, a first target feature block corresponding to the category feature block is determined from the intermediate feature sequence, then binary classification is performed on the first target feature block to obtain a first classification result, and normalization is performed on the first classification result to determine the detection category of the image to be detected. Specifically, in the embodiments of the present disclosure, binary classification can be performed on the first target feature block through a first adaptive gating unit to obtain a first classification result, and then the first classification result is normalized through a first normalization exponential function to determine the detection category of the image to be detected. Among them, the first adaptive gating unit can be a fully connected layer, and the first normalization exponential function can be a Softmax function. The embodiments of the present disclosure do not make specific limitations on this.

[0079] After determining the detection category of the image to be detected, in the adaptive mixture-of-experts module, the target detection module can be determined from the live detection module and the synthesized image detection module included in the current detection model according to the detection category, and then the output feature sequence of the image to be detected can be obtained through the target detection module. Specifically, in the embodiments of the present disclosure, after obtaining the output feature sequence of the image to be detected, the output feature sequence can be subjected to residual connection with the second sequence to be processed as the output feature sequence finally obtained by the current detection model.

[0080] Based on the above description, it can be understood that in the embodiments of the present disclosure, the adaptive mixture-of-experts module includes a first adaptive gating unit for performing binary classification on the first target feature block to obtain a first classification result, a normalization processing module for normalizing the first classification result through a first sigmoid function, a live detection module, and a synthesized image detection module.

[0081] In addition, it should be noted that in the embodiments of the present disclosure, in the process of "training the initial detection network with a training data set to obtain an image detection network", for the first adaptive gating unit, a binary cross-entropy loss function can be specifically used to train it so that it can perform accurate binary classification on the first target feature block to obtain a first classification result.

[0082] Through the above steps included in "determining the detection category of the image to be detected based on the input feature sequence", in the embodiments of the present disclosure, on the one hand, the input feature sequence can be subjected to feature extraction to obtain an intermediate feature sequence, which can improve the representability of the intermediate feature sequence to further improve the accuracy of the detection category, thereby improving the accuracy of the detection result. On the other hand, the first target feature block corresponding to the category feature block can be determined from the intermediate feature sequence, and then the first target feature block can be subjected to binary classification to obtain a first classification result, and the first classification result can be normalized to determine the detection category of the image to be detected, without calculating the mean value of the corresponding feature points of each feature block in the intermediate feature sequence to obtain a third target feature block; performing binary classification on the third target feature block to obtain a third classification result; normalizing the third classification result to determine the detection category of the image to be detected, which can not only reduce the computing power consumption when determining the detection category, but also reduce the calculation time consumption, thereby further improving the image detection efficiency.

[0083] As described above, the above steps included in "determining the detection category of the image to be detected based on the input feature sequence" can be implemented based on a Transformer model. Based on this, "performing feature extraction on the input feature sequence to obtain an intermediate feature sequence" can include the following steps:

[0084] Through the multi-head attention module, based on the attention relationships among the feature blocks in the input feature sequence, feature extraction is performed on the input feature sequence to obtain an intermediate feature sequence.

[0085] Among them, the multi-head attention module can be a feature extraction model implemented based on the multi-head attention mechanism, that is, the Multi-Head Attention mechanism. It can perform feature extraction on the input feature sequence based on the attention relationships among the feature blocks in the input feature sequence to obtain an intermediate feature sequence.

[0086] Through the above steps included in "performing feature extraction on the input feature sequence to obtain an intermediate feature sequence", in the embodiments of the present disclosure, the multi-head attention module can perform feature extraction on the input feature sequence based on the attention relationships among the feature blocks in the input feature sequence to obtain an intermediate feature sequence. The multi-head attention model expands the model's ability to focus on different image positions and provides multiple "representation subspaces" of the attention layer. Therefore, it helps to capture richer features, can further improve the representability of the intermediate feature sequence, and further improve the accuracy of the detected categories, thereby improving the accuracy of the detection results.

[0087] In addition, it should be noted that in the embodiments of the present disclosure, in the process of "training the initial detection network with a training dataset to obtain an image detection network", since the multi-head attention module is common to both the live detection task and the synthetic graph detection task, similar features of the positive samples used in the two tasks can be learned. That is, for the positive samples used in the two tasks, feature complementarity can be performed, which can further improve the representability of the intermediate feature sequence, and further improve the accuracy of the detected categories, thereby improving the accuracy of the detection results.

[0088] In some alternative embodiments, "obtaining an output feature sequence of the image to be detected through the object detection module" may include the following steps:

[0089] Through the object detection module, based on the correlation between each feature block in the intermediate feature sequence and the detected category, feature extraction is performed on the intermediate feature sequence to obtain an output feature sequence.

[0090] In the embodiments of the present disclosure, the live detection module and the synthetic graph detection module can be two expert networks with the same network structure. For example, they can both be multi-layer perceptrons, that is, both the live detection module and the synthetic graph detection module can include multiple fully connected layers.

[0091] In addition, in the embodiments of the present disclosure, through training, the live detection module can perform feature extraction on the intermediate feature sequence based on the correlation between each feature block in the intermediate feature sequence and the face attack image detection category to obtain an output feature sequence, where the intermediate feature sequence is obtained by performing feature extraction on the input feature sequence. For example, through training, the live detection module can perform feature extraction on the intermediate feature sequence based on the correlation between each feature block in the intermediate feature sequence and the face attack image features such as moiré patterns, imaging deformities, and abnormal reflectivity, to obtain unique features related to the live detection task as the output feature sequence.

[0092] Similarly, in the embodiments of the present disclosure, through training, the synthetic image detection module can perform feature extraction on the intermediate feature sequence based on the correlation between each feature block in the intermediate feature sequence and the synthetic face image detection category to obtain unique features related to the synthetic image detection task as the output feature sequence. For example, through training, the synthetic image detection module can perform feature extraction on the intermediate feature sequence based on the correlation between each feature block in the intermediate feature sequence and the face synthetic image features such as artifacts, to obtain unique features related to the synthetic image detection task as the output feature sequence.

[0093] Through the above steps included in "obtaining the output feature sequence of the image to be detected through the target detection module", in the embodiments of the present disclosure, the target detection module can perform feature extraction on the intermediate feature sequence based on the correlation between each feature block in the intermediate feature sequence and the detection category to obtain the output feature sequence, so as to ensure that the output feature sequence has a high correlation with the detection category, facilitating further improvement of the accuracy of the detection result when obtaining the detection result of the image to be detected based on the target feature sequence in the subsequent process.

[0094] In some alternative embodiments, "obtaining the detection result of the image to be detected based on the target feature sequence" includes the following steps:

[0095] Determine the second target feature block corresponding to the category feature block from the target feature sequence;

[0096] Perform binary classification processing on the second target feature block to obtain a second classification result;

[0097] Perform normalization processing on the second classification result to obtain the detection result of the image to be detected.

[0098] In an embodiment of the present disclosure, after determining a second target feature block corresponding to a category feature block from a target feature sequence, a second adaptive gating unit may be used to perform binary classification processing on the second target feature block to obtain a second classification result, and then a second normalization exponential function may be used to perform normalization processing on the second classification result to obtain a detection result of the image to be detected. Among them, the second adaptive gating unit may be a fully connected layer, and the second normalization exponential function may be a Softmax function. The embodiments of the present disclosure do not make specific limitations thereto.

[0099] In addition, it should be noted that in an embodiment of the present disclosure, in the process of "training an initial detection network with a training data set to obtain an image detection network", for the second adaptive gating unit, a binary cross-entropy loss function may be specifically used to train it so that it can accurately perform binary classification processing on the second target feature block to obtain a second classification result.

[0100] Through the above steps included in "obtaining a detection result of the image to be detected based on the target feature sequence", in an embodiment of the present disclosure, a second target feature block corresponding to a category feature block may be determined from the target feature sequence, then binary classification processing may be performed on the second target feature block to obtain a second classification result, and normalization processing may be performed on the second classification result to obtain a detection result of the image to be detected, without calculating the mean value of the corresponding feature points of each feature block in the target feature sequence to obtain a fourth target feature block; performing binary classification processing on the fourth target feature block to obtain a fourth classification result; and performing normalization processing on the fourth classification result to obtain a detection result of the image to be detected, which can not only reduce the computing power consumption when obtaining the detection result to be detected, but also reduce the calculation consumption time, thereby further improving the image detection efficiency.

[0101] The following will be combined with Figure 5 to illustrate the complete process of an image detection method provided by an embodiment of the present disclosure.

[0102] Step S501, input the image to be detected into the image detection network. When the image detection network includes at least one image detection model and the image detection network includes at least two image detection models, the output of the previous image detection model in two adjacent image detection models is used as the input of the next image detection model.

[0103] Step S502: For the current detection model among at least one image detection model, when the current detection model is the first image detection model in the image detection network, the image to be detected is segmented to obtain at least two image feature blocks, and an input feature sequence corresponding to the current detection model is formed by the at least two image feature blocks and the category feature block; when the current detection model is a non-first image detection model in the image detection network, the output feature sequence obtained by the previous image detection model of the current detection model is acquired as the input feature sequence corresponding to the current detection model.

[0104] Step S503: The input feature sequence is subjected to feature extraction through the current detection model to obtain an intermediate feature sequence, then the first target feature block corresponding to the category feature block is determined from the intermediate feature sequence, and binary classification processing is performed on the first target feature block to obtain a first classification result, and normalization processing is performed on the first classification result to determine the detection category of the image to be detected.

[0105] Step S504: According to the detection category, a target detection module is determined from the live detection module and the synthetic image detection module included in the current detection model, and the intermediate feature sequence is subjected to feature extraction through the target detection module based on the correlation between each feature block in the intermediate feature sequence and the detection category to obtain an output feature sequence.

[0106] Step S505: The second target feature block corresponding to the category feature block is determined from the target feature sequence, then binary classification processing is performed on the second target feature block to obtain a second classification result, and normalization processing is performed on the second classification result to obtain the detection result of the image to be detected. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0107] Please refer to Figure 6 , which is a schematic diagram of an application scenario of an image detection method provided by an embodiment of the present disclosure.

[0108] As mentioned above, the image detection method provided by an embodiment of the present disclosure is applied to an electronic device. Among them, the electronic device is intended to represent various forms of digital computers, such as servers, blade servers, workstations, mainframe computers, desktop computers, laptop computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0109] In an embodiment of the present disclosure, the electronic device can be used to execute the image detection method:

[0110] Input the image to be detected into an image detection network. When the image detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model;

[0111] In each image detection model in sequence, determine the target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module;

[0112] Through the target detection module, obtain the output feature sequence of the image to be detected;

[0113] Based on the target feature sequence, obtain the detection result of the image to be detected. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0114] It should be noted that in the embodiments of the present disclosure, the image to be detected can be obtained through an image acquisition device, and the image acquisition device can be built into an electronic device or independent of the electronic device. The embodiments of the present disclosure do not make specific limitations on this. In addition, in the embodiments of the present disclosure, the image detection method can be applied to scenarios such as face image detection, and specifically can involve many scenarios such as security, attendance, finance, access control, etc.

[0115] In addition, it should also be noted that in the embodiments of the present disclosure, Figure 6 The scene schematic diagram shown is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 6 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0116] The embodiments of the present disclosure provide a model training method, which can be applied to an electronic device to obtain the image detection network described in the embodiments of the foregoing image detection method. Hereinafter, a model training method provided by the embodiments of the present disclosure will be described with reference to Figure 7 the flow schematic diagram shown. It should be noted that although the logical order is shown in the flow schematic diagram, in some cases, the steps shown or described can also be executed in other orders.

[0117] Step S701, obtain a training data set. The training data set includes multiple image samples, as well as a module selection label and an image property label corresponding to each image sample. The module selection label is used to indicate the target detection module corresponding to the image sample, and the image property label is used to indicate whether the image sample is a real face image;

[0118] Step S702: Train the initial detection network with the training dataset to obtain an image detection network. When the initial detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel.

[0119] Among them, the training dataset includes multiple positive samples. The positive samples belong to real face images. Its module selection label is used to indicate that the target detection module corresponding to the image sample is a live detection module or a synthetic image detection module, and its image property label is used to indicate that the image sample is a real face image. For example, the training dataset includes 10,000 positive samples. Among them, the module selection labels of 5,000 positive samples are used to indicate that the target detection module corresponding to the image sample is a live detection module, and its image property label is used to indicate that the image sample is a real face image. For the other 5,000 positive samples, the module selection labels are used to indicate that the target detection module corresponding to the image sample is a synthetic image detection module, and its image property label is used to indicate that the image sample is a real face image.

[0120] The training dataset also includes multiple first negative-positive samples. The first negative-positive samples belong to attack face images. Its module selection label is used to indicate that the target detection module corresponding to the image sample is a live detection module, and its image property label is used to indicate that the image sample is an attack face image.

[0121] The training dataset also includes multiple second negative-positive samples. The second negative-positive samples belong to synthetic face images. Its module selection label is used to indicate that the target detection module corresponding to the image sample is a synthetic image detection module, and its image property label is used to indicate that the image sample is a synthetic face image.

[0122] In addition, in the embodiments of the present disclosure, the specific process of "training the initial detection network with the training dataset to obtain an image detection network" can refer to the relevant part descriptions in the foregoing embodiments of the image detection method, which will not be elaborated here.

[0123] By using the model training method provided in the embodiments of the present disclosure, a training data set including multiple image samples, as well as a module selection label and an image property label corresponding to each image sample can be obtained. Then, the initial detection network is trained with the training data set to obtain an image detection network. When the initial detection network includes at least one image detection model and the initial detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model. Each image detection model includes a live detection module and a composite image detection module arranged in parallel. In this way, based on the image detection network, the output feature sequence of the image to be detected can be obtained only through one of the live detection module and the composite image detection module, and then the detection result of the image to be detected can be obtained based on the target feature sequence, rather than detecting the image to be detected through the live detection module and the composite image detection module successively as in the prior art. Therefore, the complexity of the image detection process can be reduced, thereby improving the image detection efficiency.

[0124] To better implement the above image detection method, the embodiments of the present disclosure further provide an image detection device 800, which can be integrated in an electronic device. Hereinafter, with reference to Figure 8 the structural schematic diagram shown, an image detection device 800 provided by the embodiments of the present disclosure will be described.

[0125] The image detection device 800 includes:

[0126] An image input unit 801, configured to input the image to be detected into the image detection network. When the image detection network includes at least one image detection model and the image detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model;

[0127] A module selection unit 802, configured to sequentially determine, in each image detection model, a target detection module corresponding to the image to be detected from the live detection module and the composite image detection module;

[0128] A feature output unit 803, configured to obtain the output feature sequence of the image to be detected through the target detection module;

[0129] An image detection unit 804, configured to obtain the detection result of the image to be detected based on the target feature sequence. The target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

[0130] In some alternative embodiments, the module selection unit 802 is configured to:

[0131] For the current detection model in at least one image detection model, obtain the input feature sequence corresponding to the current detection model;

[0132] Through the current detection model, based on the input feature sequence, determine the detection category of the image to be detected;

[0133] According to the detection category, determine the target detection module from the live detection module and the synthetic image detection module included in the current detection model.

[0134] In some alternative embodiments, the module selection unit 802 is configured to:

[0135] In the case where the current detection model is the first image detection model in the image detection network, split the image to be detected to obtain at least two image feature blocks, and form the input feature sequence corresponding to the current detection model through the at least two image feature blocks and the category feature block;

[0136] In the case where the current detection model is a non-first image detection model in the image detection network, obtain the output feature sequence obtained by the previous image detection model of the current detection model as the input feature sequence corresponding to the current detection model.

[0137] In some alternative embodiments, the module selection unit 802 is configured to:

[0138] Extract features from the input feature sequence to obtain an intermediate feature sequence;

[0139] Determine the first target feature block corresponding to the category feature block from the intermediate feature sequence;

[0140] Perform binary classification on the first target feature block to obtain a first classification result;

[0141] Normalize the first classification result to determine the detection category of the image to be detected.

[0142] In some alternative embodiments, the module selection unit 802 is configured to:

[0143] Through the multi-head attention module, based on the attention relationship between the feature blocks in the input feature sequence, extract features from the input feature sequence to obtain an intermediate feature sequence.

[0144] In some alternative embodiments, the feature output unit 803 is configured to:

[0145] Through the target detection module, based on the correlation between each feature block in the intermediate feature sequence and the detection category, extract features from the intermediate feature sequence to obtain an output feature sequence.

[0146] In some alternative embodiments, the image detection unit 804 is configured to:

[0147] Determine a second target feature block corresponding to the category feature block from the target feature sequence;

[0148] Perform binary classification processing on the second target feature block to obtain a second classification result;

[0149] Perform normalization processing on the second classification result to obtain a detection result of the image to be detected.

[0150] In specific implementation, each of the above modules may be implemented as an independent entity, or may be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules, reference may be made to the foregoing embodiments of the image detection method, which will not be elaborated herein.

[0151] By using the image detection device 800 provided in the embodiments of the present disclosure, in each image detection model, a target detection module corresponding to the image to be detected can be determined from the live detection module and the synthetic image detection module, and an output feature sequence of the image to be detected can be obtained through the target detection module. Then, based on the target feature sequence, a detection result of the image to be detected can be obtained to indicate whether the image to be detected is a real face image. That is, in the embodiments of the present disclosure, an output feature sequence of the image to be detected can be obtained only through one of the live detection module and the synthetic image detection module, and then based on the target feature sequence, a detection result of the image to be detected can be obtained, rather than, as in the prior art, the image to be detected needs to be detected by the live detection module and the synthetic image detection module in sequence. Therefore, the complexity of the image detection process can be reduced, thereby improving the image detection efficiency.

[0152] To better implement the above model training method, the embodiments of the present disclosure further provide a model training device 900, which may be integrated in an electronic device. Hereinafter, with reference to Figure 9 the following structural schematic diagram, a model training device 900 provided by the embodiments of the present disclosure will be described.

[0153] The model training device 900 includes:

[0154] A data acquisition unit 901, configured to acquire a training data set, where the training data set includes a plurality of image samples, as well as a module selection label and an image property label corresponding to each image sample. The module selection label is used to indicate the target detection module corresponding to the image sample, and the image property label is used to indicate whether the image sample is a real face image;

[0155] A network training unit 902 is configured to train an initial detection network using a training data set to obtain an image detection network. When the initial detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel.

[0156] In specific implementation, each of the above modules can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules, reference can be made to the embodiments of the foregoing model training method, which will not be elaborated here.

[0157] By using the model training device 900 provided in the embodiments of the present disclosure, a training data set including multiple image samples, as well as a module selection label and an image property label corresponding to each image sample can be obtained, and then the initial detection network is trained using the training data set to obtain an image detection network. When the initial detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel. In this way, based on the image detection network, an output feature sequence of the image to be detected can be obtained only through one of the live detection module and the synthetic image detection module, and then based on the target feature sequence, a detection result of the image to be detected can be obtained, rather than detecting the image to be detected through the live detection module and the synthetic image detection module successively as in the prior art. Therefore, the complexity of the image detection process can be reduced, thereby improving the image detection efficiency.

[0158] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0159] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a storage medium, and a computer program product.

[0160] Figure 10 A schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown.

[0161] As described above, in the embodiments of the present disclosure, the electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0162] As Figure 10 shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0163] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as, a keyboard, a mouse, etc.; an output unit 1007, such as, various types of displays, speakers, etc.; a storage unit 1008, such as, a disk, an optical disc, etc.; and a communication unit 1009, such as, a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0164] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, for example, the image detection method and / or the model training method. For example, in some alternative embodiments, the image detection method and / or the model training method can be respectively implemented as a computer software program, which is tangibly included in a non-transitory computer-readable storage medium, such as the storage unit 1008. In some alternative embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the image detection method and / or the model training method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the image detection method and / or the model training method by any other suitable means (such as by means of firmware).

[0165] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0166] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0167] In the context of this disclosure, a non-transitory computer-readable storage medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A non-transitory computer-readable storage medium can be a machine-readable signal medium or a machine-readable storage medium. A non-transitory computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a non-transitory computer-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0168] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0169] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0170] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0171] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the foregoing image detection method.

[0172] Embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements the foregoing image detection method and / or model training method.

[0173] By using the electronic device, storage medium, and computer program product provided by the embodiments of the present disclosure, the output feature sequence of the image to be detected can be obtained only through one of the live detection module and the synthetic image detection module, and then the detection result of the image to be detected can be obtained based on the target feature sequence, rather than detecting the image to be detected through the live detection module and the synthetic image detection module in sequence as in the prior art. Therefore, the complexity of the image detection process can be reduced, thereby improving the image detection efficiency.

[0174] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, in the present disclosure, "a plurality" can be understood as at least two.

[0175] The foregoing specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An image detection method, comprising: Inputting the image to be detected into an image detection network. When the image detection network includes at least one image detection model and includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model; Sequentially in each of the image detection models, determining a target detection module corresponding to the image to be detected from a live detection module and a synthetic image detection module; Obtaining an output feature sequence of the image to be detected through the target detection module; Obtaining a detection result of the image to be detected based on a target feature sequence, where the target feature sequence is the output feature sequence obtained by the last image detection model in the image detection network, and the detection result is used to indicate whether the image to be detected is a real face image.

2. The method according to claim 1, wherein, The sequentially determining, in each of the image detection models, a target detection module corresponding to the image to be detected from a live detection module and a synthetic image detection module includes: For the current detection model in the at least one image detection model, obtaining an input feature sequence corresponding to the current detection model; Determining a detection category of the image to be detected based on the input feature sequence through the current detection model; Determining the target detection module from the live detection module and the synthetic image detection module included in the current detection model according to the detection category.

3. The method according to claim 2, wherein, The obtaining an input feature sequence corresponding to the current detection model includes: In the case where the current detection model is the first image detection model in the image detection network, splitting the image to be detected to obtain at least two image feature blocks, and forming an input feature sequence corresponding to the current detection model through the at least two image feature blocks and a category feature block; In the case where the current detection model is a non-first image detection model in the image detection network, obtaining the output feature sequence obtained by the previous image detection model of the current detection model as the input feature sequence corresponding to the current detection model.

4. The method according to claim 3, wherein, The determining a detection category of the image to be detected based on the input feature sequence includes: Performing feature extraction on the input feature sequence to obtain an intermediate feature sequence; Determining a first target feature block corresponding to the category feature block from the intermediate feature sequence; Performing binary classification processing on the first target feature block to obtain a first classification result; Performing normalization processing on the first classification result to determine the detection category of the image to be detected.

5. The method according to claim 4, wherein The performing feature extraction on the input feature sequence to obtain an intermediate feature sequence includes: Performing feature extraction on the input feature sequence through a multi-head attention module based on the attention relationship between each feature block in the input feature sequence to obtain the intermediate feature sequence.

6. The method according to claim 4, wherein, The obtaining an output feature sequence of the image to be detected through the target detection module includes: Through the target detection module, based on the correlation between each feature block in the intermediate feature sequence and the detection category, feature extraction is performed on the intermediate feature sequence to obtain the output feature sequence.

7. The method according to claim 3, wherein Obtaining the detection result of the image to be detected based on the target feature sequence includes: Determining a second target feature block corresponding to the category feature block from the target feature sequence; Performing binary classification processing on the second target feature block to obtain a second classification result; Performing normalization processing on the second classification result to obtain the detection result of the image to be detected.

8. A model training method, including: Obtaining a training data set, the training data set including a plurality of image samples, as well as a module selection label and an image property label corresponding to each image sample, the module selection label being used to indicate the target detection module corresponding to the image sample, and the image property label being used to indicate whether the image sample is a real face image; Training an initial detection network through the training data set to obtain an image detection network. When the initial detection network includes at least one image detection model and when the initial detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel; wherein, the image detection network is used to implement the image detection method according to any one of claims 1 to 7.

9. An image detection device, including: An image input unit for inputting an image to be detected into an image detection network. When the image detection network includes at least one image detection model and when the image detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models is used as the input of the next image detection model; A module selection unit for sequentially determining, in each image detection model, a target detection module corresponding to the image to be detected from the live detection module and the synthetic image detection module; A feature output unit for obtaining an output feature sequence of the image to be detected through the target detection module; An image detection unit for obtaining the detection result of the image to be detected based on the target feature sequence, the target feature sequence being the output feature sequence obtained by the last image detection model in the image detection network, and the detection result being used to indicate whether the image to be detected is a real face image.

10. The apparatus according to claim 9, wherein, The module selection unit is used for: For the current detection model among the at least one image detection model, obtaining an input feature sequence corresponding to the current detection model; Through the current detection model, based on the input feature sequence, determining the detection category of the image to be detected; According to the detection category, determining the target detection module from the live detection module and the synthetic image detection module included in the current detection model.

11. The apparatus according to claim 10, wherein, The module selection unit is used for: When the current detection model is the first image detection model in the image detection network, the image to be detected is segmented to obtain at least two image feature blocks, and an input feature sequence corresponding to the current detection model is formed by the at least two image feature blocks and the category feature block; When the current detection model is a non-first image detection model in the image detection network, the output feature sequence obtained by the previous image detection model of the current detection model is acquired as the input feature sequence corresponding to the current detection model.

12. The apparatus according to claim 11, wherein, The module selection unit is configured to: Extract features from the input feature sequence to obtain an intermediate feature sequence; Determine a first target feature block corresponding to the category feature block from the intermediate feature sequence; Perform binary classification processing on the first target feature block to obtain a first classification result; Perform normalization processing on the first classification result to determine the detection category of the image to be detected.

13. The apparatus according to claim 12, wherein, The module selection unit is configured to: Extract features from the input feature sequence through a multi-head attention module based on the attention relationship between the feature blocks in the input feature sequence to obtain the intermediate feature sequence.

14. The apparatus according to claim 12, wherein The feature output unit is configured to: Extract features from the intermediate feature sequence through the target detection module based on the correlation between the feature blocks in the intermediate feature sequence and the detection category to obtain the output feature sequence.

15. The device according to claim 11, wherein The image detection unit is configured to: Determine a second target feature block corresponding to the category feature block from the target feature sequence; Perform binary classification processing on the second target feature block to obtain a second classification result; Perform normalization processing on the second classification result to obtain the detection result of the image to be detected.

16. A model training device, comprising: A data acquisition unit, configured to acquire a training data set, the training data set including a plurality of image samples, as well as a module selection label and an image property label corresponding to each image sample, the module selection label being used to indicate the target detection module corresponding to the image sample, and the image property label being used to indicate whether the image sample is a real face image; A network training unit, configured to train an initial detection network through the training data set to obtain an image detection network. When the initial detection network includes at least one image detection model and when the initial detection network includes at least two image detection models, the output of the previous image detection model among two adjacent image detection models serves as the input of the next image detection model. Each image detection model includes a live detection module and a synthetic image detection module arranged in parallel; wherein, the image detection network is used to implement the image detection method according to any one of claims 1 to 7.

17. An electronic device, comprising: At least one processor; A memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

19. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Face living body detection method and device, electronic equipment and storage medium

    CN114511908A