Image Detection Method, Device, Computer Equipment and Storage Medium

Through the multi-task image detection model, the problem of vulnerability to forgery and fraudulent attacks in the prior art is solved, and more accurate and safer vital detection results are achieved.

CN112308035BActive Publication Date: 2025-08-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011337405.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-25
Publication Date
2025-08-01
Estimated Expiration
2040-11-25

AI Technical Summary

Technical Problem

The existing live detection technology is susceptible to counterfeiting and deceptive methods, resulting in inaccurate detection results and poses security risks.

Method used

A multi-task image detection model is used to extract different types of image features through feature extraction networks, and a branch network is used to perform live detection and image consistency verification, and the final detection results are determined based on the live detection results and image consistency verification results.

Benefits of technology

Improve the security and accuracy of image detection results, effectively defend against forgery and spoofing attacks, and ensure the reliability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112308035B_ABST
    Figure CN112308035B_ABST
Patent Text Reader

Abstract

This application relates to an image detection method, apparatus, computer device, and storage medium. The method includes: obtaining an image to be detected, where the image to be detected includes at least two different types of images; inputting the image to be detected into an image detection model, and the image detection model extracts image features corresponding to at least two different types of images respectively through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to at least two different types of images to obtain image consistency verification results corresponding to at least two different types of images; determining an image detection result corresponding to the image to be detected based on the liveness detection results and the image consistency verification results. Using this method can improve the security and accuracy of the image detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an image detection method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of image detection technology, a live detection technology has emerged. Live detection is a method for determining the true physiological characteristics of an object in some identity verification scenarios. Currently, when performing live detection, it is usually by collecting RGB images and performing live detection through the RGB images. However, the method of using RGB images for live detection is vulnerable to attacks such as forgery and deception, such as photos and 3D masks, resulting in inaccurate live detection results and causing security risks. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide an image detection method, apparatus, computer device, and storage medium that can improve the accuracy of detection results.

[0004] An image detection method, the method comprising:

[0005] Obtaining an image to be detected, the image to be detected including at least two different types of images;

[0006] Inputting the image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images respectively through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for live detection to obtain live detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to at least two different types of images to obtain image consistency verification results corresponding to at least two different types of images;

[0007] Determining an image detection result corresponding to the image to be detected based on the live detection result and the image consistency verification result.

[0008] An image detection apparatus, the apparatus comprising:

[0009] An image acquisition module, configured to obtain an image to be detected, the image to be detected including at least two different types of images;

[0010] An image detection module, configured to input an image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to at least two different types of images to obtain image consistency verification results corresponding to at least two different types of images;

[0011] A result determination module, configured to determine the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result.

[0012] A computer device, comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0013] Obtain an image to be detected, where the image to be detected includes at least two different types of images;

[0014] Input the image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to at least two different types of images to obtain image consistency verification results corresponding to at least two different types of images;

[0015] Determine the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result.

[0016] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0017] Obtain an image to be detected, where the image to be detected includes at least two different types of images;

[0018] Input the image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to at least two different types of images to obtain image consistency verification results corresponding to at least two different types of images;

[0019] Determine the image detection result corresponding to the image to be detected based on the live detection result and the image consistency verification result.

[0020] The above image detection method, device, computer device, and storage medium obtain an image to be detected, where the image to be detected includes at least two different types of images; input the image to be detected into an image detection model, and the image detection model extracts the image features corresponding to at least two different types of images respectively through a feature extraction network, input the image features corresponding to at least two different types of images into the corresponding branch network for live detection to obtain the live detection results corresponding to at least two different types of images, and perform image consistency verification on the basis of the image features corresponding to at least two different types of images using the corresponding branch network to obtain the image consistency verification results corresponding to at least two different types of images; determine the image detection result corresponding to the image to be detected based on the live detection result and the image consistency verification result. By performing live detection on different types of images through different branch networks and performing image consistency verification on different types of images using the branch network, the data consistency of different types of images is ensured, and attacks such as photos can be effectively defended, thereby improving the security and accuracy of the image detection result. Description of the Drawings

[0021] Figure 1 It is an application environment diagram of the image detection method in an embodiment;

[0022] Figure 2 It is a flowchart of the image detection method in an embodiment;

[0023] Figure 3 It is a flowchart of obtaining the first image consistency verification result in an embodiment;

[0024] Figure 4 It is a flowchart of determining the image detection result in an embodiment;

[0025] Figure 5 It is a flowchart of obtaining the target image consistency verification result in an embodiment;

[0026] Figure 6 It is a flowchart of obtaining the image consistency verification result in an embodiment;

[0027] Figure 7 It is a flowchart of training the image detection model in an embodiment;

[0028] Figure 8 It is a flowchart of obtaining the updated image detection model in an embodiment;

[0029] Figure 9 It is a flowchart of the image detection method in a specific embodiment;

[0030] Figure 10 Schematic diagram of mirror attack in a specific embodiment;

[0031] Figure 11 is Figure 10 Schematic diagram of the structure of the image detection model in a specific embodiment;

[0032] Figure 12 Block diagram of the structure of the image detection device in an embodiment;

[0033] Figure 13 Internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0034] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0035] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses cameras and computers to replace human eyes to identify and measure targets, and further performs graphic processing to make the computer process images that are more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0036] The solution provided in the embodiment of the present application relates to technologies such as biometric recognition in artificial intelligence, and is specifically described through the following embodiments:

[0037] The image detection method provided by the present application can be applied to such as Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The server 102 obtains the image to be detected collected by the terminal 102, and the image to be detected includes at least two different types of images; the server 102 inputs the image to be detected into the image detection model, and the image detection model extracts the image features corresponding to at least two different types of images respectively through the feature extraction network, inputs the image features corresponding to at least two different types of images into the corresponding branch network for liveness detection, obtains the liveness detection results corresponding to at least two different types of images, and performs image consistency verification on the basis of the image features corresponding to at least two different types of images respectively using the corresponding branch network, obtains the image consistency verification results corresponding to at least two different types of images; the server 102 determines the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0038] In one embodiment, as Figure 2 shown, a method for image detection is provided. Taking the server in Figure 1 as an example for description, it can be understood that this method can also be applied to the terminal. In this embodiment, the method includes the following steps:

[0039] Step 204, obtain the image to be detected, where the image to be detected includes at least two different types of images.

[0040] Among them, the image to be detected refers to the image that needs to be subjected to liveness detection, and this image can be an image collected by different types of cameras. The liveness refers to an object with real physiological characteristics, including people, animals, plants, etc. For example, in the application of face recognition, liveness detection can use images of combined actions such as blinking, opening the mouth, shaking the head, and nodding, and use technologies such as face key point positioning to verify whether the user is the real liveness person operating.

[0041] Different types of images refer to images with different image information. For example, infrared type images with specific infrared information can be collected by an infrared camera, RGB type images with RGB color information can be collected by an RGB (a color standard) camera, depth type images with specific depth information can be collected by a depth camera, and other types of images can be collected by other types of cameras.

[0042] Specifically, the server can obtain the images to be detected uploaded by the terminal at the same time. The images to be detected include at least two different types of images. For example, when the terminal captures different types of images through different types of cameras at the same time, the captured different types of images are uploaded to the server. The server obtains the images to be detected uploaded by the terminal, which may include infrared type images, RGB type images, and depth type images. It may also include infrared type images and RGB type images, or RGB type images and depth type images, or infrared type images and depth type images, etc. In one embodiment, the server can also obtain the images to be detected captured at the same time point and saved in the database. The images to be detected include at least two different modality images, that is, the server can obtain multiple modality images as the images to be detected.

[0043] Step 206: Input the images to be detected into the image detection model. The image detection model extracts the image features corresponding to at least two different types of images respectively through the feature extraction network, inputs the image features corresponding to at least two different types of images into the corresponding branch network for live detection, obtains the live detection results corresponding to at least two different types of images, and performs image consistency verification on the at least two different types of images respectively based on the image features corresponding to them using the corresponding branch network, and obtains the image consistency verification results corresponding to at least two different types of images.

[0044] Among them, the image detection model refers to a multi-task image detection model, which is used to perform live detection and image consistency verification on the images to be detected. The image detection model can be an artificial intelligence model obtained by performing multi-task training on the training sample data based on the neural network algorithm. The training sample data can include positive sample data and negative sample data. The positive sample data includes images with real live labels and images with image consistent labels, and the negative sample data includes images with non-real live labels and images with image inconsistent labels. Among them, the neural network in the neural network algorithm includes an input layer, a convolutional layer, a normalization (BN) layer, a fully connected layer, and an output layer. Among them, the activation function can use an S-type activation function or an R-type activation function, etc., and the loss function uses a classification task loss function, which can include a cross-entropy loss function, etc. The neural network algorithm can include an RNN feedforward neural network algorithm, a CNN convolutional neural network algorithm, a Long Short-Term Memory (LSTM), etc.

[0045] The feature extraction network refers to the network in the image detection model used to extract image features. This feature extraction network can extract image features of different types of images, that is, the network parameters in this feature extraction network are shared. Different model tasks use different branch networks for task processing. The branch network includes at least two branch networks for liveness detection and at least one branch network for image consistency verification. Different types of images correspond to different branch networks for liveness detection. For example, RGB type images correspond to the RGB liveness detection branch network, depth type images correspond to the depth liveness detection branch network, and infrared type images correspond to the infrared liveness detection branch network. Any two different types of images correspond to a branch network for image consistency verification. For example, RGB type images and depth type images correspond to the RGB and depth image consistency verification branch network, RGB type images and infrared type images correspond to the RGB and infrared image consistency verification branch network, and depth type images and infrared type images correspond to the depth and infrared image consistency verification branch network.

[0046] Specifically, the server inputs the image to be detected into the image detection model for multi-task processing. The image detection model inputs at least two different types of images into the feature extraction network for feature extraction simultaneously, and obtains the image features corresponding to at least two different types of images respectively. Each type of image has a corresponding image feature. Then, the image features corresponding to at least two different types of images are respectively input into the corresponding liveness detection branch network for liveness detection, and the liveness detection results corresponding to at least two different types of images are obtained. Each type of image has a corresponding liveness detection result. Then, the image features corresponding to at least two different types of images are input into the corresponding image consistency verification branch network for image consistency verification, and the image consistency verification results corresponding to at least two different types of images are obtained. Each pair of types of images has a corresponding image consistency verification result. The liveness detection result is used to characterize whether the image to be detected is a live image. The image consistency verification result is used to verify the consistency of different types of images. The liveness detection result includes that the image to be detected is a live image and a non-live image, and the image consistency verification result includes that the image consistency verification passes and the image consistency verification fails.

[0047] Step 208, determine the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result.

[0048] Specifically, when the live detection results corresponding to all different types of images are live images and the image consistency verification results corresponding to all different types of images all pass the image consistency verification, the image detection result corresponding to the image to be detected is obtained as a live image. When there is a non-live image in the body detection result or the image consistency verification result fails the image consistency verification, the image detection result corresponding to the image to be detected is obtained as a non-live image.

[0049] In the above image detection method, by obtaining the image to be detected, the image to be detected includes at least two different types of images; inputting the image to be detected into the image detection model, the image detection model extracts the image features corresponding to at least two different types of images respectively through the feature extraction network, inputs the image features corresponding to at least two different types of images into the corresponding branch network for live detection, obtains the live detection results corresponding to at least two different types of images, and performs image consistency verification using the corresponding branch network based on the image features corresponding to at least two different types of images respectively, obtains the image consistency verification results corresponding to at least two different types of images; determines the image detection result corresponding to the image to be detected based on the live detection result and the image consistency verification result. By performing live detection on different types of images through different branch networks, and performing image consistency verification on different types of images using the branch network, the data consistency of different types of images is ensured, and attacks such as forgery and deception can be effectively defended, thereby improving the security and accuracy of the image detection result.

[0050] In one embodiment, as Figure 3 shown, step 204, inputting the image features corresponding to at least two different types of images into the corresponding branch network for live detection, obtaining the live detection results corresponding to at least two different types of images, and performing image consistency verification using the corresponding branch network based on the image features corresponding to at least two different types of images respectively, obtaining the image consistency verification results corresponding to at least two different types of images, includes:

[0051] Step 302, inputting the image features corresponding to the first type of image among at least two different types of images into the corresponding first detection branch network for live detection, obtaining the first live detection result corresponding to the first type of image.

[0052] Among them, the first type of image refers to an RGB type image with RGB information. The first detection branch network is used for live detection of RGB type images. The first live detection result includes that the RGB type image is a live image and a non-live image.

[0053] Specifically, the server inputs the image features corresponding to the first type of image into the first detection branch network for live detection, and obtains the output first live detection result.

[0054] Step 304: Input the image features corresponding to the second type of image in at least two different types of images into the corresponding second detection branch network for liveness detection, and obtain the second liveness detection result corresponding to the second type of image.

[0055] Among them, the second type of image refers to a depth type image with depth information. The second detection branch network is used to perform liveness detection on the depth type image. The second liveness detection result includes that the depth type image is a live image and a non-live image.

[0056] Specifically, the server inputs the image features corresponding to the second type of image into the second detection branch network for liveness detection, and obtains the output second liveness detection result.

[0057] Step 306: Input the image features corresponding to the first type of image and the image features corresponding to the second type of image into the first verification branch network for image consistency verification, and obtain the first image consistency verification result corresponding to the first type of image and the second type of image.

[0058] Among them, the first verification branch network is used to verify the consistency between the first type of image and the second type of image. The first image consistency verification result is used to represent whether the consistency verification between the first type of image and the second type of image passes, and may include that the consistency verification between the first type of image and the second type of image passes and the consistency verification between the first type of image and the second type of image fails.

[0059] Specifically, the server splices the image features corresponding to the first type of image and the image features corresponding to the second type of image to obtain the spliced image features, and inputs the spliced image features into the first verification branch network for image consistency verification. Obtain the first image consistency verification result of the book.

[0060] Step 206: Determine the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result, including the steps:

[0061] Determine the image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, and the first image consistency verification result.

[0062] Specifically, when the first liveness detection result is a live image, the second liveness detection result is a live body, and the first image consistency verification result is that the image consistency verification passes, the server obtains that the image detection result corresponding to the image to be detected is a live image.

[0063] In the above embodiments, through a multi-task model, i.e., an image detection model, different types of images are used for liveness detection using different branch networks, and image consistency verification is performed through the first verification branch network. Then, according to the liveness detection results and the image consistency verification results, the image detection results corresponding to the images to be detected are determined, improving the accuracy and efficiency of obtaining the image detection results.

[0064] In one embodiment, as Figure 4 shown, step 304, that is, after inputting the image features corresponding to the second type of image among at least two different types of images into the corresponding second detection branch network for liveness detection to obtain the second liveness detection result corresponding to the second type of image, further includes:

[0065] Step 402, inputting the image features corresponding to the third type of image among at least two different types of images into the corresponding third detection branch network for liveness detection to obtain the third liveness detection result corresponding to the third type of image.

[0066] Among them, the third type of image refers to an infrared type image with infrared information. The third detection branch network is used for liveness detection of infrared type images. The third liveness detection result includes that the infrared type image is a live image and a non-live image.

[0067] Specifically, the server image detection model also includes a third detection branch network. When using the feature extraction network for feature extraction, the image features corresponding to the third type of image are extracted simultaneously, and then the image features corresponding to the third type of image are input into the corresponding third detection branch network for liveness detection simultaneously to obtain the third liveness detection result output by the third detection branch network.

[0068] The image detection method further includes:

[0069] Step 404, inputting the image features corresponding to the first type of image and the image features corresponding to the third type of image into the second verification branch network for image consistency verification to obtain the second image consistency verification result corresponding to the first type of image and the third type of image.

[0070] Among them, the second verification branch network is used for verifying the consistency of the first type of image and the third type of image. The second image consistency verification result is used to represent whether the consistency verification of the first type of image and the third type of image passes, and may include that the consistency verification of the first type of image and the third type of image passes and the consistency verification of the first type of image and the third type of image fails.

[0071] Specifically, when the third type of image is the target type of image, which is used for image recognition after obtaining the image detection result corresponding to the image to be detected, the server splices the image features corresponding to the first type of image and the image features corresponding to the third type of image to obtain the spliced features, and inputs the spliced features into the second verification branch network for image consistency verification to obtain the second image consistency verification result corresponding to the first type of image and the third type of image.

[0072] Step 406: Determine the image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, and the second image consistency verification result.

[0073] Among them, the first liveness detection result refers to the detection result of whether the first type of image is a live image. The second liveness detection result refers to the detection result of whether the second type of image is a live image, and the third liveness detection result refers to the detection result of whether the third type of image is a live image.

[0074] Specifically, when the first type of image is a live image, the second type of image is a live image, the third type of image is a live image, the first image consistency verification result is verification passed, and the second image consistency verification result is verification passed, the image detection result corresponding to the image to be detected is a live image. When any of the types of images is a non-live image or the verification result is verification failed, the image detection result corresponding to the image to be detected is a non-live image.

[0075] In one embodiment, the server inputs the image features corresponding to the second type of image and the image features corresponding to the third type of image into the third verification branch network for image consistency verification to obtain the third image consistency verification result corresponding to the second type of image and the third type of image. Then, based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, the second image consistency verification result, and the third image consistency verification result, determine the image detection result corresponding to the image to be detected.

[0076] In the above embodiment, when the image to be detected includes the third type of image, the third type of image is used for liveness detection and image consistency verification. When the liveness detection and image consistency verification of all branch networks pass, the detection result that the image to be detected is a live image is obtained, making the obtained image detection result more accurate.

[0077] In one embodiment, step 202, that is, obtaining the image to be detected, includes the steps of:

[0078] Obtain an initial video, perform image quality detection on each initial image in the initial video to obtain each initial image quality information; based on each initial image quality information, perform image screening from each initial image quality to obtain an image to be detected.

[0079] Among them, the initial video refers to the video for which video images need to be detected. Image quality refers to the subjective evaluation of people's visual perception of an image. It is generally considered that image quality refers to the degree of error generated by the measured image (i.e., the target image) relative to the standard image (i.e., the original image) in the human visual system. Image quality can be further divided into image fidelity and image intelligibility. Image fidelity describes the deviation degree between the processed image and the original image; while image intelligibility represents the degree to which people or machines can extract relevant feature information from the image. The initial image refers to the image for which image quality detection needs to be performed. The initial image quality information is used to characterize the image quality corresponding to the image in the initial video.

[0080] Specifically, the server obtains the initial video, which can be that the terminal captures the video in real time through the camera and uploads it to the server. It can also be that the server obtains the video pre - saved in the database. It can also be that the server obtains the initial video from the Internet. Then the server can perform image quality detection on each initial image in the initial video to obtain each initial image quality information. For example, the preset image quality evaluation index can be used to detect each initial image in the initial video. The preset image quality evaluation index can include resolution, color depth, and image distortion, etc. For example, a pre - trained image quality evaluation model can also be used to perform image quality detection on each initial image in the initial video, and this image quality evaluation model can be a model trained using the convolutional neural network algorithm. Finally, the server selects the best initial image quality information from each initial image quality information, and takes the initial image corresponding to the best initial image quality information as the image to be detected.

[0081] In the above - mentioned embodiment, by detecting the image quality of the initial image and then screening the initial image, the image to be detected is obtained, so that when detecting the image subsequently, a more accurate detection result can be obtained.

[0082] In one embodiment, the image detection model includes at least two feature extraction networks; as Figure 5 shown in step 204, that is, extracting the image features corresponding to at least two different types of images through the feature extraction network, including:

[0083] Step 502, inputting at least two different types of images into the corresponding feature extraction networks respectively for feature extraction to obtain the target image features corresponding to at least two different types of images; among them, the loss value between the network parameters in at least two feature extraction networks is less than the preset threshold.

[0084] Among them, at least two feature extraction networks mean that each different type of image corresponds to a feature extraction network, that is, one type of image corresponds to one feature extraction network. The target image feature refers to the image feature output after different types of images are input into the corresponding feature extraction network. The preset threshold is a pre-set loss value threshold.

[0085] Specifically, the server inputs at least two different types of images into the corresponding feature extraction networks respectively for feature extraction, obtains the target image features corresponding to at least two different types of images, and the loss value between the network parameters in the at least two feature extraction networks is less than the preset threshold. That is, during training, a trained feature extraction network can be obtained only when the loss value between the network parameters in the feature extraction network is less than the preset threshold.

[0086] Step 204, that is, input the image features corresponding to at least two different types of images into the corresponding branch networks for liveness detection, obtain the liveness detection results corresponding to at least two different types of images, and perform image consistency verification on the image features corresponding to at least two different types of images using the corresponding branch networks to obtain the image consistency verification results corresponding to at least two different types of images, including:

[0087] Step 504, input the target image features corresponding to at least two different types of images into the corresponding branch networks for liveness detection, and obtain the target liveness detection results corresponding to at least two different types of images.

[0088] Among them, the target liveness detection result refers to the liveness detection result obtained according to the target image feature, and the target liveness detection result includes non-live images and live images.

[0089] Specifically, the server simultaneously inputs the target image features corresponding to at least two different types of images into the corresponding branch networks for liveness detection, and obtains the output of the branch networks, that is, the target liveness detection results corresponding to at least two different types of images.

[0090] Step 506, perform image consistency verification on the target image features corresponding to at least two different types of images using the corresponding branch networks, and obtain the target image consistency verification results corresponding to at least two different types of images. [[ID=2,1]]

[0091] Specifically, the server splices the target image features corresponding to at least two different types of images in pairs, that is, the server splices the target image features corresponding to any two different types of images, and then inputs the spliced image features into the corresponding branch networks for image consistency verification, and obtains the target image consistency verification results output by each branch network. The target image consistency verification result is used to characterize the consistency of two different types of images.

[0092] In the above embodiments, when the image detection model includes at least two feature extraction networks, image features of different types of images are extracted through different feature extraction networks, so that the extracted image features are more accurate, and further the image consistency verification result of the subsequent live detection result is more accurate.

[0093] In one embodiment, as Figure 6 shown, step 204, performing image consistency verification on the image features corresponding to at least two different types of images using corresponding branch networks to obtain image consistency verification results corresponding to at least two different types of images, including:

[0094] Step 602, calculating the feature similarity between the image features corresponding to the first type of image in at least two different types of images and the image features corresponding to the second type of image in at least two different types of images through the first detection branch network.

[0095] Wherein, the feature similarity is used to characterize the similarity degree between different types of images, and the higher the feature similarity, the more similar the different types of images are.

[0096] Specifically, the server calculates the feature similarity between the image features corresponding to the first type of image in at least two different types of images and the image features corresponding to the second type of image in at least two different types of images through the first detection branch network. In one embodiment, the server can also directly use a similarity algorithm to calculate the feature similarity between the image features corresponding to the first type of image and the image features corresponding to the second type of image, and the similarity algorithm can include a distance similarity algorithm, a cosine similarity algorithm, a Pearson correlation coefficient algorithm, and so on.

[0097] Step 604, when the feature similarity exceeds a preset image consistency threshold, obtaining that the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification passes, and the image consistency verification passing is used to indicate that different types of images are images collected by the same device at the same time.

[0098] Specifically, the preset image consistency threshold refers to the threshold preset for indicating image consistency. When collecting images of the same living object at the same time using different types of cameras on the same device, the image features at each position of the obtained different types of images are basically aligned. There are only modal differences between different types of images, while the image features of attack images using forgery and deception methods such as photos and 3D masks are significantly different from those of other types of images. Based on this, the server compares the feature similarity with the preset image consistency threshold. When the feature similarity exceeds the preset image consistency threshold, it indicates that the first type of image and the second type of image are images of the same living object collected at the same time using the same device. Therefore, the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification passes. When the feature similarity does not exceed the preset image consistency threshold, it indicates that the first type of image and the second type of image are not images of the same living object collected at the same time using the same device, and the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification fails.

[0099] In the above embodiment, the feature similarity between the image features of the first type of image and the second type of image is calculated through the branch network, and then the image consistency verification result corresponding to the first type of image and the second type of image is determined according to the feature similarity, which improves the accuracy of the image consistency verification result and effectively ensures the consistency of different types of images, so as to effectively defend against attack images and avoid security problems.

[0100] In one embodiment, the image detection model includes an image conversion network; before step 204, that is, before extracting the image features corresponding to at least two different types of images through the feature extraction network, inputting the image features corresponding to at least two different types of images into the corresponding branch network for live detection to obtain the live detection results corresponding to at least two different types of images, and performing image consistency verification on the basis of the image features corresponding to at least two different types of images using the corresponding branch network to obtain the image consistency verification results corresponding to at least two different types of images, the method further includes the step of:

[0101] The image detection model inputs the second type of image among at least two different types of images into the image conversion network for image conversion to obtain a converted image with the same type as the first type of image among at least two different types of images, and uses the converted image as the second type of image.

[0102] Among them, the image conversion network is used to convert images of different types into images of the same type. For example, it can convert RGB images into depth images or convert RGB images into infrared images. The converted image refers to an image that is obtained after converting the second type of image and has the same type as the first type of image.

[0103] Specifically, the image detection model in the server may also include an image conversion network. Then, the second type of image that needs to be converted is input into the image conversion network for image conversion to obtain the converted image that is output and has the same type as the first type of image. This converted image is used as the second type of image, and then the first type of image and the second type of image are used to perform liveness detection through the image detection model.

[0104] In one embodiment, after obtaining the image to be detected, it is also possible to directly convert the image that needs to be converted, and then input the converted image into the image detection model for image liveness detection.

[0105] In one embodiment, before performing image conversion, it is also possible to preprocess the image to be detected to obtain the preprocessed image, and use the preprocessed image for image detection. For example, all different types of images in the image to be detected are scaled to the same size, and the different types of images of the same size are used for image detection.

[0106] In the above embodiments, by converting different types of images into images of the same type and then using the images of the same type for image liveness detection, the efficiency of image liveness detection is improved, and it is more convenient to perform liveness detection on images.

[0107] In one embodiment, step 206, that is, determining the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result, includes the steps:

[0108] When at least two different types of images are liveness images and the image consistency verification result is that the image consistency verification passes, the image detection result is that the liveness detection passes.

[0109] Among them, a liveness image refers to an object in the image to be detected that is a live object.

[0110] Specifically, the server determines that when the live detection results of different types of images are all live images, and all image consistency verification results pass the image consistency verification, the image detection result is obtained as a live detection pass, that is, the image to be detected is a live image. The server determines that when there is a non-live image in the live detection results corresponding to different types of images or there is an image consistency verification failure in the image consistency verification results, the image detection result is obtained as a live detection failure, that is, there are spoofed or forged type images in the image to be detected. At this time, the server can generate information indicating that the live image detection fails for prompting.

[0111] In one embodiment, when the image detection result is obtained as a live detection pass, the RGB image in the input of the image to be detected is input into an image recognition model for image recognition. Among them, the image recognition model can be a face recognition model, which is used to recognize the identity information of the face in the image, and then based on the recognized face identity information, face payment verification, face unlocking verification, etc. can be performed.

[0112] In the above embodiment, when at least two different types of images are live images and the image consistency verification result is a pass of the image consistency verification, the server obtains the image detection result as a live detection pass. Only when the live detection results of all branch networks are live images and the image consistency verification passes can the detection result of live detection pass be obtained, so that the obtained live detection result is updated accurately.

[0113] In one embodiment, as Figure 7 shown, the training of the image detection model includes the following steps:

[0114] Step 702, obtain a training data set, where the training data set includes at least two different types of sample images and corresponding class labels.

[0115] Among them, the training data set refers to the image data used to train the image detection model. The class labels include live detection class labels and image consistency detection labels. The live detection class labels are used to represent the live detection classes corresponding to the images, including live image classes and non-live image classes. The image consistency labels are used to represent the classes corresponding to the image consistency verification, including the class of passing the image consistency verification and the class of failing the image consistency verification.

[0116] Specifically, the server can obtain a training dataset from a database, which includes at least two different types of sample images and corresponding class labels. For example, the training dataset includes two types of samples and corresponding class labels, and a set of training samples can be stored in the database in the storage format of (the first type of image, the first type of image liveness detection label, the second type of image, the second type of image liveness detection label, the image consistency label corresponding to the first type of image and the second type of image). The server directly retrieves the training samples from the database. The server can also obtain the training dataset from the service provider server that provides the training data. The server can also collect the training dataset in real time through a collection device.

[0117] In a specific embodiment, the server can use the depth-type image, infrared-type image, and RGB-type image obtained by the same device at the same moment facing the same live body to obtain positive sample training data. Then, the RGB-type image, depth-type image, and infrared-type image corresponding to the mirror attack image are obtained to obtain negative sample training data, and then the training dataset is obtained according to the positive sample training data and the negative sample training data.

[0118] In an embodiment, the server can only obtain positive sample training data, that is, obtain the depth-type image, infrared-type image, and RGB-type image obtained by the same device at the same moment facing the same live body. Then, the positive sample training data is used to train the target image detection model, which is used to detect live body category images, and other non-live body category images are uniformly recognized as abnormal images.

[0119] Step 704, input at least two different types of sample images into the initialized image detection model. The initialized image detection model extracts the sample image features corresponding to at least two different types of sample images through the initialized feature extraction network, inputs the sample image features corresponding to at least two different types of sample images into the corresponding initialized branch network for liveness detection, obtains the initial liveness detection results corresponding to at least two different types of sample images, and performs image consistency verification using the corresponding initialized branch network based on the sample image features corresponding to at least two different types of sample images, and obtains the initial image consistency verification results corresponding to at least two different types of sample images.

[0120] Among them, the initialized image detection model refers to an image detection model with all model parameters initialized. The initialized feature extraction network refers to the initialization of the network parameters in the feature extraction network. The initialized branch network refers to the initialization of the network parameters in the branch network. The initial liveness detection result refers to the liveness detection result obtained based on the initialized parameters. The initial image consistency verification result refers to the image consistency verification result obtained based on the initialized parameters.

[0121] Specifically, when training an image detection model, the server performs forward propagation, that is, inputs at least two different types of sample images into the initialized image detection model. The initialized image detection model extracts sample image features corresponding to at least two different types of sample images through the initialized feature extraction network. Different types of sample images have different sample image features. Input the sample image features corresponding to at least two different types of sample images into the corresponding initialized branch network for liveness detection to obtain initial liveness detection results corresponding to at least two different types of sample images. Different types of sample images have different initial liveness detection results. And based on the sample image features corresponding to at least two different types of sample images, use the corresponding initialized branch network to perform image consistency verification to obtain initial image consistency verification results corresponding to at least two different types of sample images. That is, when the server trains a multi-task model, it performs forward propagation of multiple tasks simultaneously in the model.

[0122] Step 706: Calculate the loss based on the initial liveness detection results and the corresponding liveness detection class labels in the class labels to obtain liveness detection loss information corresponding to at least two different types of sample images, and calculate the loss based on the initial image consistency verification results and the corresponding image consistency class labels in the class labels to obtain image consistency verification loss information corresponding to at least two different types of sample images.

[0123] Among them, the liveness detection loss information is used to characterize the error between the initial liveness detection result and the corresponding liveness detection class label. The image consistency verification loss information is used to characterize the error between the initial image consistency verification result and the corresponding image consistency class label.

[0124] Specifically, the server uses a pre-set classification loss function to calculate the error between the initial liveness detection result and the corresponding liveness detection class label in the class labels to obtain liveness detection loss information corresponding to at least two different types of sample images. At the same time, use a pre-set classification loss function to calculate the error information between the initial image consistency verification result and the corresponding image consistency class label to obtain image consistency verification loss information. Among them, the classification loss function can be a cross-entropy loss function.

[0125] Step 708: Based on the liveness detection loss information corresponding to at least two different types of sample images, reversely update the corresponding initialized branch network and initialized feature extraction network, and reversely update the corresponding initialized branch network and initialized feature extraction network according to the image consistency verification loss information to obtain an updated image detection model.

[0126] Among them, the updated image detection model refers to the image detection model with updated parameters in the model.

[0127] Specifically, the server uses the liveness detection loss information to inversely update the corresponding initialized branch network and initialized feature extraction network based on the backpropagation algorithm, that is, to update each model parameter that affects the liveness detection result. And the server inversely updates the corresponding initialized branch network and initialized feature extraction network by using the image consistency verification loss information based on the backpropagation algorithm, that is, to update each model parameter that affects the image consistency verification result, so as to obtain an updated image detection model. Among them, the backpropagation algorithm can be a gradient descent algorithm, a conjugate gradient algorithm, and so on.

[0128] Step 710: Use the updated image detection model as the initial image detection model, and return to the step of inputting at least two different types of sample images into the initialized image detection model until the training is completed, and then obtain the image detection model.

[0129] Specifically, the server uses the updated image detection model as the initial image detection model, and returns to the step of inputting at least two different types of sample images into the initialized image detection model for execution until the training completion condition is reached. Then, the updated image detection model obtained in the last training is used as the image detection model after training. Among them, the training completion condition can be that the training reaches the maximum number of iterations, the model parameters do not change significantly, or the loss information reaches a preset threshold. Then the server can deploy and use the obtained image detection model.

[0130] In the above embodiment, the image detection model is obtained by pre-training with a training data set and then deployed and used, which improves the efficiency of image detection.

[0131] In one embodiment, the initialized image detection model includes at least two initialized feature extraction networks;

[0132] As Figure 8 shown, step 708, that is, inversely updating the corresponding initialized branch network and initialized feature extraction network based on the liveness detection loss information corresponding to at least two different types of sample images, and inversely updating the corresponding initialized branch network and initialized feature extraction network according to the image consistency verification loss information to obtain an updated image detection model, includes:

[0133] Step 802: Update the corresponding initialized feature extraction network and the corresponding initialized branch network based on the liveness detection loss information corresponding to at least two different types of sample images, and inversely update the corresponding initialized branch network and the corresponding initialized feature extraction network according to the image consistency verification loss information corresponding to at least two different types of sample images to obtain at least two to-be-confirmed feature extraction networks.

[0134] Among them, the to-be-confirmed feature extraction network refers to a feature extraction network whose network parameter update needs to be confirmed.

[0135] Specifically, the server updates the corresponding initialized feature extraction network and the corresponding initialized branch network using the liveness detection loss information. For example, if the liveness detection loss information is the liveness detection loss information corresponding to an RGB-type image, then the initialized feature extraction network for extracting this RGB-type image and the initialized branch network for performing liveness detection on the RGB-type image are updated using this liveness detection loss information. At the same time, the image consistency verification loss information is used to reversely update the corresponding initialized branch network and the corresponding initialized feature extraction network. For example, if the image consistency verification loss information is the image consistency verification loss information corresponding to an RGB-type image and a depth-type image, then at this time, the initialized branch network for verifying the consistency of the RGB-type image and the depth-type image, the initialized feature extraction network for extracting the RGB-type image, and the initialized feature extraction network for extracting the depth-type image are updated using this image consistency verification loss information. After the update is completed, at least two to-be-confirmed feature extraction networks are obtained.

[0136] Step 804: Calculate the parameter loss value between the network parameters in at least two to-be-confirmed feature extraction networks.

[0137] Specifically, the server calculates the parameter loss value between the network parameters in at least two to-be-confirmed feature extraction networks, that is, the network parameters in the to-be-confirmed feature extraction network are input into a pre-set loss function to calculate the error value between the network parameters. For example, the mean squared error loss function or the mean absolute error loss function can be used to calculate the error value between each network parameter in the first to-be-confirmed feature extraction network and each network parameter in the second to-be-confirmed feature extraction network.

[0138] Step 806: When the parameter loss value is less than the preset parameter loss threshold, obtain at least two updated feature extraction networks, and obtain an updated image detection model based on the at least two updated feature extraction networks.

[0139] Specifically, the preset parameter loss threshold refers to the error threshold between the pre-set network parameters. The server compares the parameter loss value with the preset parameter loss threshold. When the parameter loss value is less than the preset parameter loss threshold, at least two updated feature extraction networks are obtained, and an updated image detection model is obtained based on the at least two updated feature extraction networks. When the parameter loss value is not less than the preset parameter loss threshold, the network parameters are updated again or adjusted until the parameter loss value is less than the preset parameter loss threshold.

[0140] In the above embodiments, by calculating the parameter loss values between the network parameters in at least two feature extraction networks to be confirmed, at least two updated feature extraction networks are obtained, and then an image detection model is obtained based on the at least two updated feature extraction networks, so that the accuracy of the trained image detection model is improved.

[0141] In a specific embodiment, as Figure 9 shown, an image detection method is provided, which specifically includes the following steps:

[0142] Step 902: Obtain an initial video, perform image quality detection on each initial image in the initial video to obtain the image quality information of each initial image, and perform image screening based on the image quality information of each initial image to obtain the images to be detected. The images to be detected include at least two different types of images.

[0143] Step 904: Input the images to be detected into the image detection model. The image detection model inputs the second type of images among the at least two different types of images into the image conversion network for image conversion to obtain a converted image with the same type as the first type of images among the at least two different types of images, and uses the converted image as the second type of images.

[0144] Step 906: The image detection model extracts the image features corresponding to the at least two different types of images respectively through the feature extraction network, and inputs the image features corresponding to the first type of images among the at least two different types of images into the corresponding first detection branch network for live detection to obtain the first live detection result corresponding to the first type of images.

[0145] Step 908: The image detection model inputs the image features corresponding to the second type of images among the at least two different types of images into the corresponding second detection branch network for live detection to obtain the second live detection result corresponding to the second type of images.

[0146] Step 910: The image detection model inputs the image features corresponding to the third type of images among the at least two different types of images into the corresponding third detection branch network for live detection to obtain the third live detection result corresponding to the third type of images.

[0147] Step 912: The image detection model inputs the image features corresponding to the first type of images and the image features corresponding to the second type of images into the first verification branch network for image consistency verification to obtain the first image consistency verification result corresponding to the first type of images and the second type of images.

[0148] Step 914, the image detection model inputs the image features corresponding to the first type of image and the image features corresponding to the third type of image into the second verification branch network for image consistency verification, and obtains the second image consistency verification results corresponding to the first type of image and the third type of image.

[0149] Step 916, determine the image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, and the second image consistency verification result.

[0150] The present application also provides an application scenario, which applies the above image detection method. Specifically, the application of the image detection method in this application scenario is as follows:

[0151] In a face identity verification system, a depth (or infrared) camera and an RGB camera are used, and face liveness is judged using a depth (or infrared) map, and then face identity verification is performed according to the face liveness judgment result. At this time, if a mirror attack is used to attack the face identity verification system, the accuracy of face image liveness detection will be reduced, and security risks will be brought to subsequent face identity recognition, such as Figure 10 shown in the schematic diagram of using a mirror attack. Among them, the attacker uses a mirror to provide a forged photo to the RGB camera for shooting. Since face recognition uses RGB images for face recognition and passes the liveness detection of the forged RGB image, it will cause major security risks. However, in the present application, the collected depth face image, RGB face image, and infrared face image are input into the image detection model. As Figure 11The following is a schematic structural diagram of the image detection model in this application. It includes an input layer, a convolutional module, a fully connected layer, an output layer, and uses a loss function for backpropagation update during training. The convolutional module includes a convolutional layer, a BN layer, and an activation function. At this time, the image detection model extracts features through a feature extraction network, then inputs the deep face image features into the deep liveness detection network for liveness detection to obtain the deep liveness detection result. At the same time, the RGB face image features are input into the RGB liveness detection network for liveness detection to obtain the RGB liveness detection result. At the same time, the RGB face image features and the deep face image features are input into the image consistency verification network for cross-verification to obtain the image consistency verification result of the RGB face image and the deep face image. At the same time, the infrared face image features are input into the infrared liveness detection network for liveness detection to obtain the infrared detection result. At the same time, the infrared face image features and the RGB face image features are input into the image consistency verification network for cross-verification to obtain the image consistency verification result of the infrared face image and the RGB face image. Among them, the liveness detection network is a model based on a convolutional neural network. For example, it can use networks such as VGG16 (Visual Geometry Group Network 16), GoogleNet (a new deep learning structure), ResNet (residual network), and MobileNet (neural network applied to mobile devices), etc.

[0152] At this time, since the RGB image is an attack image, and the features of the attack image are quite different from those of the infrared image and the depth image, the image consistency verification result of the obtained infrared face image and the RGB face image fails, and the image consistency verification result of the obtained infrared face image and the RGB face image fails. However, the liveness detection result is a liveness detection failure, making the obtained image detection result more accurate, thus avoiding attacks by attackers on the face authentication system and ensuring the security of the face authentication system.

[0153] This application also provides an application scenario that applies the above image detection method. Specifically, the application of the image detection method in this application scenario is as follows:

[0154] Applied in the animal liveness image detection platform, specifically:

[0155] At the same time, infrared animal images, RGB animal images, and depth animal images are captured by different types of cameras. The infrared animal images, RGB animal images, and depth animal images are input into an animal image detection model. The animal image detection model extracts the image features corresponding to the infrared animal images, RGB animal images, and depth animal images through a feature extraction network, inputs the infrared animal images, RGB animal images, and depth animal images into corresponding branch networks for liveness detection to obtain the liveness detection results corresponding to the infrared animal images, RGB animal images, and depth animal images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to the infrared animal images, RGB animal images, and depth animal images to obtain the image consistency verification results corresponding to the infrared animal images, RGB animal images, and depth animal images. When all the liveness detection results are animal liveness and the image consistency verification results are passed, it is determined that the liveness detection of the animal image passes, that is, the animal image is a live animal image.

[0156] It should be understood that although Figures 2 - 9 the steps in the flowchart of Figures 2 - 9 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0157] In one embodiment, as Figure 12 shown, an image detection device 1200 is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: an image acquisition module 1202, an image detection module 1204, and a result determination module 1206, where:

[0158] The image acquisition module 1202 is configured to acquire an image to be detected, and the image to be detected includes at least two different types of images;

[0159] An image detection module 1204 is configured to input an image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images through a feature extraction network, inputs the image features corresponding to at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to at least two different types of images, and performs image consistency verification on the image features corresponding to at least two different types of images using the corresponding branch networks to obtain image consistency verification results corresponding to at least two different types of images.

[0160] A result determination module 1206 is configured to determine the image detection result corresponding to the image to be detected based on the liveness detection result and the image consistency verification result.

[0161] In one embodiment, the image detection module 1204 includes:

[0162] A first liveness detection unit is configured to input the image features corresponding to the first type of image among at least two different types of images into a corresponding first detection branch network for liveness detection to obtain a first liveness detection result corresponding to the first type of image.

[0163] A second liveness detection unit is configured to input the image features corresponding to the second type of image among at least two different types of images into a corresponding second detection branch network for liveness detection to obtain a second liveness detection result corresponding to the second type of image.

[0164] A first verification unit is configured to input the image features corresponding to the first type of image and the image features corresponding to the second type of image into a first verification branch network for image consistency verification to obtain a first image consistency verification result corresponding to the first type of image and the second type of image.

[0165] The result determination module 1206 is further configured to determine the image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, and the first image consistency verification result.

[0166] In one embodiment, the image detection module 1204 further includes:

[0167] A third liveness detection unit is configured to input the image features corresponding to the third type of image among at least two different types of images into a corresponding third detection branch network for liveness detection to obtain a third liveness detection result corresponding to the third type of image.

[0168] The image detection device 1200 further includes:

[0169] A third verification unit is configured to input the image features corresponding to the first type of image and the image features corresponding to the third type of image into the second verification branch network for image consistency verification, so as to obtain a second image consistency verification result corresponding to the first type of image and the third type of image;

[0170] The result determination module 1206 is further configured to determine an image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, and the second image consistency verification result.

[0171] In one embodiment, the image acquisition module 1202 is further configured to acquire an initial video, perform image quality detection on each initial image in the initial video to obtain each initial image quality information; and perform image screening from each initial image quality based on each initial image quality information to obtain an image to be detected.

[0172] In one embodiment, the image detection model includes at least two feature extraction networks; the image detection module 1204 includes:

[0173] A target feature extraction unit is configured to input at least two different types of images into corresponding feature extraction networks for feature extraction respectively, so as to obtain target image features corresponding to at least two different types of images; wherein, the loss value between the network parameters in at least two feature extraction networks is less than a preset threshold;

[0174] A target detection result obtaining unit is configured to input the target image features corresponding to at least two different types of images into corresponding branch networks for liveness detection, so as to obtain target liveness detection results corresponding to at least two different types of images;

[0175] A target verification result obtaining unit is configured to use corresponding branch networks to perform image consistency verification on the target image features corresponding to at least two different types of images, so as to obtain target image consistency verification results corresponding to at least two different types of images.

[0176] In one embodiment, the image detection module 1204 is further configured to calculate a feature similarity between the image features corresponding to the first type of image in at least two different types of images and the image features corresponding to the second type of image in at least two different types of images through a first detection branch network; when the feature similarity exceeds a preset image consistency threshold, it is obtained that the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification passes, and the image consistency verification passing is used to indicate that different types of images are images collected by the same device at the same time.

[0177] In one embodiment, the image detection model includes an image conversion network; the image detection module 1204 further includes:

[0178] An image conversion unit is configured to input a second type of image among at least two different types of images into an image conversion network by an image detection model for image conversion, so as to obtain a converted image with the same type as the first type of image among the at least two different types of images, and use the converted image as the second type of image.

[0179] In one embodiment, the result determination module 1206 is further configured to, when the at least two different types of images are live images and the image consistency verification result is that the image consistency verification is passed, obtain that the image detection result is that the live detection is passed.

[0180] In one embodiment, the image detection device 1200 further includes:

[0181] A data acquisition module is configured to acquire a training data set, where the training data set includes at least two different types of sample images and corresponding class labels;

[0182] An initial training module is configured to input the at least two different types of sample images into an initialized image detection model. The initialized image detection model extracts sample image features corresponding to the at least two different types of sample images respectively through an initialized feature extraction network, inputs the sample image features corresponding to the at least two different types of sample images into corresponding initialized branch networks for live detection, obtains initial live detection results corresponding to the at least two different types of sample images, and performs image consistency verification on the basis of the sample image features corresponding to the at least two different types of sample images respectively by using the corresponding initialized branch networks, so as to obtain initial image consistency verification results corresponding to the at least two different types of sample images;

[0183] A loss calculation module is configured to calculate losses based on the initial live detection results and the corresponding live detection class labels in the class labels, so as to obtain live detection loss information corresponding to the at least two different types of sample images, and calculate losses based on the initial image consistency verification results and the corresponding image consistency class labels in the class labels, so as to obtain image consistency verification loss information corresponding to the at least two different types of sample images;

[0184] An update module is configured to reversely update the corresponding initialized branch network and initialized feature extraction network based on the live detection loss information corresponding to the at least two different types of sample images, and reversely update the corresponding initialized branch network and initialized feature extraction network according to the image consistency verification loss information, so as to obtain an updated image detection model;

[0185] A model acquisition module is configured to use the updated image detection model as the initial image detection model, and return to the step of inputting the at least two different types of sample images into the initialized image detection model, until the training is completed, so as to obtain an image detection model.

[0186] In one embodiment, initializing the image detection model includes at least two initialized feature extraction networks; the updating module includes:

[0187] The network to be confirmed obtaining unit is configured to update the corresponding initialized feature extraction network and the corresponding initialized branch network based on the liveness detection loss information corresponding to at least two different types of sample images, and reversely update the corresponding initialized branch network and the corresponding initialized feature extraction network according to the image consistency verification loss information corresponding to at least two different types of sample images, so as to obtain at least two feature extraction networks to be confirmed;

[0188] The network confirmation module is configured to calculate the parameter loss value between the network parameters in at least two feature extraction networks to be confirmed; when the parameter loss value is less than the preset parameter loss threshold, at least two updated feature extraction networks are obtained, and an updated image detection model is obtained based on the at least two updated feature extraction networks.

[0189] For the specific limitations on the image detection device, reference may be made to the limitations on the image detection method in the above text, which will not be elaborated here. Each module in the above image detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0190] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 13 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as images to be detected or training data sets. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image detection method.

[0191] Those skilled in the art can understand that Figure 13 the structure shown in

[0192] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0193] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0194] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the steps in the above method embodiments.

[0195] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0196] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0197] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An image detection method, characterized in that, The method includes: Obtaining an image to be detected, where the image to be detected includes at least two different types of images; Inputting the image to be detected into an image detection model. The image detection model extracts the image features corresponding to the at least two different types of images respectively through a feature extraction network, inputs the image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain the liveness detection results corresponding to the at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to the at least two different types of images to obtain the image consistency verification results corresponding to the at least two different types of images, including: inputting the image features corresponding to the first type of image in the at least two different types of images into the corresponding first detection branch network for liveness detection to obtain the first liveness detection result corresponding to the first type of image, inputting the image features corresponding to the second type of image in the at least two different types of images into the corresponding second detection branch network for liveness detection to obtain the second liveness detection result corresponding to the second type of image, inputting the image features corresponding to the third type of image in the at least two different types of images into the corresponding third detection branch network for liveness detection to obtain the third liveness detection result corresponding to the third type of image, inputting the image features corresponding to the first type of image and the image features corresponding to the second type of image into the first verification branch network for image consistency verification to obtain the first image consistency verification result corresponding to the first type of image and the second type of image, and inputting the image features corresponding to the first type of image and the image features corresponding to the third type of image into the second verification branch network for image consistency verification to obtain the second image consistency verification result corresponding to the first type of image and the third type of image; Determining the image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, and the second image consistency verification result.

2. The method according to claim 1, wherein The obtaining of the image to be detected includes: Obtaining an initial video, performing image quality detection on each initial image in the initial video to obtain the quality information of each initial image; Performing image screening from the quality of each initial image based on the quality information of each initial image to obtain the image to be detected.

3. The method according to claim 1, wherein The image detection model includes at least two feature extraction networks; The extracting of the image features corresponding to the at least two different types of images respectively through the feature extraction network includes: Inputting the at least two different types of images into the corresponding feature extraction networks respectively for feature extraction to obtain the target image features corresponding to the at least two different types of images; wherein, the loss value between the network parameters in the at least two feature extraction networks is less than a preset threshold. Inputting the image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain the liveness detection results corresponding to the at least two different types of images, and performing image consistency verification using the corresponding branch networks based on the image features corresponding to the at least two different types of images to obtain the image consistency verification results corresponding to the at least two different types of images, includes: Inputting the target image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain the target liveness detection results corresponding to the at least two different types of images; Performing image consistency verification on the target image features corresponding to the at least two different types of images using the corresponding branch networks to obtain the target image consistency verification results corresponding to the at least two different types of images.

4. The method according to claim 1, characterized in that, The performing image consistency verification using the corresponding branch networks based on the image features corresponding to the at least two different types of images to obtain the image consistency verification results corresponding to the at least two different types of images, includes: Calculating the feature similarity between the image features corresponding to the first type of image and the image features corresponding to the second type of image in the at least two different types of images through a first detection branch network; When the feature similarity exceeds a preset image consistency threshold, obtaining that the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification passes, and the image consistency verification passing is used to indicate that different types of images are images collected by the same device at the same time.

5. The method according to claim 1, wherein The image detection model includes an image conversion network; Before extracting the image features corresponding to the at least two different types of images through the feature extraction network, inputting the image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain the liveness detection results corresponding to the at least two different types of images, and performing image consistency verification using the corresponding branch networks based on the image features corresponding to the at least two different types of images to obtain the image consistency verification results corresponding to the at least two different types of images, further includes: The image detection model inputs the second type of image in the at least two different types of images into the image conversion network for image conversion to obtain a converted image with the same type as the first type of image in the at least two different types of images, and uses the converted image as the second type of image.

6. The method according to claim 1, characterized in that, The method further includes: When the at least two different types of images are liveness images and the image consistency verification result is that the image consistency verification passes, obtaining that the image detection result is that the liveness detection passes.

7. The method according to claim 1, characterized in that The training of the image detection model includes the following steps: Obtaining a training data set, where the training data set includes at least two different types of sample images and corresponding class labels; Input the at least two different types of sample images into an initialized image detection model. The initialized image detection model extracts sample image features corresponding to the at least two different types of sample images respectively through an initialized feature extraction network, inputs the sample image features corresponding to the at least two different types of sample images into corresponding initialized branch networks for liveness detection to obtain initial liveness detection results corresponding to the at least two different types of sample images, and performs image consistency verification using the corresponding initialized branch networks based on the sample image features corresponding to the at least two different types of sample images to obtain initial image consistency verification results corresponding to the at least two different types of sample images; Calculate loss based on the initial liveness detection results and the corresponding liveness detection class labels in the class labels to obtain liveness detection loss information corresponding to the at least two different types of sample images, and calculate loss based on the initial image consistency verification results and the corresponding image consistency class labels in the class labels to obtain image consistency verification loss information corresponding to the at least two different types of sample images; Update the corresponding initialized branch network and the initialized feature extraction network in the reverse direction based on the liveness detection loss information corresponding to the at least two different types of sample images, and update the corresponding initialized branch network and the initialized feature extraction network in the reverse direction according to the image consistency verification loss information to obtain an updated image detection model; Use the updated image detection model as the initial image detection model, and return to the step of inputting the at least two different types of sample images into the initialized image detection model until the training is completed to obtain the image detection model.

8. The method according to claim 7, wherein The initialized image detection model includes at least two initialized feature extraction networks; The updating the corresponding initialized branch network and the initialized feature extraction network in the reverse direction based on the liveness detection loss information corresponding to the at least two different types of sample images, and updating the corresponding initialized branch network and the initialized feature extraction network in the reverse direction according to the image consistency verification loss information to obtain an updated image detection model includes: Update the corresponding initialized feature extraction network and the corresponding initialized branch network based on the liveness detection loss information corresponding to the at least two different types of sample images, and update the corresponding initialized branch network and the corresponding initialized feature extraction network in the reverse direction according to the image consistency verification loss information corresponding to the at least two different types of sample images to obtain at least two to-be-confirmed feature extraction networks; Calculate the parameter loss value between the network parameters in the at least two to-be-confirmed feature extraction networks; When the parameter loss value is less than a preset parameter loss threshold, obtain at least two updated feature extraction networks, and obtain the updated image detection model based on the at least two updated feature extraction networks.

9. An image detection device, characterized in that, The device includes: An image acquisition module for acquiring an image to be detected, where the image to be detected includes at least two different types of images; An image detection module, configured to input the image to be detected into an image detection model. The image detection model extracts image features corresponding to at least two different types of images through a feature extraction network, inputs the image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain liveness detection results corresponding to the at least two different types of images, and performs image consistency verification using the corresponding branch networks based on the image features corresponding to the at least two different types of images to obtain image consistency verification results corresponding to the at least two different types of images, including: inputting the image features corresponding to the first type of image among the at least two different types of images into a corresponding first detection branch network for liveness detection to obtain a first liveness detection result corresponding to the first type of image; inputting the image features corresponding to the second type of image among the at least two different types of images into a corresponding second detection branch network for liveness detection to obtain a second liveness detection result corresponding to the second type of image; inputting the image features corresponding to the third type of image among the at least two different types of images into a corresponding third detection branch network for liveness detection to obtain a third liveness detection result corresponding to the third type of image; inputting the image features corresponding to the first type of image and the image features corresponding to the second type of image into a first verification branch network for image consistency verification to obtain a first image consistency verification result corresponding to the first type of image and the second type of image; inputting the image features corresponding to the first type of image and the image features corresponding to the third type of image into a second verification branch network for image consistency verification to obtain a second image consistency verification result corresponding to the first type of image and the third type of image; A result determination module, configured to determine an image detection result corresponding to the image to be detected based on the first liveness detection result, the second liveness detection result, the third liveness detection result, the first image consistency verification result, and the second image consistency verification result.

10. The device according to claim 9, characterized in that, The image acquisition module is further configured to acquire an initial video, perform image quality detection on each initial image in the initial video to obtain quality information of each initial image; and perform image screening on each initial image based on the quality information of each initial image to obtain an image to be detected.

11. The device according to claim 9, wherein, The image detection model includes at least two feature extraction networks; The image detection module includes: A target feature extraction unit, configured to input the at least two different types of images into corresponding feature extraction networks for feature extraction respectively to obtain target image features corresponding to the at least two different types of images respectively; wherein, a loss value between network parameters in the at least two feature extraction networks is less than a preset threshold; A target detection result obtaining unit, configured to input the target image features corresponding to the at least two different types of images into corresponding branch networks for liveness detection to obtain target liveness detection results corresponding to the at least two different types of images; A target verification result obtaining unit, configured to perform image consistency verification on the target image features corresponding to the at least two different types of images respectively using corresponding branch networks, and obtain the target image consistency verification results corresponding to the at least two different types of images.

12. The device according to claim 9, characterized in that, An image detection module, comprising: A similarity calculation unit, configured to calculate the feature similarity between the image features corresponding to the first type of image in the at least two different types of images and the image features corresponding to the second type of image in the at least two different types of images through a first detection branch network; A verification unit, configured to, when the feature similarity exceeds a preset image consistency threshold, obtain that the image consistency verification result corresponding to the first type of image and the second type of image is that the image consistency verification passes, and the image consistency verification passing is used to indicate that the different types of images are images collected by the same device at the same time.

13. The device according to claim 9, characterized in that, The image detection model includes an image conversion network; The image detection module further includes: An image conversion unit, configured to input the second type of image in the at least two different types of images into the image conversion network by the image detection model for image conversion, obtain a converted image with the same type as the first type of image in the at least two different types of images, and use the converted image as the second type of image.

14. The device according to claim 9, characterized in that, The result determination module is further configured to, when the at least two different types of images are live images and the image consistency verification result is that the image consistency verification passes, obtain that the image detection result is that the live detection passes.

15. The device according to claim 9, characterized in that, The device further includes: A data acquisition module, configured to acquire a training data set, where the training data set includes at least two different types of sample images and corresponding class labels; An initial training module, configured to input the at least two different types of sample images into an initialized image detection model, where the initialized image detection model extracts sample image features corresponding to the at least two different types of sample images respectively through an initialized feature extraction network, input the sample image features corresponding to the at least two different types of sample images respectively into corresponding initialized branch networks for live detection, obtain initial live detection results corresponding to the at least two different types of sample images, and perform image consistency verification on the sample image features corresponding to the at least two different types of sample images respectively using corresponding initialized branch networks, and obtain initial image consistency verification results corresponding to the at least two different types of sample images; A loss calculation module, configured to calculate a loss based on the initial live detection results and the corresponding live detection class labels in the class labels, obtain live detection loss information corresponding to the at least two different types of sample images, and calculate a loss based on the initial image consistency verification results and the corresponding image consistency class labels in the class labels, and obtain image consistency verification loss information corresponding to the at least two different types of sample images; An update module, configured to reversely update the corresponding initialized branch network and the initialized feature extraction network based on the liveness detection loss information corresponding to the at least two different types of sample images, and reversely update the corresponding initialized branch network and the initialized feature extraction network according to the image consistency verification loss information, so as to obtain an updated image detection model; A model obtaining module, configured to use the updated image detection model as an initial image detection model, and return the step of inputting the at least two different types of sample images into the initialized image detection model, until the training is completed, to obtain the image detection model.

16. The device according to claim 15, characterized in that, The initialized image detection model includes at least two initialized feature extraction networks; The update module includes: A to-be-confirmed network obtaining unit, configured to update the corresponding initialized feature extraction network and the corresponding initialized branch network based on the liveness detection loss information corresponding to the at least two different types of sample images, and reversely update the corresponding initialized branch network and the corresponding initialized feature extraction network according to the image consistency verification loss information corresponding to the at least two different types of sample images, so as to obtain at least two to-be-confirmed feature extraction networks; A network confirmation module unit, configured to calculate a parameter loss value between network parameters in the at least two to-be-confirmed feature extraction networks; when the parameter loss value is less than a preset parameter loss threshold, obtain at least two updated feature extraction networks, and obtain the updated image detection model based on the at least two updated feature extraction networks.

17. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Living body detection method and device and storage medium

    CN111444744A

  • Human face recognition living body detection method and device, computer equipment and storage medium

    CN111931594A