Living body detection method, device, storage medium and electronic device

By fusing the features and scores of multimodal images in the face recognition system and constraining them with the evaluation results, the problem of poor stability of multimodal input is solved, the accuracy and stability of live detection are improved, and the security of the system is enhanced.

CN116311547BActive Publication Date: 2025-06-06ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310166477.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-06-06
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

It is difficult to achieve good stability in multimodal inputs in facial recognition systems, and problems of missing a certain mode or multimodal misalignment often occur, resulting in damage to the classification results and reducing the security capabilities of the system.

Method used

By obtaining the feature information corresponding to the images of multiple modalities, performing feature fusion and fractional fusion, and combining the evaluation results to improve the accuracy and stability of live detection.

Benefits of technology

The accuracy and stability of multimodal system for live detection is enhanced, the negative impact of damaged mode on the final detection results is reduced, and the security of the image recognition system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311547B_ABST
    Figure CN116311547B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a living body detection method, device, storage medium and electronic device. Feature fusion is performed on feature information corresponding to images of multiple modalities through evaluation results corresponding to images of multiple modalities to obtain fused feature information, and score fusion is performed on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to the fused feature information to obtain a fused score. Liveness verification is performed on the image based on the fused score to determine whether the target corresponding to the image is alive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method, device, storage medium and electronic device for detecting a living body. Background Art

[0002] With the continuous development of face recognition systems in recent years, "liveness detection" technology has become an indispensable part of face recognition systems. "Liveness detection" technology can effectively intercept non-live attack samples, including mobile phone attacks, paper attacks, head models, etc. Offline face recognition technology based on fixed equipment often has problems such as large sample size, different hardware devices, and large differences in the environment in which face recognition is located. Therefore, the robustness requirements of the face recognition link are relatively high. In order to further improve the security capabilities of the face recognition system, multimodal acquisition systems and recognition systems are usually used.

[0003] However, it is difficult to achieve good stability for multimodal inputs for face recognition. Problems such as missing a certain modality or multi-modal misalignment often occur. The damaged modality will have a serious impact on the final classification results, ultimately reducing the security capabilities of the face recognition system. Summary of the invention

[0004] The embodiments of this specification provide a method, device, storage medium and electronic device for detecting a living body, which can enhance the accuracy of detecting a living body and improve the security of an image recognition system. The technical solution is as follows:

[0005] In a first aspect, an embodiment of the present specification provides a method for detecting a living body, the method comprising:

[0006] Obtain feature information corresponding to images of multiple modalities;

[0007] According to evaluation results of the images of the multiple modalities, the plurality of feature information are subjected to feature fusion to obtain fused feature information;

[0008] Score-fusing the plurality of first prediction scores and the second prediction scores to obtain a fusion score; wherein the plurality of first prediction scores are scores obtained by respectively predicting the images of the plurality of modalities, and the second prediction score is a score obtained by predicting the fusion feature information;

[0009] Whether the objects corresponding to the images of the multiple modalities are living objects is detected according to the fusion scores.

[0010] In a second aspect, an embodiment of the present specification provides a living body detection device, the device comprising:

[0011] A feature acquisition module is used to acquire feature information corresponding to images of multiple modalities;

[0012] A feature fusion module, configured to fuse the plurality of feature information according to evaluation results of the plurality of modal images respectively, to obtain fused feature information;

[0013] A score fusion module, used to score-fuse a plurality of first prediction scores and a second prediction score to obtain a fusion score; wherein the plurality of first prediction scores are scores obtained by respectively predicting the images of the plurality of modalities, and the second prediction score is a score obtained by predicting the fusion feature information;

[0014] A living body detection module is used to detect whether the target corresponding to the images of the multiple modalities is a living body according to the fusion score.

[0015] In a third aspect, an embodiment of the present specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0016] In a fourth aspect, an embodiment of the present specification provides a computer program product, wherein the computer program product stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0017] In a fifth aspect, an embodiment of the present specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0018] The beneficial effects brought by the technical solutions provided in some embodiments of the present specification include at least:

[0019] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 It is a flowchart of a liveness detection method provided in an embodiment of this specification;

[0022] Figure 2A - Figure 2C is a multimodal image provided by an embodiment of this specification;

[0023] Figure 3 It is a flowchart of a liveness detection method provided in an embodiment of this specification;

[0024] Figure 4 It is a flowchart of another living body detection method provided in an embodiment of this specification;

[0025] Figure 5 It is a flowchart of a liveness detection method provided in an embodiment of this specification;

[0026] Figure 6 It is a flowchart of another living body detection method provided in an embodiment of this specification;

[0027] Figure 7 is a structural schematic diagram of a living body detection device provided in an embodiment of this specification;

[0028] Figure 8 It is a structural schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0030] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood in specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are an "or" relationship.

[0031] The present specification is described in detail below with reference to specific embodiments.

[0032] With the continuous development of face recognition systems in recent years, "liveness detection" technology has become an indispensable part of face recognition systems. "Liveness detection" technology can effectively intercept non-liveness attack samples, including mobile phone attacks, paper attacks, head models, etc., to ensure the safety of users. For example, when a user completes payment through a mobile terminal by performing face verification, or when a user completes door opening through face verification on a door lock, the collected image including the user's facial feature information must first be subjected to liveness detection. After passing the liveness detection, it is then tested whether the user's facial feature information included in the image is legal and has passed identity authentication.

[0033] Offline face recognition technology based on fixed equipment types often has problems such as large sample size, different hardware equipment, and large differences in the face recognition environment. Therefore, high robustness requirements are placed on the face recognition link.

[0034] In order to further improve the security capabilities of the face recognition system, a multimodal image acquisition system and a recognition system are usually used. The multimodal image acquisition system acquires the multimodal image corresponding to the target image, and the recognition system of the multimodal image acquisition system performs recognition processing on the multimodal image, and determines whether the target image is alive based on the recognition result.

[0035] However, it is difficult to achieve good stability for multimodal inputs for face recognition. The first reason is that the multimodal image acquisition system is composed of multiple high-precision single-modal image acquisition devices. When one of the single-modal image acquisition devices fails, the multimodal image acquisition system will have one or more modes missing or damaged. The second reason is that misalignment problems are prone to occur when processing multimodal images. Ultimately, the damaged mode will have a serious impact on the classification results of the multimodal image, resulting in misjudgment of whether the target corresponding to the multimodal image is alive, resulting in reduced security capabilities of the face recognition system.

[0036] Therefore, in view of the problems in the above-mentioned multi-modal image acquisition system in the detection of living bodies, this specification provides a method for detecting living bodies to solve them. Figure 1 As shown, a method for detecting a living body is proposed in an embodiment of this specification. The method can be implemented by a computer program and can be run on a living body detection device based on a von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0037] Specifically, the living body detection method includes:

[0038] S102, obtaining feature information corresponding to images of multiple modalities respectively;

[0039] An image is a similar and vivid description or portrait of natural things or objective objects (people, animals, plants, landscapes, etc.), or an image can be understood as a way of expressing natural things or objective objects (people, animals, plants, landscapes, etc.), which contains relevant information about the described object. Usually an image is a picture with visual effects.

[0040] Images of multiple modalities can be understood as multiple images obtained by acquiring and processing the same picture through a multimodal image acquisition system, each image corresponds to a modality, and the number of modalities is at least two.

[0041] In one embodiment, the plurality of modalities include at least two of the following modalities: infrared image modality, depth image modality, and color image modality. Figure 2A As shown in, it is a color image acquired by the multimodal image acquisition system, such as Figure 2B As shown, it is the image acquired by the multi-modal image acquisition system. Figure 2A The depth image corresponding to the color image shown is Figure 2C As shown, it is the image acquired by the multi-modal image acquisition system. Figure 2AInfrared image corresponding to the color image shown. It can be understood that in this embodiment, the hand is taken as the target, and the image including the hand is collected by the multi-modal image acquisition system to obtain images of multiple modalities. However, the embodiments of this specification do not limit the type and number of targets, and the multi-modal image acquisition system shown can also collect images including other modalities.

[0042] Acquire the feature information corresponding to each modality image. The feature information of images of different modalities is different. For example, if the image of the modality is a depth image, the feature information of the image includes depth information. The depth information can be understood as the distance value or depth value from any point in the image to the depth modality image acquisition device, that is, the three-dimensional coordinate information of any point. Figure 2B As shown, it is a schematic diagram of a modality image provided by the present embodiment as a depth image, and the areas with different shades of color in the figure correspond to different depth values. For another example, the image of the modality is a color image, which can be a color image in a hue saturation value (HSV) space or a color image in a red green blue (RGB) space, and the feature information of the image includes the color value of each pixel on the image, such as an RGB value or an HSV value.

[0043] The method for obtaining the feature information corresponding to the image of each modality can be to use a pre-trained feature extraction module to extract features from the image of each modality. For example, N single-modality image processing models can be pre-trained, and each single-modality image processing model has a good recognition rate for its corresponding single-modality image. The feature extraction module in such a single-modality image processing model can be used to extract the single-modality feature information in the image of the corresponding modality. For example, the image of this modality is a depth image, and the feature information corresponding to the depth image includes depth information, which is obtained by a single-modality model composed of stereo image matching, a depth camera or a deep learning neural network.

[0044] S104, according to the evaluation results of the images of the multiple modalities, performing feature fusion on the multiple feature information to obtain fused feature information;

[0045] According to the characteristic information of the image of each modality, the image of each modality is evaluated respectively to obtain the evaluation results corresponding to the images of multiple modalities. The evaluation of the image of each modality can be based on confidence or acquisition quality or other standards to obtain the evaluation results.

[0046] In one embodiment, the images of each modality are evaluated based on the confidence of the images of each modality to obtain an evaluation result. The confidence of the image of the modality can be understood as the credibility of the image of the modality, which can usually be represented by a corresponding value on the confidence label. In other words, the value on the confidence label corresponds to the credibility of a single modality image.

[0047] First, the confidence of each modality image is obtained. Classification models such as Convolutional Neural Networks (CNN), KNN (K-nearest neighbors), logistic regression, LDA (Linear discriminant analysis), and QDA (Quadratic discriminant analysis) can be used to predict the feature vector of each modality image input, thereby generating the confidence corresponding to each modality image, that is, obtaining multiple confidences corresponding to multi-modal images.

[0048] Secondly, the images of each modality are evaluated separately based on the confidence of the images of each modality to obtain an evaluation result. For example, the fully connected layer of the CNN model is used to connect multiple confidence outputs for the evaluation value of each modality image as the evaluation result, and the images of the M modalities with evaluation values ​​higher than the preset value are determined as the images of the target modality finally adopted, and the feature information corresponding to the images of the target modality finally adopted is used for feature fusion. For another example, after the CNN model is used to obtain the evaluation value corresponding to the image of each modality, different weights T are assigned to the images of N modalities with different evaluation values ​​based on preset rules. The preset rule can be that the higher the evaluation value, the higher the weight T corresponding to. 1 For another example, by using classification models such as KNN, logistic regression, LDA, and QDA, the input prediction confidence is directly classified into a certain category, and the M modal images corresponding to the category are the images of the target modality finally adopted, or the weight ratio corresponding to the category is the weight ratio adopted when the N modal images are subjected to feature fusion.

[0049] In this embodiment, the evaluation result of each modality is obtained by the confidence of each modality, so that when feature fusion is performed according to the evaluation result, the influence of the feature information of the image of the modality with low confidence on the fused feature information is reduced, thereby improving the credibility of the fused feature information.

[0050] In another embodiment, the image of each modality is evaluated separately based on the quality of the image of each modality, thereby obtaining an evaluation result. The acquisition devices corresponding to the images of different modalities are all affected by the acquired working or acquisition environment. The quality of the image of the modality can usually be characterized by a corresponding value in the quality dimension. In other words, the value in the quality dimension corresponds to the degree of damage on a single modality image. The higher the quality value, the lower the degree of damage corresponding to the image.

[0051] Specifically, the quality of each modality image is evaluated by preset rules to obtain an evaluation result. For example, the preset rules include the clarity, brightness, etc. of each modality image and the noise and information content corresponding to multiple pixels of each modality image. In one embodiment, the evaluation result is and the M modality images with quality values ​​higher than the preset value are determined as the final target modality images, and the feature information corresponding to the final target modality images is used for feature fusion. In another embodiment, the evaluation result includes the quality values ​​corresponding to the images of multiple modalities, and further assigns different weights T to the N modal images with different quality values. 2 , the rule for assigning weight parameter distribution can be that the higher the quality value, the higher the weight T 2 The bigger.

[0052] In this embodiment, the evaluation result of each modality is obtained by the quality of each modality, so that when feature fusion is performed according to the evaluation result, the influence of feature information of the low-quality modality image on the fused feature information is reduced, thereby improving the credibility of the fused feature information.

[0053] In one embodiment, the images of different modalities can be segmented before quality assessment to obtain segmented images, including the target object. The segmented images can be obtained by any one or more methods such as histogram threshold method, region growing method, image-based random field model method, relaxed labeling region segmentation method, etc., and this specification embodiment does not impose any limitation on this.

[0054] In this embodiment, quality assessment is only performed on the segmented image, which can reduce the amount of computation consumed by the processor during quality assessment, improve assessment efficiency, and reduce the impact of the quality of non-target areas on the final assessment result.

[0055] Furthermore, according to the evaluation results of the images of the multiple modalities, the multiple feature information is fused to obtain fused feature information. The fusion method can be direct splicing, or the fusion method can also be weighted fusion based on weight distribution parameters to generate fused feature information.

[0056] For example, the evaluation results of the images of multiple modalities are evaluated according to the confidence or quality. The evaluation result is to determine that the images corresponding to M modalities among N modalities are the images of the target modality finally adopted, M=3, and the feature information of the images of each modality is represented in the form of a vector. The feature information corresponding to the images of the M modalities is respectively a 1 、a 2 、a 3 , then the corresponding fusion feature information can be obtained by direct splicing (a 1 , a 2 , a 3 ).

[0057] In another embodiment, the evaluation results of the images of multiple modalities are evaluated separately according to the confidence or quality, and the evaluation result is to assign different weights to the images of N modalities, and then the feature information can be weighted based on the weight distribution parameter. For example, assuming that a preset weight distribution parameter (w 1 , w 2 , w 3 ), the vector corresponding to the feature information of the images of multiple modalities is b 1 、b 2 、b 3 , then the fusion feature information after weighted fusion can be (w 1 ×b 1 , w 2 ×b 2 , w 3 ×b 3 ).

[0058] The evaluation results of the images of multiple modalities include the weight distribution parameters during fusion or the image of the target modality, so that the multiple feature information are fused through the evaluation results to obtain fused feature information, which can highlight the feature information of certain corresponding modalities or reduce the feature information of certain modalities, so that the final fused feature information is more in line with the actual situation.

[0059] S106, fusing the multiple first prediction scores and the second prediction scores;

[0060] The multiple first prediction scores are scores obtained by predicting the images of multiple modalities respectively, and the second prediction score is a score obtained by predicting the fused feature information. The first prediction score of the image of the modality can be understood as the possibility that the target corresponding to the image of the modality is a living body, and the second prediction score can be understood as the possibility that the target corresponding to the fused feature information is a living body. The first prediction score and the second prediction score can usually be represented by corresponding values ​​in the dimension of whether the living body detection is passed.

[0061] In other words, the value of the first prediction score corresponds to the possibility that the target corresponding to a single-modality image is a living body. For example, if the first prediction score of an image of a certain modality is 1, then the target corresponding to the image of this modality is very likely to be a living body, and if the first prediction score of an image of a certain modality is 0.5, then the target part corresponding to the image of this modality is likely to be a living body.

[0062] First, a first prediction score corresponding to the image of each modality and a second prediction score corresponding to the fused feature information are obtained. Classification models such as CNN network CNN and KNN, logistic regression, LDA, QDA, etc. can be used to predict the feature vector and fused feature information of each modality image input, thereby generating a first prediction score corresponding to the image of each modality and a second prediction score corresponding to the fused feature information.

[0063] Secondly, multiple first prediction scores and second prediction scores are score-fused. The score fusion method can be to fuse multiple first prediction scores and second prediction scores based on weight distribution parameters, wherein the weight distribution parameters are obtained by being preset in a storage unit, after evaluating the image of each modality, or through the evaluation result of S106 or other methods. The score fusion method can also be to take the average of multiple first prediction scores and second prediction scores or any other calculation method.

[0064] For example, the weight distribution parameter is the same as the weight distribution parameter when the features of images of multiple modalities are fused, and the multiple first prediction scores and the second prediction scores are score fused based on the weight distribution parameter. Specifically, the image of each modality is evaluated to obtain the evaluation result; based on other preset rules, the images of N modalities are assigned different weights T 2 , assuming that the preset weight distribution parameter (w 1 , w 2 , w 3 ), the vector corresponding to the feature information of the images of multiple modalities is b 1 、b 2 、b 3 , then the fusion feature information after weighted fusion can be (w 1 ×b 1 , w 2 ×b 2 , w 3 ×b 3 ); The first prediction score corresponding to the images of multiple modalities is c 1 、c 2 、c 3 , to fusion feature information (w 1 ×b 1 , w 2 ×b 2, w 3 ×b 3) The second prediction score is c 0 , and finally the fusion score is (w 1 ×c 1 +w 2 ×c 2 +w 3 ×c 3 +w 0 ×c 0 ) / N, where the weight w corresponding to the fusion feature information 0 is the preset weight value, for example, w 0 =1.

[0065] It is understandable that the weight distribution parameters may be different from the weight distribution parameters when performing feature fusion on images of multiple modalities. For example, when performing feature fusion, images of multiple modalities are evaluated by confidence to obtain a first evaluation result, and when performing score fusion, images of multiple modalities are evaluated by quality evaluation to obtain a second evaluation result, then the weight distribution parameters corresponding to the first evaluation result and the second evaluation result are different.

[0066] In this embodiment, weight values ​​are assigned to a plurality of first prediction scores and second prediction scores to reduce the influence of the first prediction scores of low-quality or low-credibility modal images on the fusion score, thereby improving the credibility of the fusion score.

[0067] S108. Detect whether the targets corresponding to the images of multiple modalities are living objects according to the fusion scores.

[0068] The fusion score represents the possibility of whether the target corresponding to the images of multiple modalities is alive. The higher the score, the higher the possibility of the target being alive. FIG. 2A to FIG. 2C As shown in the figure, the images of multiple modes are RGB images, depth images and infrared images. The feature information corresponding to the images of the above three modes are fused to obtain fused feature information; the image of each mode is predicted to obtain three first prediction scores, and the fused feature information is predicted to obtain a second prediction score. The three first prediction scores and the second prediction scores are further fused to obtain a fused score. The fused score is used to judge FIG. 2A to FIG. 2C Whether the target is alive, for example, if the fusion score is 0.8, then FIG. 2A to FIG. 2C The targets shown have passed in vivo validation.

[0069] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0070] In one embodiment, Figure 3 As shown, another method for detecting live bodies proposed in the embodiment of this specification can be realized by a computer program and can be run on a live body detection device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0071] Specifically, the living body detection method includes:

[0072] S202, obtaining feature information corresponding to images of multiple modalities respectively;

[0073] See the above S102, which will not be described again here.

[0074] S204, performing quality assessment on the image of each modality according to feature information corresponding to the images of multiple modalities, to obtain an assessment result corresponding to the image of each modality;

[0075] The acquisition devices corresponding to the images of different modalities are all affected by many aspects such as the acquired work or acquisition environment, which may result in different quality of images of different modalities. For example, a multimodal image acquisition system includes an RGB image acquisition device, an infrared image acquisition device, and a depth image acquisition device. The infrared image acquisition device has a lens wear problem, so the infrared modality image acquired by the infrared image acquisition device is an image of a damaged modality, and the quality of the infrared modality image is lower than the depth modality image acquired by the depth image acquisition device and lower than the RGB modality image acquired by the RGB image acquisition device.

[0076] The image quality of a modality can usually be characterized by a corresponding value in the quality dimension. In other words, the value in the quality dimension corresponds to the degree of damage on a single modality image. The higher the quality value, the lower the corresponding degree of damage.

[0077] Specifically, the image quality of each modality is evaluated based on the characteristic information of the image of each modality, thereby obtaining an evaluation result. For example, the image quality of each modality is evaluated based on the clarity, brightness, etc. of the image of each modality, as well as the noise and information amount corresponding to the multiple pixels of the image of each modality, thereby obtaining an evaluation result.

[0078] In one embodiment, according to the modality type corresponding to the image of each modality, feature extraction of the depth corresponding to the modality type is performed on the image of each modality to obtain feature information corresponding to the image of each modality; the depths of feature extraction corresponding to different modality types are not completely the same; according to the feature information corresponding to the image of each modality, the image quality is evaluated to obtain the evaluation result corresponding to the image of each modality. In other words, images of different modalities correspond to feature information extraction of different depths.

[0079] For example, a multimodal image acquisition system includes an RGB image acquisition device, an infrared image acquisition device, and a depth image acquisition device. The acquired multimodal images include images of RGB modality, images of infrared modality, and images of depth modality, wherein the depth of feature information extraction of RGB images is deeper than that of infrared modality images and depth modality images. The reason is that the information density of feature information of a single pixel in an RGB modality image is higher than that of images of other modalities, that is, the texture information of images of other modalities is weaker than that of images of RGB modality. It is necessary to perform a deeper expansion of the feature information of the RGB modality image to extract the feature information of the RGB modality image and judge the quality of the RGB modality image, while images of other modalities only require shallow feature information extraction.

[0080] Specifically, the three resblock layers in the convolutional neural network are used to extract features of the RGB modality images, infrared modality images, and depth modality images respectively. For the infrared modality images and the depth modality images, the feature information of 64×64 dimensions is extracted after two resblocks for quality assessment, and for the RGB modality images, the feature information of 128×128 dimensions is extracted after three resblocks for quality assessment.

[0081] In this embodiment, feature extraction of different depths is adopted for images of different modalities, so that quality assessment is performed based on the extracted features. Considering that the textures of images of different modalities are different, the feature information extracted during quality assessment is more reasonable.

[0082] S206, performing feature fusion on multiple feature information according to the evaluation results corresponding to the image of each modality to obtain fused feature information;

[0083] According to the evaluation results obtained by performing quality evaluation on images of multiple modalities, multiple feature information is feature fused to obtain fused feature information. For example, the evaluation result is to determine the images of M modalities with quality values ​​higher than a preset value as the images of the target modality finally adopted, and the feature information corresponding to the images of the target modality finally adopted is used for feature fusion, that is, the feature information of the images of the M target modalities is fused, and the fusion method is direct splicing.

[0084] For another example, the evaluation result includes quality values ​​corresponding to images of multiple modalities, and different weight values ​​are assigned to images of N modalities with different quality values. The rule for assigning weight parameter distribution may be that the higher the quality value, the greater the weight corresponding to it. The fusion method may be to generate fusion feature information based on weighted fusion of weight distribution parameters.

[0085] S208, fusing the multiple first prediction scores and the second prediction scores to obtain a fusion score;

[0086] See the above S106, which will not be repeated here.

[0087] S210: Detect whether the target corresponding to the images of multiple modalities is a living body according to the fusion score.

[0088] See the above S108, which will not be repeated here.

[0089] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0090] In one embodiment, Figure 4As shown, another method for detecting live bodies proposed in the embodiment of this specification can be realized by a computer program and can be run on a live body detection device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0091] Specifically, the living body detection method includes:

[0092] S302, obtaining feature information corresponding to images of multiple modalities respectively;

[0093] See the above S102, which will not be described again here.

[0094] S304, according to the evaluation results of the images of the multiple modalities, the multiple feature information is fused to obtain fused feature information;

[0095] See the above S104, which will not be described again here.

[0096] S306, according to the evaluation result corresponding to the image of each modality, the first prediction score and the second prediction score corresponding to the image of each modality are score-fused to obtain a fusion score;

[0097] The multiple first prediction scores are scores obtained by predicting images of multiple modalities respectively, and the second prediction score is a score obtained by predicting the fused feature information. The first prediction score of the image of the modality can be understood as the possibility that the target corresponding to the image of the modality is a living body, and the second prediction score can be understood as the possibility that the target corresponding to the fused feature information is a living body. The first prediction score and the second prediction score can usually be characterized by corresponding values ​​in the dimension of whether the living body detection is passed. The method for obtaining the first prediction score and the second prediction score is shown in S106, which will not be repeated here.

[0098] Furthermore, after obtaining multiple first prediction scores and second prediction scores, based on the evaluation result of the image of each modality in S304, the first prediction score and the second prediction score corresponding to the image of each modality are score-fused to obtain a fusion score.

[0099] For example, in S304, images of multiple modalities are evaluated by quality value, confidence or other evaluation criteria, and the evaluation result is to determine that images corresponding to M modalities out of N modalities are images of the target modality finally adopted, and the first prediction scores and second prediction scores corresponding to the images corresponding to the M modalities are averaged to obtain the fusion score.

[0100] For another example, in S304, images of multiple modalities are evaluated by quality value or confidence or other evaluation criteria, and the evaluation result is to assign different weights to the images of N modalities, that is, to obtain the weight distribution parameters corresponding to the images of N modalities. For example, the weight distribution parameter is (h 1 ,h 2 ,h 3 ), the first prediction score corresponding to the images of multiple modalities is d 1 ,d 2 ,d 3 , the second prediction score for the fused feature information is d 0 , and finally the fusion score is (h 1 ×d 1 +h 2 ×d 2 +h 3 ×d 3 +h 0 ×d 0 ) / N, where the weight d corresponding to the fusion feature information 0 is the preset weight value, for example, d 0 =1.

[0101] It is understandable that the embodiments of this specification also include other content or form evaluation results and a calculation method for obtaining a fusion score based on the evaluation results.

[0102] S308: Detect whether the target corresponding to the images of multiple modalities is a living body according to the fusion score.

[0103] See the above S108, which will not be repeated here.

[0104] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0105] In one embodiment, Figure 5 As shown, another method for detecting live bodies proposed in the embodiment of this specification can be realized by a computer program and can be run on a live body detection device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0106] Specifically, the living body detection method includes:

[0107] S402, obtaining feature information corresponding to images of multiple modalities respectively;

[0108] See the above S102, which will not be described again here.

[0109] S404, obtaining a second weight corresponding to the image of each modality according to the evaluation result corresponding to the image of each modality;

[0110] The quality of images of multiple modalities is evaluated to obtain evaluation results of the images of each modality, wherein the multiple evaluation results include multiple levels, and the value of the second weight is associated with the level of the evaluation result. For example, the number of evaluation levels is the same as the number of modalities of the multimodal image acquisition system, and the evaluation levels include high quality, medium quality, and low quality. The high quality evaluation level corresponds to the second weight x 1 , the medium quality evaluation level corresponds to the second weight x 2 , the low quality evaluation level corresponds to the second weight x 3 , x 1 >x 2 >x 3 .

[0111] S406, performing feature fusion on multiple feature information according to second weights corresponding to images of multiple modalities to obtain fused feature information;

[0112] According to the obtained parameter distribution of the second weight, multiple feature information are fused to obtain fused feature information. For example, the distribution of the second weight is (x 1 , x 2 , x 3 ), the vector corresponding to the feature information of the images of multiple modalities is e 1 、e 2 、e 3 , then the fusion feature information after weighted fusion is (x 1 ×e 1 , x 2 ×e 2 , x 3 ×e 3 ).

[0113] S408, predicting images of multiple modalities respectively to obtain a first prediction score corresponding to the image of each modality;

[0114] Specifically, before predicting images of multiple modalities, the images of the multiple modalities are respectively subjected to region enhancement processing according to the attention mechanism to obtain region enhanced images of multiple modalities; the region enhanced images of multiple modalities are respectively predicted to obtain first prediction scores corresponding to the images of each modality. The attention mechanism can make the processor focus more on the target area of ​​the image when predicting the image, predict whether the target area of ​​the image is alive, and thus obtain a prediction score.

[0115] In liveness detection, it is hoped that the processor will focus more on the facial features in the image for analysis and prediction. Therefore, by using the attention mechanism to perform regional enhancement processing on the image of each modality, the accuracy of judging whether the image of each modality is live can be improved.

[0116] S410, obtaining a first weight corresponding to the image of each modality according to the evaluation result corresponding to the image of each modality;

[0117] The quality of images of multiple modalities is evaluated to obtain evaluation results of images of each modality. The multiple evaluation results include multiple levels, and the value of the first weight is associated with the level of the evaluation result. For example, the number of evaluation levels is the same as the number of modalities of the multimodal image acquisition system, and the evaluation levels include high quality, medium quality and low quality. The high quality evaluation level corresponds to the second weight y 1 , the medium quality evaluation level corresponds to the second weight y 2 , the low quality evaluation level corresponds to the second weight y 3 ,y 1 >y 2 >y 3 .

[0118] In this embodiment, the values ​​of the first weight and the second weight may be the same or different. For example, according to actual computing requirements, although the feature fusion and the score fusion are graded, the value of the second weight corresponding to the preset feature fusion is different from the value of the first weight corresponding to the score fusion.

[0119] S412, fusing the plurality of first prediction scores and second prediction scores according to the first weight corresponding to each first prediction score to obtain a fusion score;

[0120] According to the first weight corresponding to each first prediction score, the multiple first prediction scores and the second prediction scores are weightedly calculated and fused. The first prediction scores corresponding to the images of the multiple modalities are f 1 、f 2、f 3 , to fusion feature information (x 1 ×e 1 , x 2 ×e 2 , x 3 ×e 3) The second prediction score is f 0 , and finally the fusion score is (y 1 ×f 1 +y 2 ×f 2 +y 3 ×f 3 +y 0 ×f 0 ) / N, where the weight y corresponding to the fusion feature information 0 is the preset weight value, for example, y 0 =1.

[0121] S414: Detect whether the target corresponding to the images of multiple modalities is a living body according to the fusion score.

[0122] See the above S108, which will not be repeated here.

[0123] like Figure 6 As shown, Figure 6 It is a flowchart of another liveness detection method provided in an embodiment of this specification. Figure 6 The image includes three modes, namely, RGB mode image, infrared mode image and depth mode image.

[0124] The three resblock layers in the convolutional neural network are used to extract features of the RGB, infrared and depth images respectively. The feature information of 64×64 dimensions is extracted from the infrared and depth images after two resblocks for quality assessment, and the feature information of 128×128 dimensions is extracted from the RGB images after three resblocks for quality assessment.

[0125] Furthermore, according to the weight parameter distribution obtained by quality assessment, the images of the three modalities are fused to obtain a fused feature vector.

[0126] Furthermore, the regional features of the three modal images are enhanced based on the attention block in the convolutional neural network, and then the first prediction scores of the three modal images are obtained through one resblock feature extraction and global average pooling (GAP). And the second prediction score is obtained after one resblock feature extraction and GAP processing are performed on the fused feature vector.

[0127] Furthermore, based on the weight parameter distribution, the first prediction scores corresponding to the three modal images and the second prediction scores of the fused feature vectors are score fused to obtain a fusion score.

[0128] Finally, classification is performed based on the fusion score to obtain the classification result, that is, whether the target corresponding to the image of the three modalities is a living object.

[0129] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0130] The following are device embodiments of this specification, which can be used to implement the method embodiments of this specification. For details not disclosed in the device embodiments of this specification, please refer to the method embodiments of this specification.

[0131] See also Figure 7 , which shows a schematic diagram of the structure of a liveness detection device provided by an exemplary embodiment of this specification. The liveness detection device can be implemented as all or part of the device through software, hardware or a combination of both. The liveness detection device includes a request acquisition module 701, an address resolution module 702, a sending authentication module 703, and a data acquisition module 704.

[0132] A feature acquisition module 701 is used to acquire feature information corresponding to images of multiple modalities;

[0133] A feature fusion module 702 is used to fuse the plurality of feature information according to evaluation results of the plurality of modal images to obtain fused feature information;

[0134] The score fusion module 703 is used to score-fuse the multiple first prediction scores and the second prediction scores to obtain a fusion score; wherein the multiple first prediction scores are scores obtained by respectively predicting the images of the multiple modalities, and the second prediction score is a score obtained by predicting the fusion feature information;

[0135] The living body detection module 704 is used to detect whether the target corresponding to the images of the multiple modalities is a living body according to the fusion score.

[0136] In one embodiment, the score fusion module 703 further includes:

[0137] The score fusion unit is used to score-fuse the first prediction score and the second prediction score corresponding to each image of the modality according to the evaluation result corresponding to the image of each modality to obtain a fusion score.

[0138] In one embodiment, the score fusion unit includes:

[0139] A first prediction subunit is used to predict the images of the multiple modalities respectively to obtain a first prediction score corresponding to the image of each modality;

[0140] A first weight subunit, configured to obtain a first weight corresponding to the image of each modality according to the evaluation result corresponding to each image of each modality;

[0141] A second prediction subunit is used to predict the fused feature information to obtain a second prediction score corresponding to the fused feature information;

[0142] The score fusion subunit is used to score-fuse multiple first prediction scores and the second prediction scores according to the first weight corresponding to each first prediction score to obtain a fusion score.

[0143] In one embodiment, the first prediction subunit is used to perform region enhancement processing on the images of the multiple modalities according to the attention mechanism to obtain region enhanced images of the multiple modalities;

[0144] And used to predict the regional enhanced images of the multiple modalities respectively, to obtain the first prediction scores corresponding to the images of each modality respectively.

[0145] In one embodiment, the feature fusion module 702 includes:

[0146] A quality assessment unit, configured to perform quality assessment on the image of each modality according to feature information corresponding to the image of each modality, and obtain an assessment result corresponding to the image of each modality;

[0147] The feature fusion unit is used to fuse the plurality of feature information according to the evaluation results corresponding to the images of each modality to obtain fused feature information.

[0148] In one embodiment, the feature fusion unit includes:

[0149] An evaluation subunit, configured to obtain a second weight corresponding to the image of each modality according to an evaluation result corresponding to the image of each modality;

[0150] The fusion subunit is used to fuse the plurality of feature information according to the second weight corresponding to the image of each modality to obtain fused feature information.

[0151] In one embodiment, the plurality of evaluation results include a plurality of grades, and the value of the second weight is associated with the grades of the evaluation results.

[0152] In one embodiment, the evaluation subunit is specifically used to perform feature extraction of the depth corresponding to the modality type on the image of each modality according to the modality type corresponding to the image of each modality, so as to obtain feature information corresponding to the image of each modality; the depth of feature extraction corresponding to different modality types is not completely the same;

[0153] And it is used to evaluate the quality of the image according to the feature information corresponding to the image of each modality, so as to obtain the evaluation result corresponding to the image of each modality.

[0154] In one embodiment, the plurality of modalities include at least two of the following modalities: an infrared image modality, a depth image modality, and a color image modality.

[0155] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0156] It should be noted that the liveness detection device provided in the above embodiment only uses the division of the above functional modules as an example when executing the liveness detection method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the liveness detection device provided in the above embodiment and the liveness detection method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0157] The serial numbers of the embodiments of this specification are for description only and do not represent the advantages or disadvantages of the embodiments.

[0158] The present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figure 1 - Figure 6 The liveness detection method of the embodiment shown in the figure can be specifically implemented by referring to Figure 1 - Figure 6 The specific description of the illustrated embodiment will not be repeated here.

[0159] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figure 1 - Figure 6 The liveness detection method of the embodiment shown in the figure can be specifically implemented by referring to Figure 1 - Figure 6 The specific description of the illustrated embodiment will not be repeated here.

[0160] See also Figure 8 , is a schematic diagram of the structure of an electronic device provided in the embodiment of this specification. Figure 8 As shown, the electronic device 800 may include: at least one processor 801 , at least one network interface 804 , a user interface 803 , a memory 805 , and at least one communication bus 802 .

[0161] The communication bus 802 is used to realize the connection and communication between these components.

[0162] The user interface 803 may include a display screen (Display) and a camera (Camera), and the optional user interface 803 may also include a standard wired interface and a wireless interface.

[0163] The network interface 804 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0164] Among them, the processor 801 may include one or more processing cores. The processor 801 uses various interfaces and lines to connect various parts within the entire server 800, and executes various functions and processes data of the server 800 by running or executing instructions, programs, code sets or instruction sets stored in the memory 805, and calling data stored in the memory 805. Optionally, the processor 801 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 801 can integrate one or more combinations of a processor (Central Processing Unit, CPU), an image processor (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 801, and it can be implemented separately through a chip.

[0165] Among them, the memory 805 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 805 includes a non-transitory computer-readable storage medium. The memory 805 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 805 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 805 may also be optionally at least one storage device located away from the aforementioned processor 801. As Figure 8 As shown, the memory 805 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a living body detection application.

[0166] exist Figure 8In the electronic device 800 shown, the user interface 803 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 801 can be used to call the liveness detection application stored in the memory 805 and specifically perform the following operations:

[0167] Obtain feature information corresponding to images of multiple modalities;

[0168] According to evaluation results of the images of the multiple modalities, the plurality of feature information are subjected to feature fusion to obtain fused feature information;

[0169] Score-fusing the plurality of first prediction scores and the second prediction scores to obtain a fusion score; wherein the plurality of first prediction scores are scores obtained by respectively predicting the images of the plurality of modalities, and the second prediction score is a score obtained by predicting the fusion feature information;

[0170] Whether the objects corresponding to the images of the multiple modalities are living objects is detected according to the fusion scores.

[0171] In one embodiment, the processor 801 performs the score fusion of the plurality of first prediction scores and the second prediction scores to obtain the fusion score, specifically performing:

[0172] According to the evaluation result corresponding to the image of each modality, the first prediction score and the second prediction score corresponding to the image of each modality are score-fused to obtain a fusion score.

[0173] In one embodiment, the processor 801 executes the evaluation result corresponding to the image of each modality, and fuses the first prediction score and the second prediction score corresponding to the image of each modality to obtain a fusion score, specifically performing:

[0174] Predicting the images of the multiple modalities respectively to obtain first prediction scores corresponding to the images of each modality;

[0175] According to the evaluation results corresponding to the images of each modality, obtaining the first weight corresponding to the images of each modality;

[0176] Predicting the fused feature information to obtain a second prediction score corresponding to the fused feature information;

[0177] According to a first weight corresponding to each of the first prediction scores, a plurality of the first prediction scores and the second prediction scores are score-fused to obtain a fusion score.

[0178] In one embodiment, the processor 801 performs the prediction of the images of the multiple modalities to obtain the first prediction scores corresponding to the images of each modality, specifically performing:

[0179] Performing region enhancement processing on the images of the multiple modalities according to the attention mechanism to obtain region enhanced images of the multiple modalities;

[0180] Predict the regional enhanced images of the multiple modalities respectively to obtain first prediction scores corresponding to the images of each modality respectively.

[0181] In one embodiment, the processor 801 performs the feature fusion of the plurality of feature information according to the evaluation results of the images of the plurality of modalities respectively to obtain the fused feature information, specifically performing:

[0182] Performing a quality assessment on the image of each modality according to feature information corresponding to the image of each modality, to obtain an assessment result corresponding to the image of each modality;

[0183] According to the evaluation results corresponding to the images of each modality, multiple feature information are fused to obtain fused feature information.

[0184] In one embodiment, the processor 801 performs the feature fusion of the plurality of feature information according to the evaluation result corresponding to the image of each modality to obtain the fused feature information, specifically performing:

[0185] Obtaining a second weight corresponding to the image of each modality according to the evaluation result corresponding to the image of each modality;

[0186] According to the second weight corresponding to the image of each modality, the plurality of feature information are fused to obtain fused feature information.

[0187] In one embodiment, the plurality of evaluation results include a plurality of grades, and the value of the second weight is associated with the grades of the evaluation results.

[0188] In one embodiment, the processor 801 performs the quality assessment of the image of each modality according to the feature information corresponding to the image of each modality to obtain the assessment result corresponding to the image of each modality, specifically performing:

[0189] According to the modality type corresponding to the image of each modality, feature extraction of the depth corresponding to the modality type is performed on the image of each modality to obtain feature information corresponding to the image of each modality; the depth of feature extraction corresponding to different modality types is not completely the same;

[0190] The image quality is evaluated according to the feature information corresponding to the image of each modality to obtain an evaluation result corresponding to the image of each modality.

[0191] In one embodiment, the plurality of modalities include at least two of the following modalities: an infrared image modality, a depth image modality, and a color image modality.

[0192] This specification performs feature fusion on feature information corresponding to images of multiple modalities based on evaluation results corresponding to images of multiple modalities to obtain fused feature information, and performs score fusion on first prediction scores corresponding to images of multiple modalities and second prediction scores corresponding to fused feature information to obtain fused scores, and performs liveness verification on images based on the fused scores to determine whether the target corresponding to the image is live. In other words, the embodiment of this specification performs an evaluation before performing feature fusion on images of multiple modalities, and the evaluation results play a constraining role in the feature fusion process and the score fusion process, thereby achieving two-way constraint effects in the multimodal system. Therefore, when a single modality has a problem of poor image quality in the multimodal system, the damaged modal image is evaluated to reduce the contribution of the modal image in the feature fusion link and the score fusion link, thereby reducing the negative impact of the damaged modal image on the final liveness detection, enhancing the accuracy and stability of the multimodal system for liveness detection, and improving the security of the image recognition system.

[0193] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the object features, interactive behavior features and user information involved in this specification are all obtained with full authorization.

[0194] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0195] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for detecting a living body, the method comprising: include: Obtain feature information corresponding to images of multiple modalities; According to evaluation results of the images of the multiple modalities, the plurality of feature information are subjected to feature fusion to obtain fused feature information; Score-fusing the plurality of first prediction scores and the second prediction scores to obtain a fusion score; wherein the plurality of first prediction scores are scores obtained by respectively predicting the images of the plurality of modalities, and the second prediction score is a score obtained by predicting the fusion feature information; Detecting whether the targets corresponding to the images of the multiple modalities are living objects according to the fusion scores; The step of fusing the multiple first prediction scores and the second prediction scores to obtain a fusion score includes: According to the evaluation result corresponding to the image of each modality, the first prediction score and the second prediction score corresponding to the image of each modality are score-fused to obtain a fusion score; The step of fusing the first prediction score and the second prediction score corresponding to each of the images of the modalities according to the evaluation results corresponding to the images of each modality to obtain a fusion score includes: Predicting the images of the multiple modalities respectively to obtain first prediction scores corresponding to the images of each modality; According to the evaluation result corresponding to the image of each modality, obtaining the first weight corresponding to the image of each modality; Predicting the fused feature information to obtain a second prediction score corresponding to the fused feature information; fusing the first prediction scores and the second prediction scores according to a first weight corresponding to each of the first prediction scores to obtain a fusion score; The step of fusing the plurality of feature information based on the evaluation results of the images of the plurality of modalities to obtain the fused feature information includes: Obtaining a second weight corresponding to the image of each modality according to the evaluation result corresponding to the image of each modality; According to the second weight corresponding to the image of each modality, the plurality of feature information are fused to obtain fused feature information.

2. The method for detecting living bodies according to claim 1, wherein the images of the multiple modalities are predicted respectively to obtain a first prediction score corresponding to the image of each modality. include: Performing region enhancement processing on the images of the multiple modalities according to the attention mechanism to obtain region enhanced images of the multiple modalities; Predict the regional enhanced images of the multiple modalities respectively to obtain first prediction scores corresponding to the images of each modality respectively.

3. The method for detecting living bodies according to claim 1, before obtaining the second weight corresponding to the image of each modality according to the evaluation result corresponding to the image of each modality, further comprising: include: The quality of the image of each modality is evaluated according to the feature information corresponding to the image of each modality, so as to obtain the evaluation result corresponding to the image of each modality. 4 . The liveness detection method according to claim 1 , wherein the plurality of evaluation results include a plurality of levels, and the value of the second weight is associated with the level of the evaluation result.

5. The method for detecting living bodies according to claim 3, wherein the quality of the image of each modality is evaluated according to the feature information corresponding to the image of each modality to obtain the evaluation result corresponding to the image of each modality, include: According to the modality type corresponding to the image of each modality, feature extraction of the depth corresponding to the modality type is performed on the image of each modality to obtain feature information corresponding to the image of each modality; the depth of feature extraction corresponding to different modality types is not completely the same; The image quality is evaluated according to the feature information corresponding to the image of each modality to obtain an evaluation result corresponding to the image of each modality. 6 . The living body detection method according to claim 1 , wherein the multiple modalities include at least two of the following modalities: an infrared image modality, a depth image modality, and a color image modality.

7. A living body detection device, the device include: A feature acquisition module is used to acquire feature information corresponding to images of multiple modalities; A feature fusion module, configured to fuse the plurality of feature information according to evaluation results of the plurality of modal images respectively, to obtain fused feature information; A score fusion module, used to score-fuse a plurality of first prediction scores and a second prediction score to obtain a fusion score; wherein the plurality of first prediction scores are scores obtained by respectively predicting the images of the plurality of modalities, and the second prediction score is a score obtained by predicting the fusion feature information; A living body detection module, used for detecting whether the target corresponding to the images of the multiple modalities is a living body according to the fusion score; Wherein, the score fusion module also includes: A score fusion unit, configured to perform score fusion on the first prediction score and the second prediction score corresponding to each image of the modality according to the evaluation result corresponding to the image of each modality, to obtain a fusion score; Wherein, the score fusion unit comprises: A first prediction subunit is used to predict the images of the multiple modalities respectively to obtain a first prediction score corresponding to the image of each modality; A first weight subunit, configured to obtain a first weight corresponding to the image of each modality according to the evaluation result corresponding to each image of each modality; A second prediction subunit is used to predict the fused feature information to obtain a second prediction score corresponding to the fused feature information; a score fusion subunit, configured to perform score fusion on a plurality of the first prediction scores and the second prediction scores according to a first weight corresponding to each of the first prediction scores, to obtain a fusion score; Wherein, the score fusion module includes: An evaluation subunit, configured to obtain a second weight corresponding to the image of each modality according to an evaluation result corresponding to the image of each modality; The fusion subunit is used to fuse the plurality of feature information according to the second weight corresponding to the image of each modality to obtain fused feature information.

8. A computer storage medium, It is characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method as claimed in any one of claims 1 to 6.

9. A computer program product, It is characterized in that The computer program product stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method as claimed in any one of claims 1 to 6.

10. An electronic device, It is characterized in that include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Face living body detection method and device, storage medium and electronic equipment

    CN113033409A

  • Human body weight identification method and device based on bimodal feature fusion network

    CN114387612A