An image live body detection method, device, storage medium and electronic equipment
By performing attention operations and modal correlation fusion on multiple image features of the target object, high-quality third-object fusion features are generated, which solves the problems of detection accuracy and generalization of image liveness detection under environmental differences and achieves better liveness attack detection results.
Patent Information
- Application Number
- CN202211660755.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-12-23
AI Technical Summary
In existing technologies, image liveness detection methods are inaccurate when faced with environmental differences, have poor generalization effects, and are difficult to effectively resist attacks such as photos, face swaps, and masks.
By acquiring images of multiple object types of the target object, determining the modal features of each object type, performing feature attention operations and feature fusion, and combining the modal correlation fusion of basic modal features and reference modal features, a high-quality third object fusion feature is generated for detection.
It improves the generalization ability of image liveness detection, achieves better liveness attack detection results, and improves the accuracy and fineness of detection.
Smart Images

Figure CN116229585B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to an image living body detection method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the continuous development of object recognition systems in recent years, image living body detection has become an indispensable part of object recognition systems. Image living body detection needs to verify whether the image to be detected is a real living body object operation when collected, and needs to be able to effectively resist common attack means such as photos, face changing, masks, shielding and screen flipping, so as to effectively intercept non-living body type attack images, including mobile phone attacks, paper attacks, head models and the like. SUMMARY
[0003] The present specification provides an image living body detection method, device, storage medium and electronic device, and the technical solution is as follows:
[0004] In a first aspect, the present specification provides an image living body detection method, and the method comprises:
[0005] Obtaining at least two types of object images of a target object, and determining object modal features corresponding to each type of object image;
[0006] Performing feature attention operation processing on each object modal feature to obtain first object modal features corresponding to each object modal feature, and performing feature fusion based on each first object modal feature to obtain a first object fusion feature;
[0007] Obtaining at least one basic modal feature and at least one reference modal feature in each object modal feature, and performing modal correlation fusion processing based on the basic modal feature and the reference modal feature to obtain a second object fusion feature;
[0008] Performing screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and performing image living body detection processing on the target object based on the third object fusion feature.
[0009] In a second aspect, the present specification provides an image living body detection device, and the device comprises:
[0010] An image acquisition module is configured to obtain at least two types of object images of a target object, and determine object modal features corresponding to each type of object image;
[0011] The operation processing module is configured to perform feature attention operation processing on each object modal feature to obtain a first object modal feature corresponding to each object modal feature, and perform feature fusion based on each first object modal feature to obtain a first object fusion feature;
[0012] The operation processing module is configured to obtain at least one basic modal feature and at least one reference modal feature in each object modal feature, perform modal correlation fusion processing based on the basic modal feature and the reference modal feature to obtain a second object fusion feature;
[0013] The living body detection module is configured to perform screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and perform image living body detection processing on the target object based on the third object fusion feature.
[0014] In a third aspect, the present specification provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and performing the method steps of one or more embodiments of the present specification.
[0015] In a fourth aspect, the present specification provides an electronic device, which can include a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and performing the method steps of one or more embodiments of the present specification.
[0016] In a fifth aspect, the present specification provides a computer program product, which stores at least one instruction, and the instruction is suitable for being loaded by a processor and performing the method steps of one or more embodiments of the present specification.
[0017] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:
[0018] In one or more embodiments of the present specification, the electronic device determines the object modal features corresponding to the object images of different image modalities of the target object, on the one hand, performs attention operation processing on each object modal feature to obtain a first object modal feature, performs feature fusion based on each first object modal feature to obtain a first object fusion feature, on the other hand, performs modal correlation fusion based on the basic modal feature and the reference modal feature in the object modal feature to obtain a second object fusion feature, and then performs screening fusion based on the first object fusion feature and the second object fusion feature to obtain a high-quality and fine-grained third object fusion feature. The third object fusion feature has high separability, and makes full use of the image characteristics between different image modalities to make the image living body detection have good feature representation effect, which can assist subsequent accurate living body detection classification, achieve better living body attack detection effect, and improve the generalization ability of image living body detection. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a scene diagram of an image liveness detection system provided in this manual;
[0021] Figure 2 This is a flowchart illustrating an image liveness detection method provided in this specification;
[0022] Figure 3 This is a flowchart illustrating another image-based liveness detection method provided in this manual;
[0023] Figure 4 This is a schematic diagram illustrating the generation of a first object fusion feature in the image liveness detection method provided in this specification;
[0024] Figure 5 This is a flowchart illustrating a correlation fusion method for image liveness detection provided in this specification;
[0025] Figure 6 This is a schematic diagram of cross-attention processing provided in this specification;
[0026] Figure 7 This is a schematic diagram of an image liveness detection method provided in this specification;
[0027] Figure 8 This is a schematic diagram of the structure of an image liveness detection device provided in this specification;
[0028] Figure 9 This is a schematic diagram of the structure of an operation processing module provided in this manual;
[0029] Figure 10 This is a schematic diagram of the structure of an electronic device provided in this specification;
[0030] Figure 11 This is a schematic diagram of the operating system and user space provided in this manual;
[0031] Figure 12 yes Figure 11 Architecture diagram of the Android operating system in China;
[0032] Figure 13 yes Figure 11Architecture diagram of the IOS operating system. DETAILED DESCRIPTION
[0033] The technical solutions in the specification will be described clearly and completely in combination with the drawings in the specification. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0034] In the description of the present application, it should be understood that the terms "first", "second" and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance. In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units not listed, or optionally includes other steps or units inherent to the process, method, product or device. The specific meaning of the above terms in the present application can be understood by those skilled in the art according to the specific circumstances. In addition, in the description of the present application, "multiple" means two or more, unless otherwise specified. The association relationship between the associated objects is described, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0035] In the related art, "image liveness detection" needs to effectively intercept attack image data of non-liveness types, including mobile phone attacks, paper attacks, head model attacks, etc. However, in actual application, there are often large differences in the environment where "image liveness detection" is located, and the "image liveness detection" method in the related art has the problems of inaccurate actual application liveness detection results and poor generalization effect. Therefore, the "image liveness detection" in the related art has certain limitations, and one or more of the aforementioned limitations can be improved or even solved by performing the image liveness detection method in one or more embodiments of the specification.
[0036] The present application will be described in detail below in combination with specific embodiments.
[0037] Please refer to Figure 1 , a scene schematic diagram of an image liveness detection system provided in the specification. As Figure 1 shown, the image liveness detection system can at least include a client cluster and a service platform 100.
[0038] The client cluster can include at least one client, such as Figure 1 As shown, specifically including client 1 corresponding to user 1, client 2 corresponding to user 2, …, client n corresponding to user n, n is an integer greater than 0.
[0039] Each client in the client cluster can be an electronic device with communication function, including but not limited to: wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices, or other processing devices connected to wireless modems, etc. In different networks, electronic devices can be called different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in 5G network or future evolution network, etc.
[0040] The service platform 100 can be a separate server device, such as: rack-mounted, blade, tower, or cabinet server device, or using workstations, mainframe computers, etc. Hardware devices with strong computing power; It can also be a server cluster composed of multiple servers. Each server in the service cluster can be composed in a symmetrical manner, where each server is functionally equivalent and positionally equivalent in the transaction link. Each server can independently provide services externally. The independent service can be understood as not requiring the assistance of another server.
[0041] In one or more embodiments of the present specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and based on the communication connection, complete the interaction of data in the image living body detection process, such as online transaction data interaction, such as data interaction of at least two types of object images, illustratively, the client can collect at least two types of object images of the target object and send them to the service platform 100, and the service platform 100 performs the image living body detection method related in the present specification to perform image living body detection to obtain the living body detection category and feedback to the client; For example, the service platform 100 can distribute the related image living body detection model for image living body detection to several clients to instruct the client to perform the image living body detection method related in the present specification to perform image living body detection to obtain the living body detection category; For example, the service platform 100 can obtain training sample data for training the related image living body detection model from the client, such as living body detection sample images, etc.
[0042] Further, the step of "determining object modal feature corresponding to each of the object images, performing feature attention processing on each of the object modal features to obtain first object modal feature corresponding to each of the object modal features, performing feature fusion based on each of the first object modal features to obtain first object fusion feature, obtaining at least one basic modal feature and at least one reference modal feature in each of the object modal features, performing modal correlation fusion processing based on the basic modal feature and the reference modal feature to obtain second object fusion feature, performing screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain third object fusion feature, and performing image live body detection processing on the target object based on the third object fusion feature" can be performed by controlling the related image live body detection model.
[0043] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, where the network can be a wireless network or a wired network. The wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network. The wired network includes but is not limited to an Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML) and the like are used to represent data (such as the target compression package) exchanged through the network. In addition, all or some links can be encrypted using conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) and the like. In other embodiments, custom and / or dedicated data communication technologies can be used instead of or in addition to the above data communication technologies.
[0044] The image live body detection system provided by the specification and the image live body detection method in one or more embodiments belong to the same concept. The execution subject of the image live body detection method corresponding to the one or more embodiments of the specification can be the service platform 100 described above. The execution subject of the image live body detection method corresponding to the one or more embodiments of the specification can also be a client, which is determined based on the actual application environment. The implementation process of the image live body detection system embodiment can be seen from the method embodiment described below, which will not be described here.
[0045] Based on Figure 1 The following describes a model processing method provided by one or more embodiments of the present specification in detail based on the scenario schematic diagram shown in the figure.
[0046] Please refer to Figure 2 A flowchart of a model processing method provided by one or more embodiments of the present specification is shown, which can be implemented by a computer program and can run on an image living body detection device based on the Von Neumann system. The computer program can be integrated in an application or run as a standalone tool application. The model processing device can be an electronic device.
[0047] Specifically, the image living body detection method comprises:
[0048] S102: Obtain at least two types of object images of a target object, and determine object modal characteristics corresponding to each type of object image;
[0049] The object image can be understood as image data collected for a target object (such as a user, an animal, etc.) in an image living body detection scenario;
[0050] Each object image in the at least two types of object images has a different image modal type;
[0051] Further, the image living body detection scenario can be a scenario in actual application, such as finance, insurance, security, etc., which needs to detect the object image living body. For example, when a user performs user object verification, the user generally needs to upload an object image on an application program or website to verify whether it is operated by the user himself;
[0052] The object image can have different image modal types in actual application. The object image can be one or more of the following image modal types: video modal type, color picture modal (rgb) type, small video modal type, animation modal type, depth image (depth) modal type, infrared image (ir) modal type, etc.
[0053] In one or more embodiments of the present specification, the image modal type of the object image for image living body detection is at least two.
[0054] Illustratively, with the popularity and rapid development of electronic devices, electronic devices can support the collection of object images of several image models, and can collect color modal object images, infrared modal object images, and depth modal object images through the execution of a multi-modal collection component of a subject;
[0055] The object modal feature is an object modal feature obtained after image feature extraction is performed on an object image of a corresponding image modality, and the object modal feature can include one or more of modal features such as color features, shape features, depth features, texture features, and spatial relationship features, based on different image modalities;
[0056] It can be understood that the object modal features extracted from object images of different image modalities are different;
[0057] In one or more embodiments of the present specification, a feature conversion network can be used to convert each object image into a corresponding object modal feature. The object modal feature can be an image feature represented in the form of a modal feature map by mapping the object image to a high-dimensional vector space. The object modal feature can be represented by a character F in the vector space;
[0058] For example, an infrared modal object image of an infrared image modality (IR) type can determine a corresponding infrared object modal feature F IR ;
[0059] For example, a depth modal object image of a depth image modality (depth) type can determine a corresponding depth object modal feature F Depth ;
[0060] Optionally, the feature conversion network can be implemented using one or more of a VGG feature processing network, an AlexNet feature processing network, an Overfeat feature processing network, and a Resnet feature processing network in a classical feature conversion network. For example, in some embodiments, the feature processing network is based on an open-source Resnet feature processing network to perform feature processing on an input image to obtain an object modal feature;
[0061] In one or more embodiments of the present specification, object images of at least two types of different image modalities are used for image living body detection. The imaging characteristics of the corresponding image modalities under different image modalities can be fully utilized. Through subsequent corresponding processing, the separability characteristics brought by the extraction and fusion of several modalities are played, and the granularity, image quality, and other indicators of the modal image features are significantly improved, which is helpful for image living body detection to classify and identify images.
[0062] Optionally, determining the object modal features corresponding to each type of object image can be implemented based on a pre-trained image living body detection model. After obtaining at least two types of object images for a target object, the at least two types of object images are input into the image living body detection model, and the image living body detection method related in the present specification can be performed by the image living body detection model.
[0063] S104: performing feature attention operation processing on each of the object modal features to obtain a first object modal feature corresponding to each of the object modal features, and performing feature fusion based on the first object modal features to obtain a first object fusion feature;
[0064] The feature attention operation processing is used to adaptively focus on the object modal features of the corresponding image modal to obtain a first object modal feature that is more suitable for the characteristics of the corresponding image modal. The first object modal feature has richer modal characteristic representation than the original object modal feature before processing, sufficient interaction between features on the corresponding feature map channel, deeper feature granularity in the corresponding modal dimension, and higher feature quality. Furthermore, the feature attention operation processing on the object modal features of different object modalities will obtain corresponding first object modal features.
[0065] Illustratively, the feature attention processing can be implemented based on one or more of channel attention operation, spatial attention operation, and accumulation excitation attention operation.
[0066] Optionally, the channel attention operation processing can be performed on each of the object modal features to obtain an improved first object modal feature for each of the object modal features. The spatial attention operation processing can be performed on each of the object modal features to obtain an improved first object modal feature for each of the object modal features, and so on.
[0067] It can be understood that if the object modal features are i, then the first object modal features are i. The i first object modal features reflect the modal characteristics of the same target object in different modal dimensions and after attention adaptive focusing. After obtaining the first object modal features, the feature fusion is performed on the first object modal features of different image modal dimensions to obtain a first object fusion feature.
[0068] Illustratively, the feature fusion processing mode can be implemented using feature splicing operation, feature sampling operation, feature fitting operation, and the like.
[0069] In some embodiments, the feature fusion of the first object modal features of different image modal dimensions can be implemented based on a feature fusion network. The feature fusion network can be used to implement the feature fusion of the first object modal features to obtain a first object fusion feature.
[0070] Optionally, the feature fusion network can be part of a trained image living body detection model.
[0071] S106: Obtain at least one base modal feature and at least one reference modal feature in each of the object modal features, and perform modal correlation fusion processing based on the base modal feature and the reference modal feature to obtain a second object fusion feature;
[0072] The base modal feature can be a base image feature (or can be regarded as a main image feature) as a reference for modal feature fusion in the fusion process, and the reference modal feature is used as an auxiliary reference in the subsequent fusion process based on the base modal feature.
[0073] The reference modal feature can be understood as an auxiliary reference image feature (or can be regarded as an auxiliary image feature) determined based on the base modal feature.
[0074] In a feasible implementation, the obtaining of at least one base modal feature and at least one reference modal feature in each of the object modal features can be: selecting at least one base modal feature and at least one reference modal feature from each of the object modal features.
[0075] Illustratively, the priority of different image modalities can be set, and based on the priority of the image modalities corresponding to each of the object modal features, a target number of object modal features are selected as base modal features, and the features other than the base modal features are selected as reference modal features.
[0076] Illustratively, the target image modality that is predicted to be better can be determined in combination with device environment parameters (such as brightness parameters, environmental location parameters, and electromagnetic interference parameters), and the object modal feature corresponding to the target image modality is selected as the base modal feature, and the features other than the base modal feature are selected as the reference modal feature.
[0077] The base modal mapping relationship between a plurality of reference device environment parameters and the target image modalities corresponding thereto can be set in advance, the device environment parameters during image acquisition are obtained, and the target image modality corresponding to the device environment parameters is queried based on the base modal mapping relationship.
[0078] In a feasible implementation, the obtaining of at least one base modal feature and at least one reference modal feature in each of the object modal features can be: obtaining a preset at least one base modal feature and at least one reference modal feature from each of the object modal features.
[0079] Illustratively, the object modal feature corresponding to the specified base image modality is selected as the base modal feature, and the features other than the base modal feature are selected as the reference modal feature.
[0080] Specifically, then, the modal correlation fusion processing is performed based on the base modal feature and the reference modal feature to obtain a second object fusion feature corresponding to each of the object modal features;
[0081] The modal correlation fusion processing is performed by taking the base modal feature as a reference, measuring the correlation degree of the channel feature in the reference modal feature and the base modal feature, and performing attention region feature fusion based on the correlation degree by using an attention operation, so as to obtain the second object fusion feature.
[0082] S108: performing screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and performing image living body detection processing on the target object based on the third object fusion feature.
[0083] Illustratively, a feature importance screening network can be constructed and trained in advance based on a machine learning model, and the first object fusion feature and the second object fusion feature are screened and fused by the feature importance screening network to obtain a final fusion feature, that is, the third object fusion feature.
[0084] Optionally, the importance feature screening can be performed by using one or more of convolution operation, pooling operation, normalization operation, etc. in the feature importance screening network to perform feature enhancement and fusion.
[0085] It can be understood that the third object fusion feature fuses the image characteristics of multiple image modalities, and the third object fusion feature has higher feature quality and better feature granularity, so that the third object fusion feature used for image living body detection classification has higher separability, so as to distinguish the living body category and the attack category, and achieve better living body attack detection effect.
[0086] Further, the image living body detection processing is performed by using the third object fusion feature, and a living body detection category of the target object is output.
[0087] The living body detection category includes one of a living body image category and an attack image category.
[0088] Optionally, a living body detection classifier can be constructed and trained in advance based on a machine learning model, and the image living body detection processing is performed by using the living body detection classifier based on the third object fusion feature, and the living body detection category of the target object is output.
[0089] In one or more embodiments of the present specification, the electronic device determines object modal features corresponding to object images of different image modalities of a target object, on the one hand, performs attention operation processing on each object modal feature to obtain first object modal features, performs feature fusion based on each first object modal feature to obtain first object fusion features, on the other hand, performs modal correlation fusion based on a basic modal feature and a reference modal feature in the object modal features to obtain second object fusion features, and then performs screening fusion based on the first object fusion features and the second object fusion features to obtain third object fusion features with high quality and fine granularity. The third object fusion features have high separability, and make full use of the image characteristics between different image modalities to make the image liveness detection have good feature representation effect, which can assist subsequent accurate liveness detection classification, achieve better liveness attack detection effect, and improve the generalization ability of image liveness detection.
[0090] Please refer to Figure 3 , Figure 3 is a flowchart of another embodiment of an image liveness detection method proposed in one or more embodiments of the present specification. Specifically:
[0091] S202: Obtain at least two types of object images for a target object, and determine object modal features corresponding to each type of object image;
[0092] For details, refer to the method steps of other embodiments of the present specification, which will not be repeated here.
[0093] S204: Perform channel attention operation processing on each object modal feature to obtain first object modal features corresponding to each object modal feature, perform feature concatenation processing on each first object modal feature to obtain a concatenation processing feature, and perform convolution aggregation processing on the concatenation processing feature to obtain first object fusion features.
[0094] The first object modal feature is a modal feature obtained by performing channel attention operation processing on a plurality of object modal features.
[0095] It can be understood that the channel attention mechanism can be used to perform channel attention operation on each object modal feature to obtain the first object modal feature. Generally, the channel attention mechanism focuses on the importance of the feature map channel corresponding to each object modal feature, and assigns feature weights based on the importance to perform feature fusion.
[0096] In a feasible implementation, the channel attention operation processing on each object modal feature to obtain the first object modal feature corresponding to each object modal feature can include the following solutions:
[0097] The electronic device can determine channel importance of each object modal feature in at least one image feature channel dimension through a squeeze-and-excitation processing network, and perform attention operation processing on the object modal feature in the image feature channel dimension based on the channel importance, to obtain a first object modal feature corresponding to each object modal feature.
[0098] Illustratively, the channel attention operation processing can be implemented based on a squeeze-and-excitation processing network. The squeeze-and-excitation processing network generally includes two parts, a squeeze part and a compression part. The squeeze part compresses global spatial information of a feature map corresponding to the object modal feature, and then performs feature learning in the feature map channel dimension to form the importance of each channel, i.e., the channel importance. Finally, the excitation part assigns different weights to each feature map channel to generate the first object modal feature.
[0099] Specifically, after determining a plurality of first object modal features, feature splicing processing is performed on each first object modal feature using a feature splicing operation to obtain spliced processing features, and convolution aggregation processing is performed on the spliced processing features to obtain a first object fusion feature.
[0100] Illustratively, as shown in Figure 4 , Figure 4 is a schematic diagram related to generation of a first object fusion feature. It is assumed that the dimension of a feature map corresponding to n object modal features (f1, f2,..., fn, respectively) can be represented as HxWxC, where H is the height of the feature map, W is the width of the feature map, and C represents the number of feature map channels. The squeeze part can compress the global spatial information of the feature map HxWxC corresponding to the object modal feature using a pooling operation, for example, to compress the dimension from HxWxC to 1x1xC (commonly represented by a 1x1xC weight tensor). Then, in the excitation part, the 1x1xC weight tensor is predicted through a fully connected layer to obtain the channel importance of each image feature channel dimension, and then the channel importance is excited to the original feature map HxWxC corresponding image feature channel to perform channel fusion operation, thereby obtaining the first object modal feature. It can be understood that a plurality of first object modal features (F1, F2,..., Fn) can be obtained for a plurality of object modal features through the foregoing steps. Then, feature splicing processing is performed on each first object modal feature (F1, F2,..., Fn) using a feature splicing operation to obtain spliced processing features, and convolution aggregation processing (such as Figure 4 conv processing) is performed on the spliced processing features to obtain a first object fusion feature (which can be represented as F x1 ).
[0101] S206: Obtain at least one base modal feature and at least one reference modal feature in each object modal feature.
[0102] In an implementable embodiment, at least one base modal feature and at least one reference modal feature can be selected from each object modal feature;
[0103] In an implementable embodiment, at least one base modal feature and at least one reference modal feature can be selected from each object modal feature;
[0104] S208: performing modal correlation fusion processing based on the base modal feature and the reference modal feature to obtain a second object fusion feature;
[0105] Specifically, the cross-attention processing network can be used to determine the modal correlation information between the base modal feature and the reference modal feature from the modal correlation dimension;
[0106] Specifically, the modal correlation information is obtained by measuring the correlation between the channel feature in the reference modal feature and the base modal feature, with the base modal feature as the reference.
[0107] Specifically, the base modal feature and the reference modal feature can be fused based on the modal correlation information to obtain a second object fusion feature, and the object modal feature includes at least one base modal feature and at least one reference modal feature.
[0108] Specifically, the cross-attention processing network can be a neural network module trained based on a machine learning model in advance. The cross-attention processing network can perform cross-attention operation on other reference modal features with the base modal feature having the highest separability as the base modal. The cross-attention mechanism (CA) is used to calculate the direct correlation between the base modal feature and the reference modal feature, so as to obtain the modal correlation information. In some embodiments, the modal correlation information is represented in the form of a relationship mapping vector. After normalization based on the modal correlation information, the base modal feature and the reference modal feature can be fused to obtain a second object fusion feature.
[0109] For example, as shown in Figure 5 , Figure 5 is a flowchart of correlation fusion. The base modal feature and the reference modal feature can be fused based on the modal correlation information to obtain a second object fusion feature, which can be:
[0110] A2: determining the modal correlation information between the base modal feature and the reference modal feature from the modal correlation dimension by using the cross-attention processing network, and performing dot product processing on the base modal feature and each reference modal feature based on the modal correlation information to obtain at least one cross-attention modal feature;
[0111] Understandably, modal correlation information is represented in the form of interaction relationship mapping vectors;
[0112] like Figure 6 As shown, Figure 6 This is a diagram involving cross-attention processing. Figure 6 In this context, the basic modal feature can be represented as Fa, and there are usually multiple reference modal features, which can be represented as F. b1 F b2 ..., F bn Where n is a positive integer, the cross-attention processing network processes the basic modality feature Fa and the reference modality feature F. b1 F b2 ..., F bn Cross-attention processing network cross-attention mechanism (e.g.) Figure 6 The CA part shown is used to calculate the fundamental modal feature Fa and the reference modal feature F. b1 F b2 ..., F bn The modal correlation information between them can usually be represented as an interaction relationship mapping vector. Then, based on each interaction relationship mapping vector, after normalization, it is multiplied with the original reference modal features (i.e., dot product operation) to form at least one cross-attention modal feature Ft1, Ft2. t2 ..., Ft n ;
[0113] A4: The basic modal features and each of the cross-attention modal features are summed to obtain the cross-attention fusion features;
[0114] Indicatively, during the cross-attention processing stage of a network, such as Figure 6 As shown, at least one cross-attention modal feature Ft1, F t2 ..., Ft n The original basic modal features Fa are superimposed and summed to obtain several reference modalities and basic modalities after cross-fusion. The fused cross-attention features carry the basic image modal features and cross-attention characteristics.
[0115] A6: Perform convolution processing on the cross-attention fusion features to obtain the second object fusion features.
[0116] Schematic, then the multiple cross-attention fusion features are convolutionally processed through convolutional layers (e.g.) Figure 6 The conv processing shown above yields the second object fusion feature, and the feature map resolution of the second object fusion feature is consistent with the feature map resolution of the original object modality feature.
[0117] Illustratively, taking the at least two types of object images as color type object images, infrared type object images, and depth type object images as an example, taking the determined basic modal feature as a color modal feature, and taking the reference modal features as infrared modal features and depth modal features as an example for interpretation, as follows,
[0118] The point multiplication processing of the basic modal feature and each reference modal feature based on the modal correlation information to obtain at least one cross-attention modal feature, and the sum processing of the basic modal feature and each cross-attention modal feature to obtain a cross-attention fusion feature can include the following scheme:
[0119] Specifically, the color modal feature is multiplied with the infrared modal feature and the depth modal feature based on the modal correlation information to obtain cross-attention color-infrared modal features and cross-attention color-depth modal features.
[0120] Specifically, the color modal feature, the cross-attention color-infrared modal feature, and the cross-attention color-depth modal feature are summed to obtain a cross-attention fusion feature.
[0121] S210: performing screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and performing image living body detection processing on the target object based on the third object fusion feature.
[0122] Illustratively, as shown in Figure 7 , Figure 7 is a schematic diagram of image living body detection involved in the present specification, in Figure 7 , the channel importance of each object modal feature in at least one image feature channel dimension is determined from the image feature channel dimension by the squeezing excitation processing network, and the object modal feature is subjected to attention operation processing in the image feature channel dimension based on the channel importance, to obtain a first object modal feature F x1; and determining, by a cross-attention processing network, modal correlation information of the base modal feature and the reference modal feature from a modal correlation dimension, performing dot product processing on the base modal feature and each of the reference modal features based on the modal correlation information to obtain at least one cross-attention modal feature, and performing convolution processing on the at least one cross-attention modal feature to obtain a second object fusion feature; then performing region screening and fusion based on the first object fusion feature and the second object fusion feature by using a feature importance screening network to obtain a final third object fusion feature, and then performing image live body detection processing based on the third object fusion feature by using a live body detection classification network to output a live body detection category of the target object, the live body detection category including one of an image live body category and an image attack category;
[0123] The first object fusion feature and the second object fusion feature are usually in the form of an object fusion feature map.
[0124] In a feasible implementation, the feature importance screening network can be used to perform region screening processing on the first object fusion feature and the second object fusion feature to obtain a plurality of object region features, and the plurality of object region features are fused to obtain a third object fusion feature.
[0125] Illustratively, the region screening processing can be a feature block-based screening mechanism. The feature block-based screening mechanism can screen a plurality of object region features with good separability from the first object fusion feature and the second object fusion feature, and then fuse the plurality of object region features to obtain the third object fusion feature.
[0126] Illustratively, the feature block-based screening mechanism can be a neural network module constructed and trained based on a machine learning model.
[0127] It should be noted that the machine learning model involved in one or more embodiments of the present specification includes but is not limited to one or more of fitting of a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN), an embedding model, a gradient boosting decision tree (GBDT) model, a logistic regression (LR) model, and the like.
[0128] In one or more embodiments of the present specification, an initial image liveness detection model and at least one network module included in the initial image liveness detection model, such as a squeeze excitation processing network, a cross-attention processing network, a feature importance screening network, a liveness detection classification network, and the like, can be pre-constructed based on a machine learning model. A large amount of sample object image data is obtained, the sample object image data being at least two types of sample object image data for a sample object. Then, the sample object image data is used to train the initial image liveness detection model. After the initial image liveness detection model satisfies a model end training condition, a trained image liveness detection model can be obtained. In the model training stage, the image liveness detection method can refer to the image liveness detection method of one or more embodiments of the present specification, which will not be described here.
[0129] In one or more embodiments of the present specification, the model end training condition can include, for example, a loss function value less than or equal to a preset loss function threshold, an iteration number reaching a preset number threshold, and the like. The specific model end training condition can be determined based on actual conditions, which will not be specifically limited here.
[0130] In one or more embodiments of the present specification, the image liveness detection method can be understood as a multi-modal liveness attack detection method based on cross-attention and hybrid feature fusion. The image liveness detection collects original data of image modalities such as rgb modalities, ir modalities, and depth modalities through multi-modal acquisition, and then extracts features. On the one hand, the feature extraction based on the squeeze excitation processing network and the feature concatenation operation can effectively extract original multi-modal features and fused features. On the other hand, the feature extraction based on cross-attention and hybrid feature fusion can obtain high-quality and fine-grained deep third object fusion features. The third object fusion features have high separability, and fully utilize the image characteristics between different image modalities, so that the image liveness detection can have good feature representation effect, which can assist subsequent accurate liveness detection classification, achieve better liveness attack detection effect, and improve the generalization ability of the image liveness detection.
[0131] The following will be described in detail with reference to the accompanying drawings. Figure 8 The image liveness detection device provided in the present specification will be described in detail. It should be noted that, Figure 8 The image liveness detection device shown in the figure is used to execute the method of the present application Figures 1-7 The method of the embodiment shown in the figure is only shown with parts related to the present specification, and the specific technical details not disclosed will be described with reference to the method of the embodiment shown in the figure. Figures 1-7 The embodiment shown in the figure.
[0132] Please refer to Figure 8Fig. 1 is a structural schematic diagram of an image living body detection device according to an embodiment of the present application. The image living body detection device 1 can be implemented by software, hardware or a combination of both to be all or part of a user terminal. According to some embodiments, the image living body detection device 1 comprises an image living body detection module 11, an operation processing module 12 and a living body detection module 13, which are specifically used for:
[0133] The image acquisition module 11 is configured to acquire at least two types of object images of a target object, and determine object modality features corresponding to each type of object image.
[0134] The operation processing module 12 is configured to perform feature attention operation processing on each object modality feature to obtain a first object modality feature corresponding to each object modality feature, and perform feature fusion based on each first object modality feature to obtain a first object fusion feature.
[0135] The operation processing module 12 is configured to acquire at least one basic modality feature and at least one reference modality feature in each object modality feature, and perform modality correlation fusion processing based on the basic modality feature and the reference modality feature to obtain a second object fusion feature.
[0136] The living body detection module 13 is configured to perform screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and perform image living body detection processing on the target object based on the third object fusion feature.
[0137] Optionally, the operation processing module 12 is configured to:
[0138] perform channel attention operation processing on each object modality feature to obtain a first object modality feature corresponding to each object modality feature.
[0139] Optionally, the operation processing module 12 is configured to:
[0140] determine a channel importance degree of each object modality feature in at least one image feature channel dimension through a squeeze-and-excitation processing network, and perform attention operation processing on the object modality feature in the image feature channel dimension based on the channel importance degree to obtain a first object modality feature corresponding to each object modality feature.
[0141] Optionally, the operation processing module 12 is configured to perform feature splicing processing on each first object modality feature to obtain splicing processing features, and perform convolution aggregation processing on the splicing processing features to obtain a first object fusion feature.
[0142] Optionally, the operation processing module 12 is configured to: select at least one base modal feature and at least one reference modal feature from the object modal features; or,
[0143] acquire at least one base modal feature and at least one reference modal feature from the object modal features.
[0144] Optionally, as shown in Figure 9 the operation processing module 12 comprises:
[0145] a correlation processing unit 121 configured to determine modal correlation information of the base modal feature and the reference modal feature from the modal correlation dimension by using a cross-attention processing network;
[0146] a feature fusion unit 122 configured to perform feature fusion on the base modal feature and the reference modal feature based on the modal correlation information to obtain a second object fusion feature, the object modal feature comprising at least one base modal feature and at least one reference modal feature
[0147] Optionally, the feature fusion unit 122 is configured to:
[0148] perform point multiplication processing on the base modal feature and each of the reference modal features based on the modal correlation information to obtain at least one cross-attention modal feature;
[0149] perform addition processing on the base modal feature and each of the cross-attention modal features to obtain a cross-attention fusion feature;
[0150] perform convolution processing on the cross-attention fusion feature to obtain a second object fusion feature.
[0151] Optionally, the at least two types of object images are color object images, infrared object images, and depth object images, the base modal feature is a color modal feature, the reference modal feature is an infrared modal feature and a depth modal feature, and the feature fusion unit 122 is configured to:
[0152] perform point multiplication processing on the color modal feature and the infrared modal feature and the depth modal feature based on the modal correlation information to obtain a cross-attention color-infrared modal feature and a cross-attention color-depth modal feature;
[0153] perform addition processing on the color modal feature, the cross-attention color-infrared modal feature, and the cross-attention color-depth modal feature to obtain a cross-attention fusion feature.
[0154] Optionally, the living body detection module 13 is configured to:
[0155] The first object fusion feature and the second object fusion feature are processed by a feature importance filtering network to obtain multiple object region features;
[0156] The features of each object region are fused to obtain a third object fusion feature.
[0157] Optionally, the device 1 is used for:
[0158] The image liveness detection is performed by a liveness detection classification network based on the fusion features of the third object, and the liveness detection category for the target object is output. The liveness detection category includes either an image liveness category or an image attack category.
[0159] It should be noted that the image liveness detection device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the image liveness detection method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image liveness detection device and the image liveness detection method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0160] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0161] In one or more embodiments of this specification, the electronic device determines the object modality features corresponding to object images of different image modalities of the target object. On the one hand, it performs attention operation processing on each object modality feature to obtain a first object modality feature, and performs feature fusion based on each first object modality feature to obtain a first object fusion feature. On the other hand, it performs modality correlation fusion based on the basic modality feature and reference modality feature in the object modality feature to obtain a second object fusion feature. Then, it performs screening and fusion based on the first object fusion feature and the second object fusion feature to obtain a high-quality and fine-grained third object fusion feature. The third object fusion feature has high separability and makes full use of the image characteristics between different image modalities, so that image liveness detection can have a good feature representation effect, which can assist subsequent accurate liveness detection classification, achieve better liveness attack detection effect, and improve the generalization ability of image liveness detection.
[0162] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-7 The image liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-7The specific description of the illustrated embodiments will not be repeated here.
[0163] The present application also provides a computer program product, which stores at least one instruction, the at least one instruction is loaded and executed by the processor to perform the method as described above Figures 1-7 The specific execution process of the image living body detection method of the illustrated embodiments can be referred to Figures 1-7 The specific description of the illustrated embodiments will not be repeated here.
[0164] Please refer to Figure 10 It shows the structure block diagram of the electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140 and a bus 150. The processor 110, the memory 120, the input device 130 and the output device 140 can be connected through the bus 150.
[0165] The processor 110 can include one or more processing cores. The processor 110 connects various parts in the entire electronic device by using various interfaces and lines, executes various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be realized in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), programmable logic array (PLA). The processor 110 can integrate a combination of one or several of central processing unit (CPU), graphics processing unit (GPU) and modem. Among them, the CPU mainly processes operating system, user interface and application program, etc.; the GPU is responsible for rendering and drawing display content; the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but be realized by a separate communication chip.
[0166] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing each of the following method embodiments, etc., and the operating system can be an Android system, an IOS system developed by Apple Inc., a system developed based on the Android system or the IOS system, or other systems. The data storage area can also store data created by the electronic device in use, such as a phonebook, audio and video data, chat record data, etc.
[0167] Referring to Figure 11 As shown, the memory 120 can be divided into an operating system space and a user space, and the operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good running effects, the operating system allocates corresponding system resources to different third-party applications. However, there are also differences in the demand for system resources in different application scenarios in the same third-party application, for example, in the local resource loading scenario, the third-party application has a higher requirement for the disk reading speed, and in the animation rendering scenario, the third-party application has a higher requirement for the GPU performance. However, the operating system and the third-party application are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application, resulting in that the operating system cannot perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0168] In order to enable the operating system to distinguish the specific application scenario of the third-party application, it is necessary to open up the data communication between the third-party application and the operating system, so that the operating system can obtain the current scenario information of the third-party application at any time, and then perform targeted system resource adaptation based on the current scenario.
[0169] Taking the operating system as an Android system for example, the programs and data stored in the memory 120 are as follows Figure 12As shown, the memory 120 can store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380, wherein the Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to an operating system space, and the application layer 380 belongs to a user space. The Linux kernel layer 320 provides underlying drivers for various hardware of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, and the like. The system runtime library layer 340 provides main feature support for the Android system through some C / C++ libraries. For example, an SQLite library provides database support, an OpenGL / ES library provides 3D drawing support, a Webkit library provides browser kernel support, and the like. An Android runtime is also provided in the system runtime library layer 340, which mainly provides some core libraries to allow developers to use the Java language to write Android applications. The application framework layer 360 provides various APIs that can be used when building an application, and developers can also build their own applications by using these APIs, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application program is running in the application layer 380, which can be native applications provided by the operating system, such as a contact program, a message program, a clock program, a camera application, and the like, or third-party applications developed by third-party developers, such as game applications, instant messaging programs, photo beautification programs, and the like.
[0170] For example, taking an IOS system as the operating system, the programs and data stored in the memory 120 can include an IOS kernel 320, an IOS runtime library 340, an application framework 360, and an application 380. Figure 13As shown, the IOS system includes: a core operating system layer 420, a core service layer 440, a media layer 460, and a Cocoa Touch layer 480. The core operating system layer 420 includes an operating system kernel, drivers, and low-level hardware abstractions that provide more hardware-specific functionality to program frameworks in the core service layer 440. The core service layer 440 provides system services and / or program frameworks that applications need, such as a Foundation framework, an account framework, an advertisement framework, a data storage framework, a network connection framework, a geographic location framework, a motion framework, and the like. The media layer 460 provides interfaces for applications related to audio and video, such as interfaces related to graphics images, interfaces related to audio technology, interfaces related to video technology, an AirPlay interface for wireless audio and video transmission technology, and the like. The Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development, and is responsible for user touch interaction on the electronic device. For example, a local notification service, a remote push service, an advertisement framework, a game tool framework, a message user interface (UI) framework, a user interface UIKit framework, a map framework, and the like.
[0171] In Figure 13 In the framework shown, the frameworks related to most applications include, but are not limited to, the Foundation framework in the core service layer 440 and the UIKit framework in the Cocoa Touch layer 480. The Foundation framework provides many basic object classes and data types, and provides the most basic system services for all applications, and is UI-independent. The UIKit framework provides basic UI class libraries for creating touch-based user interfaces, and iOS applications can provide UIs based on the UIKit framework, so it provides the basic framework of the application for building user interfaces, drawing, processing and user interaction events, responding to gestures, and the like.
[0172] In the IOS system, the manner and principle of implementing data communication between a third-party application and an operating system can refer to the Android system, and will not be described herein.
[0173] The input device 130 is configured to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is configured to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In an example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen configured to receive a touch operation of a user using a finger, a stylus, or any suitable object on or near the touch display screen, and display a user interface of each application. The touch display screen is usually arranged on a front panel of the electronic device. The touch display screen can be designed as a full screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, which is not limited in the present specification.
[0174] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above-described drawings does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown in the drawings, or combine certain components, or different component arrangements. For example, the electronic device further includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module, and the like, which are not described herein.
[0175] In the present specification, the execution subject of each step can be the electronic device described above. Alternatively, the execution subject of each step is an operating system of the electronic device. The operating system can be an Android system, an IOS system, or other operating systems, which are not limited in the present specification.
[0176] The electronic device of the present specification can further have a display device installed thereon. The display device can be various devices capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), and the like. A user can use the display device on the electronic device 101 to view displayed text, images, videos, and the like. The electronic device can be a smart phone, a tablet computer, a game device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook computer, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, an electronic clothing, and the like.
[0177] In Figure 10 In the electronic device shown, the processor 110 can be configured to invoke an application stored in the memory 120, and specifically perform the following operations:
[0178] Obtaining at least two types of object images for a target object, determining object modal characteristics corresponding to each type of object image;
[0179] Performing feature attention processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic, and performing feature fusion based on each first object modal characteristic to obtain a first object fusion feature;
[0180] Obtaining at least one basic modal characteristic and at least one reference modal characteristic in each object modal characteristic, and performing modal correlation fusion processing based on the basic modal characteristic and the reference modal characteristic to obtain a second object fusion feature;
[0181] Performing screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and performing image living body detection processing on the target object based on the third object fusion feature.
[0182] In one embodiment, the processor 110, in performing the feature attention processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic, further performs the following operations:
[0183] Performing channel attention operation processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic.
[0184] In one embodiment, the processor 110, in performing the channel attention operation processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic, performs the following steps:
[0185] Determining the channel importance of each object modal characteristic in at least one image feature channel dimension through extrusion excitation processing network, to perform attention operation processing on the object modal characteristic in the image feature channel dimension based on the channel importance, to obtain a first object modal characteristic corresponding to each object modal characteristic.
[0186] In one embodiment, the processor 110, in performing the feature fusion based on each first object modal characteristic to obtain a first object fusion feature, performs the following steps:
[0187] The first object modal feature is subjected to feature splicing processing to obtain splicing processing features, and the splicing processing features are subjected to convolution aggregation processing to obtain first object fusion features.
[0188] In one embodiment, the processor 110 performs the following steps in performing the obtaining of at least one base modal feature and at least one reference modal feature from each of the object modal features:
[0189] selecting at least one base modal feature and at least one reference modal feature from each of the object modal features; or,
[0190] obtaining a preset at least one base modal feature and at least one reference modal feature from each of the object modal features.
[0191] In one embodiment, the processor 110 performs the following steps in performing the obtaining of second object fusion features based on the base modal features and the reference modal features:
[0192] determining modal correlation information of the base modal features and the reference modal features from a modal correlation dimension through a cross-attention processing network;
[0193] performing feature fusion of the base modal features and the reference modal features based on the modal correlation information to obtain second object fusion features, the object modal features including at least one base modal feature and at least one reference modal feature
[0194] In one embodiment, the processor 110 performs the following steps in performing the obtaining of second object fusion features based on the modal correlation information:
[0195] performing dot product processing of the base modal features and each of the reference modal features based on the modal correlation information to obtain at least one cross-attention modal feature;
[0196] performing addition processing of the base modal features and each of the cross-attention modal features to obtain cross-attention fusion features;
[0197] performing convolution processing of the cross-attention fusion features to obtain second object fusion features.
[0198] In an embodiment, the at least two types of object images are color object images, infrared object images, and depth object images, the base modality feature is a color modality feature, the reference modality features are an infrared modality feature and a depth modality feature, and the processor 110 performs the following steps in performing the point multiplication of the base modality feature with each of the reference modality features based on the modality correlation information to obtain at least one cross-attention modality feature, and the summing of the base modality feature and each of the cross-attention modality features to obtain a cross-attention fusion feature:
[0199] point multiplication of the color modality feature with the infrared modality feature and the depth modality feature based on the modality correlation information to obtain a cross-attention color-infrared modality feature and a cross-attention color-depth modality feature;
[0200] summing of the color modality feature, the cross-attention color-infrared modality feature, and the cross-attention color-depth modality feature to obtain a cross-attention fusion feature.
[0201] In an embodiment, the processor 110 performs the following steps in performing the screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature:
[0202] region screening processing of the first object fusion feature and the second object fusion feature by a feature importance screening network to obtain a plurality of object region features;
[0203] fusion of each of the object region features to obtain a third object fusion feature.
[0204] In an embodiment, the processor 110 performs the following steps in performing the image live body detection processing on the target object based on the third object fusion feature:
[0205] image live body detection processing based on the third object fusion feature by a live body detection classification network to output a live body detection category for the target object, the live body detection category including one of an image live body category and an image attack category.
[0206] In one or more embodiments of the present specification, the electronic device determines the object modal features corresponding to the object images of different image modalities of the target object, on the one hand, performs attention operation processing on each object modal feature to obtain a first object modal feature, and performs feature fusion based on each first object modal feature to obtain a first object fusion feature, on the other hand, performs modal correlation fusion on the basis modal feature and the reference modal feature in the object modal feature to obtain a second object fusion feature, and then performs screening fusion based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature with high quality and fine granularity, the third object fusion feature has high separability, and fully utilizes the image characteristics between different image modalities to make the image liveness detection have good feature representation effect, which can assist subsequent accurate liveness detection classification, achieve better liveness attack detection effect, and improve the generalization ability of image liveness detection.
[0207] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, the processes of the above-mentioned embodiments can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory, or a random access memory.
[0208] The above only discloses the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application, so equivalent changes made according to the claims of the present application still fall within the scope of the present application.
Claims
1. An image living body detection method, the method comprising: acquiring at least two types of object images of a target object, and determining object modal characteristics corresponding to each type of object image; performing feature attention processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic, and performing feature fusion based on each first object modal characteristic to obtain a first object fusion feature; pre-setting priorities of different image modalities, selecting a target number of object modal characteristics as a basic modal feature based on the order of the priorities of the image modalities corresponding to each object modal characteristic, and selecting features other than the basic modal feature as reference modal features, or determining a target image modality from each object modal characteristic based on device environment parameters during image acquisition, and selecting an object modal feature corresponding to the target image modality as a basic modal feature, and selecting features other than the basic modal feature as reference modal features; determining modal correlation information of the basic modal feature and the reference modal feature from a modal correlation dimension through a cross-attention processing network, and performing feature fusion on the basic modal feature and the reference modal feature based on the modal correlation information to obtain a second object fusion feature; performing screening and fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and performing image living body detection processing on the target object based on the third object fusion feature.
2. The method of claim 1, wherein the feature attention processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic comprises: performing channel attention operation processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic.
3. The method of claim 2, wherein the channel attention operation processing on each object modal characteristic to obtain a first object modal characteristic corresponding to each object modal characteristic comprises: determining the channel importance of each object modal characteristic in at least one image feature channel dimension through a squeeze excitation processing network, and performing attention operation processing on the object modal characteristic in the image feature channel dimension based on the channel importance to obtain a first object modal characteristic corresponding to each object modal characteristic.
4. The method of claim 1, wherein the feature fusion based on each first object modal characteristic to obtain a first object fusion feature comprises: performing feature splicing processing on each first object modal characteristic to obtain spliced processing features, and performing convolution aggregation processing on the spliced processing features to obtain a first object fusion feature.
5. The method of claim 1, determining a target image modality from each of the object modality features based on the device environmental parameters at the time of image capture, comprising: pre-setting a basic modal mapping relationship between a plurality of reference device environment parameters and target image modalities corresponding thereto, and obtaining device environment parameters during image acquisition, and querying a target image modality corresponding to the device environment parameters based on the basic modal mapping relationship.
6. The method of claim 1, wherein the feature fusion of the base modal feature and the reference modal feature based on the modal correlation information to obtain a second object fusion feature comprises: point-wise multiplication processing of the base modal feature with each of the reference modal features based on the modal correlation information to obtain at least one cross-attention modal feature; sum processing of the base modal feature and each of the cross-attention modal features to obtain a cross-attention fusion feature; convolution processing of the cross-attention fusion feature to obtain the second object fusion feature.
7. The method of claim 6, wherein the at least two types of object images are color object images, infrared object images, and depth object images, the base modal feature is a color modal feature, and the reference modal features are infrared modal features and depth modal features, the point-wise multiplication processing of the base modal feature with each of the reference modal features based on the modal correlation information to obtain at least one cross-attention modal feature, and the sum processing of the base modal feature and each of the cross-attention modal features to obtain a cross-attention fusion feature, comprise: point-wise multiplication processing of the color modal feature with the infrared modal feature and the depth modal feature based on the modal correlation information to obtain cross-attention color-infrared modal features and cross-attention color-depth modal features; sum processing of the color modal feature, the cross-attention color-infrared modal features, and the cross-attention color-depth modal features to obtain a cross-attention fusion feature.
8. The method of claim 1, wherein the screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature comprises: region screening processing of the first object fusion feature and the second object fusion feature by a feature importance screening network to obtain a plurality of object region features; fusion of each of the object region features to obtain the third object fusion feature.
9. The method of claim 1, wherein the image live body detection processing based on the third object fusion feature on the target object comprises: image live body detection processing based on the third object fusion feature by a live body detection classification network to output a live body detection category for the target object, the live body detection category comprising one of an image live body category and an image attack category.
10. An image live body detection device, the device comprising: an image acquisition module configured to acquire at least two types of object images for a target object and determine object modal features corresponding to each type of the object images; an operation processing module configured to perform feature attention operation processing on each of the object modal features to obtain a first object modal feature corresponding to each of the object modal features, and perform feature fusion based on each of the first object modal features to obtain a first object fusion feature. The operation processing module is configured to set priorities of different image modalities, select object modal feature of a target number as a basic modal feature based on a high-low order of priorities of image modalities corresponding to each object modal feature, and select features other than the basic modal feature as reference modal features; or determine a target image modality from each object modal feature based on a device environment parameter during image acquisition, select object modal feature corresponding to the target image modality as the basic modal feature, and select features other than the basic modal feature as the reference modal features; determine modal correlation information of the basic modal feature and the reference modal features from a modal correlation dimension through a cross-attention processing network; and perform feature fusion on the basic modal feature and the reference modal features based on the modal correlation information to obtain a second object fusion feature. The living body detection module is configured to perform screening fusion processing based on the first object fusion feature and the second object fusion feature to obtain a third object fusion feature, and perform image living body detection processing on the target object based on the third object fusion feature. 11.A computer storage medium, the computer storage medium storing a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the method steps of any one of claims 1-9. 12.A computer program product, the computer program product storing at least one instruction, the at least one instruction being adapted to be loaded and executed by a processor to perform the method steps of any one of claims 1-9.
13. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method steps of any one of claims 1-9.
Citation Information
Patent Citations
RGB-D image saliency target detection method and device, equipment and storage medium
CN111967477A
Face living body detection method and device, electronic device and storage medium
CN113435408A
Illegal image recognition method, system and equipment
CN114140673A
Living body detection method and device, electronic equipment, storage medium and program product
CN114743277A