Data processing method, device, computer readable storage medium and processor
By generating a vital body recognition model, using texture feature maps and high-level semantic feature maps for vital body recognition, the problem of inaccurate vital body recognition results in the prior art is solved, and higher recognition accuracy and security are achieved.
Patent Information
- Application Number
- CN202011467170.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-14
AI Technical Summary
In the prior art, when performing live body recognition, there is a problem that the live body recognition results are inaccurate, especially when facing non-living attack samples, it is difficult to accurately identify them.
By obtaining the texture feature map and high-level semantic feature map of the sample picture, it is stitched into the target feature map, and input it into the full connection layer of the neural network model for training to generate a living body recognition model. This model is used to extract texture features and high-level semantic features of verification images for live recognition.
It effectively improves the accuracy of live body recognition, can accurately refuse to identify non-living attack samples, solves the problem of inaccurate live body recognition results, and enhances the security of the face recognition system.
Smart Images

Figure CN114627518B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of living body recognition, and in particular to a data processing method, device, computer-readable storage medium and processor. Background Art
[0002] At present, liveness recognition is an important means of distinguishing personal identity, but the equipment used for liveness recognition faces challenges from non-liveness attack samples. Criminals can attack the recognition system by holding photos or video clips of relevant persons, thereby disguising the identity of the person and passing the face recognition system. Therefore, non-liveness attacks are a major security risk of the current face recognition system, and there is a technical problem that the liveness recognition results are inaccurate when performing liveness recognition.
[0003] With respect to the above-mentioned technical problem of inaccurate liveness recognition results during liveness recognition, no effective solution has been proposed yet. Summary of the invention
[0004] Embodiments of the present invention provide a data processing method, device, computer-readable storage medium, and processor to at least solve the technical problem of inaccurate liveness recognition results during liveness recognition.
[0005] According to one aspect of an embodiment of the present invention, a data processing method is provided. The method may include: obtaining a sample image, wherein the sample image is any one of multiple images captured by a camera; extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0006] According to one aspect of an embodiment of the present invention, a method for identifying a living body is provided. The method may include: obtaining a target image to be identified as a living body, wherein the target image is any one of multiple images of a target object captured by a camera; calling a living body identification model to extract texture features and high-level semantic features of the target image, wherein the living body identification model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; based on the texture features and high-level semantic features of the target image, predicting whether the target object is a living body.
[0007] According to one aspect of an embodiment of the present invention, a data processing method is provided. The method may include: a display interface of an attendance system receives an attendance request, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera; the attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request, wherein a liveness recognition model is used to extract the texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the attendance system displays the liveness recognition result on the display interface, wherein the target object is predicted to be alive based on the texture features and high-level semantic features of the target image.
[0008] According to one aspect of an embodiment of the present invention, another data processing method is provided. The method may include: displaying a liveness authentication interface on a payment system, and displaying a target image to be subjected to liveness authentication in the liveness authentication interface, wherein the target image is any one of a plurality of images of a target object captured by a camera, and the target object is within the authentication area of the liveness authentication interface; the payment system outputs a verification instruction, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; the payment system obtains texture features and high-level semantic features of the target image based on the verification instruction, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the payment system displays the liveness recognition result on the liveness authentication interface, wherein the target object is predicted to be live based on the texture features and high-level semantic features of the target image; and the payment system performs a payment operation when the target object is confirmed to be live.
[0009] According to another aspect of an embodiment of the present invention, a data processing device is also provided. The device may include: a first acquisition unit, used to acquire a sample image, wherein the sample image is any one of multiple images captured by a camera; a first extraction unit, used to extract a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; an input unit, used to input the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0010] According to another aspect of an embodiment of the present invention, a living body recognition device is also provided. The device may include: a second acquisition unit, used to acquire a target image to be subjected to living body recognition, wherein the target image is any one of multiple images of the target object captured by a camera; a second extraction unit, used to call a living body recognition model to extract texture features and high-level semantic features of the target image, wherein the living body recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; a prediction unit, used to predict whether the target object is alive based on the texture features and high-level semantic features of the target image.
[0011] According to another aspect of an embodiment of the present invention, another data processing device is also provided. The device may include: a first receiving unit, used to receive an attendance request through a display interface of an attendance system, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera; a second receiving unit, used to receive the texture features and high-level semantic features of the target image fed back by the attendance system based on the attendance request, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; a first display unit, used to display the liveness recognition result on a display interface through the attendance system, wherein based on the texture features and high-level semantic features of the target image, it is predicted whether the target object is alive.
[0012] According to another aspect of an embodiment of the present invention, another living body recognition device is provided. The device may include: a second display unit, used to display a liveness authentication interface through a payment system, and display a target image to be subjected to liveness authentication in the liveness authentication interface, wherein the target image is any one of a plurality of images of a target object captured by a camera, and the target object is within an authentication area of the liveness authentication interface; an output unit, used to output a verification instruction through the payment system, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; a third acquisition unit, used to acquire texture features and high-level semantic features of the target image based on the verification instruction through the payment system, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; a third display unit, used to display a liveness recognition result on the liveness authentication interface through the payment system, wherein whether the target object is live is predicted based on the texture features and high-level semantic features of the target image; and an execution unit, used to execute a payment operation when the target object is confirmed to be live through the payment system.
[0013] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute the data processing method of the present invention.
[0014] According to another aspect of an embodiment of the present invention, a processor is provided, which is used to run a program, wherein the program executes the data processing method of the embodiment of the present invention when the program is run by the processor.
[0015] According to another aspect of an embodiment of the present invention, a data processing system is also provided. The system may include: a processor; a memory connected to the processor, and used to provide the processor with instructions for processing the following processing steps: obtaining a sample image, wherein the sample image is any one of multiple images captured by a camera; extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0016] In an embodiment of the present invention, a sample image is obtained, wherein the sample image is any one of the multiple images collected by the camera; a texture feature map and a high-level semantic feature map of the sample image are extracted, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; the target feature map is input into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract the texture features and high-level semantic features of the verification image, and the texture features and the high-level semantic features are used to perform liveness recognition on the verification image. In other words, the present application uses a silent liveness recognition method to extract the texture feature map and the high-level semantic feature map of the sample image and splice them to obtain a target feature map, so as to obtain a liveness recognition model through training with the target feature map, and then use the liveness recognition model to perform liveness recognition on the image to be liveness recognized, thereby refusing to recognize non-liveness attack samples, achieving the purpose of filtering non-liveness attack samples, solving the technical problem of inaccurate liveness recognition results when performing liveness recognition, and achieving the technical effect of improving the accuracy of liveness recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to an embodiment of the present invention;
[0019] Figure 2 is a flow chart of a data processing method according to an embodiment of the present invention;
[0020] Figure 3 is a flow chart of a living body identification method according to an embodiment of the present invention;
[0021] Figure 4 is a flow chart of another data processing method according to an embodiment of the present invention;
[0022] Figure 5 is a flow chart of another data processing method according to an embodiment of the present invention;
[0023] Fig. 6A is a schematic diagram of a living body recognition system according to an embodiment of the present invention;
[0024] Figure 6B is a schematic diagram of a living body recognition scenario according to an embodiment of the present invention;
[0025] Figure 6C is a schematic diagram of another living body recognition scenario according to an embodiment of the present invention;
[0026] Figure 7 is a schematic diagram of a data processing device according to an embodiment of the present invention;
[0027] Figure 8 is a schematic diagram of a living body recognition device according to an embodiment of the present invention;
[0028] Fig. 9 is a schematic diagram of another data processing device according to an embodiment of the present invention;
[0029] Fig.10 is a schematic diagram of another data processing device according to an embodiment of the present invention;
[0030] Fig.11 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0034] Live face recognition: judging whether the face image captured by the camera is real face data or forged attack face data (including printed photos, posters, A4 paper, screen attacks from mobile phones, monitors and Pads, and disguise attacks such as 3D masks);
[0035] Live, the face object directly captured by the camera is the data of real face;
[0036] Non-live attacks: the face objects collected by the camera are fake attack face data (including printed photos, posters, A4 paper, screen attacks from mobile phones, monitors and Pads, and disguise attacks such as 3D masks);
[0037] Attribute information: For live objects, the attribute information includes facial features and expression information (pointed nose, smile, etc.); for non-live attacks, the attribute information includes the type of non-live attacks (printed photos, screen attacks, etc.) and lighting conditions (low light, backlight, etc.);
[0038] Geometric information: For living objects, the geometric information includes the depth map of the face; for non-living objects, the geometric information includes the reflection map of paper, screen or mask.
[0039] Example 1
[0040] According to an embodiment of the present invention, an embodiment of a data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to an embodiment of the present invention. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0043] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data processing method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the data processing method of the above-mentioned application program is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0044] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0045] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0046] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).
[0047] exist Figure 1 In the operating environment shown, this application provides Figure 2 It should be noted that the data processing method of this embodiment can be Figure 1 The illustrated embodiment is executed by a mobile terminal.
[0048] Figure 2 is a flow chart of a data processing method according to an embodiment of the present invention. Figure 2 As shown, the data processing method may include the following steps:
[0049] Step S202, obtaining a sample image, wherein the sample image is any one of the multiple images captured by the camera.
[0050] In the technical solution provided in the above step S202 of the present invention, multiple images can be collected by a camera, for example, the multiple images are a sequence of images continuously collected by the camera within a period of time, and any one of the multiple images is determined as a sample image, that is, this embodiment relies on a sample image collected by the camera for recognition to perform silent liveness recognition. Since silent liveness recognition processes a single image, it has a faster response speed and fewer model parameters, which can improve the efficiency of liveness recognition. Among them, the sample image can include live samples and non-live attack samples.
[0051] Step S204, extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are concatenated to obtain a target feature map.
[0052] In the technical solution provided in the above step S204 of the present invention, after obtaining the sample image, a texture feature map and a high-level semantic feature map of the sample image may be extracted.
[0053] In this embodiment, the texture feature map extracted from the sample image also needs to be trained. The texture information of the local details of the sample image can be calculated first to form a feature expression, and then the feature is pooled multiple times to continuously increase the receptive field of the convolutional neurons to form texture features at different scales, which can be the underlying texture features, and then a texture feature map is generated through it.
[0054] In this embodiment, the high-level semantic information of the sample image can be fully extracted through a cascade structure of multiple residual modules, a high-level speech feature map can be generated through the high-level speech information, and then the high-level semantic feature map can be output.
[0055] Since the local details of the sample image have a significant guiding role in the live attack category and the live identification binary classification, after generating the texture feature map and the high-level semantic feature map of the sample image, this embodiment can splice the texture feature map and the high-level semantic feature map, for example, determine the channel direction of the texture feature map and the high-level semantic feature map, and splice the texture feature map and the high-level semantic feature map along the determined channel direction to obtain the target feature map.
[0056] Step S206, input the target feature map into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of the verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0057] In the technical solution provided in the above step S206 of the present invention, after the texture feature map and the high-level semantic feature map are spliced to obtain the target feature map, the target feature map can be input into the fully connected layer (FC) of the neural network model to train and obtain a living body recognition model.
[0058] In this embodiment, the neural network model includes a fully connected layer. This embodiment can input the obtained target feature map into the fully connected layer for training to obtain a liveness recognition model. The liveness recognition model extracts the texture features and high-level semantic features of the verification image to be subjected to liveness recognition. The texture features and high-level semantic features can further perform liveness recognition on the verification image, for example, to predict the type of non-liveness attack. Since the liveness recognition model of this embodiment is obtained by training the target feature map obtained by splicing the texture feature map and the high-level semantic feature map, it can avoid the defect that the learned features do not have generalization and insufficient ability to discriminate samples.
[0059] Through the above steps S202 to S206, a sample image is obtained, wherein the sample image is any one of the multiple images collected by the camera; a texture feature map and a high-level semantic feature map of the sample image are extracted, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; the target feature map is input into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract the texture features and high-level semantic features of the verification image, and the texture features and the high-level semantic features are used to perform liveness recognition on the verification image. That is to say, this embodiment uses a silent liveness recognition method to extract the texture feature map and the high-level semantic feature map of the sample image, and splices them to obtain the target feature map, so as to obtain the liveness recognition model through the target feature map training, and then uses the liveness recognition model to perform liveness recognition on the liveness recognition image to be performed, thereby refusing to recognize non-liveness attack samples, achieving the purpose of filtering non-liveness attack samples, solving the technical problem of inaccurate liveness recognition results when performing liveness recognition, and achieving the technical effect of improving the accuracy of liveness recognition.
[0060] The above method of this embodiment is further introduced below.
[0061] As an optional implementation, step S204, extracting a texture feature map of the sample image, includes: determining a region image where a target object is located in the sample image; identifying texture information in the region image; and using a central difference convolutional network model to perform a pooling operation on the texture information in the region image to generate a texture feature map.
[0062] In this embodiment, when extracting the texture feature map of the sample image, the area image where the target object is located in the sample image can be determined first, and the texture information therein is identified for the area image where the target object is located. The texture information is the texture information of the local details of the sample image, and a feature expression is formed through it. Then, a central difference convolution network model (central difference conv, referred to as CD conv) is used to perform a pooling operation on the texture information in the above-mentioned area image. For example, the texture information in the above-mentioned area image is pooled three times, and the receptive field of the convolution neurons is continuously increased, thereby forming texture features at different scales, and a texture feature map is generated through the texture features. Among them, this embodiment does not specifically limit the number of central difference convolution network models. Its residual ratio can be adjusted.
[0063] As an optional implementation, extracting a texture feature map and a high-level semantic feature map of a sample image includes: using an extraction network cascaded with multiple residual modules in a pre-trained model to analyze the sample image and extracting high-level semantic information from the sample image; and constructing a high-level semantic feature map based on the high-level semantic information.
[0064] In this embodiment, the overall network structure of the pre-trained model can be consistent with the network structure of the residual model (ResNet-18), and can be initialized using the parameters of the ResNet-18 pre-trained by the image data set (ImageNet). The pre-trained model of this embodiment includes multiple residual modules, for example, the multiple residual modules are cascaded structures to form an extraction network, and the sample image can be analyzed by the extraction network to fully extract high-level semantic information from the sample image, and then based on the high-level semantic information, a high-level semantic feature map is constructed, and the high-level semantic feature map is output.
[0065] As an optional implementation, the high-level semantic information includes at least one of the following: facial information and lighting information of a sample image, wherein a high-level semantic feature map is constructed based on the high-level semantic information, including: training the facial information in the sample image to generate attribute information of the face in the sample image; training the light information in the sample image to generate lighting information of the face in the sample image; and constructing a high-level semantic feature map based on the attribute information and lighting information of the face in the sample image.
[0066] In this embodiment, the high-level semantic information may include facial information and lighting information of the target object in the sample image, wherein the facial information and lighting information of the target object belong to the attribute information of the live sample and the non-live attack sample in the sample image, the facial information may be used to generate the attribute information of the face in the sample image, the attribute information is used to indicate the attribute of the face (human face attribute), which is a category label, and the lighting information may be the lighting condition of the environment, which is used to generate the lighting information of the face in the sample image, the lighting information is used to indicate the lighting condition of the face, and may be the lighting information of the camera when the sample image is collected as a non-live attack sample. When constructing a high-level semantic feature map based on high-level semantic information, the above-mentioned facial information in the sample image may be used to train the attribute information of the face, and the light information in the sample image may be used to train the lighting information of the face in the sample image. After obtaining the attribute information and illumination information of the face in the sample image, a high-level semantic feature map can be constructed based on the attribute information and illumination information of the face in the sample image, and then the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map. The target feature image is used to train a liveness recognition model to perform liveness recognition on the images to be subjected to liveness recognition, thereby avoiding inaccurate liveness recognition results due to insufficient consideration of facial information and illumination information collected from different faces.
[0067] Optionally, in this embodiment, considering that the facial information and lighting information of the sample image belong to high-level semantic information, this embodiment can directly use the constructed high-level semantic feature map as the input of the fully connected layer of the neural network model to further predict the facial attribute information and lighting information.
[0068] Optionally, in this embodiment, cross entropy loss can be used for liveness recognition, non-liveness attack sample classification and illumination classification of sample images, and the corresponding weights can be set to 1, 0.1, and 0.01, respectively; since even the same face may have multiple facial attributes, for the classification of facial attributes, this embodiment can use a cross entropy loss function (BCEloss), and its weight can be set to 1.
[0069] As an optional implementation, in step S204, after extracting the texture feature map and the high-level semantic feature map of the sample image, the method further includes: obtaining a depth estimation and a reflection estimation of the sample image based on the high-level semantic feature map, wherein, for the estimation functions of the depth estimation and the reflection estimation, MES loss is used as a loss function.
[0070] In this embodiment, liveness recognition of the sample image is regarded as a binary classification problem, and the depth estimation and reflection estimation are both derived from the high-level semantic feature map. Therefore, after extracting the texture feature map and the high-level semantic feature map of the sample image, this embodiment can obtain the depth estimation and reflection estimation (adaptive) of the sample image based on the high-level semantic feature map through the above-mentioned central difference convolutional network model, wherein the depth estimation can be the depth map estimation of the live sample, and the reflection estimation can be the reflection map estimation of the non-live attack sample.
[0071] Optionally, for the estimation functions of the depth estimation and reflection estimation, this embodiment may adopt MES loss as a loss function, and the weight of the loss function may be set to 0.1.
[0072] As an optional implementation, in step S206, after obtaining the liveness recognition model, the method also includes: obtaining a verification image to be subjected to liveness recognition; using the liveness recognition model to extract bottom-level texture features and high-level semantic features from the verification image; and predicting whether the target object in the verification image is alive based on the texture features and semantic features.
[0073] In this embodiment, after the liveness recognition model is trained, the liveness recognition model can be used to perform liveness recognition on the image to be subjected to liveness recognition. This embodiment can obtain a verification picture to be subjected to liveness recognition, and the verification picture can be obtained by collecting an image acquisition device. After obtaining the verification picture, the liveness recognition model can be used to extract the texture features at the bottom layer and the semantic features at the high layer from the verification picture. Optionally, when predicting different attributes, this embodiment can specifically design a bottom layer texture feature extraction network and a high layer semantic feature extraction network, and can use post-processing means to further improve its accuracy, wherein the liveness recognition model includes a bottom layer texture feature extraction network and a high layer semantic feature extraction network, the bottom layer texture feature extraction network can also be referred to as a bottom layer texture feature extractor, and the high layer semantic feature extraction network can also be referred to as a high layer semantic feature extractor. This embodiment can extract the texture features at the bottom layer from the verification picture through the bottom layer texture feature extraction network, and can extract the semantic features at the high layer from the verification picture through the high layer semantic feature extraction network, and then predict whether the target object in the verification picture is alive based on the texture features and the semantic features.
[0074] As an optional implementation, when the target object in the verification picture is non-living, the type of the verification picture is determined based on texture features and semantic features extracted from the verification picture, wherein the type of the verification picture includes at least one of the following: a flat photo, a photo of a three-dimensional mask, and a display image in a display device.
[0075] In this embodiment, when the target object in the verification image is non-living, that is, the verification image is determined to be a non-living attack sample, the type of the verification image can be determined based on the texture features and semantic features extracted from the verification image, that is, the category of the non-living attack sample can be determined, which can be a flat photo, a photo of a three-dimensional mask, or a display image in a display device. Therefore, this embodiment fully takes into account the categories of different non-living attack samples and improves the accuracy of the liveness recognition results.
[0076] Optionally, considering that the reflectivity map has a negative effect on the prediction of printed non-living attack samples, when the training image is a printed attack sample (e.g., a photo, A4 paper), this embodiment can cut off the gradient propagation from the reflectivity map, and can perform moiré detection processing on the verification image to eliminate false negative samples in one step, thereby improving the accuracy of living body recognition.
[0077] Optionally, the embodiment can intercept the face of the target object in the sample image, obtain an image block containing only skin color in the face, calculate the third-order color moment of the image block, and use the calculated result as a feature for clustering. Optionally, the embodiment can divide the sample images of non-living samples into two categories: containing and not containing moiré patterns, and use the living samples and non-living attack samples containing moiré patterns as training sets to train a moiré recognition network, so that the verification images predicted to be living can be further processed for moiré detection to further exclude false negative samples, thereby improving the accuracy of living recognition.
[0078] The embodiment of the present invention also provides a living body recognition method based on the testing process of the living body recognition model.
[0079] Figure 3 FIG. 1 is a flow chart of a method for identifying a living body according to an embodiment of the present invention. Figure 3 As shown, the method may include the following steps:
[0080] Step S302, obtaining a target image to be subjected to living body recognition, wherein the target image is any one of a plurality of images of the target object captured by a camera.
[0081] In the technical solution provided in the above step S302 of the present invention, multiple images of the target object can be captured by the camera, and any one of the multiple images can be determined as the target image to be subjected to liveness recognition. That is, this embodiment relies on a picture captured by the camera for recognition to perform silent liveness recognition.
[0082] Step S304: calling a living body recognition model to extract texture features and high-level semantic features of the target image.
[0083] In the technical solution provided in the above step S304 of the present invention, after obtaining the target image to be subjected to liveness identification, the liveness identification model can be called to extract the texture features and high-level semantic features of the target image through the liveness identification model, wherein the liveness identification model is generated by training the target feature map, and the target feature map is composed of the texture feature map and the high-level semantic feature map identified from the sample image.
[0084] In this embodiment, the liveness recognition model is a pre-trained model for liveness recognition of the target image to be liveness recognized. Optionally, this embodiment first calculates the texture information of the local details of the sample image to form a feature expression, and then performs multiple pooling operations on the feature to continuously increase the receptive field of the convolutional neurons to form texture features at different scales, and generates a texture feature map of the sample image through it; it can also fully extract the high-level semantic information of the above sample image through a cascade structure of multiple residual modules, and generate a high-level speech feature map of the sample image through the high-level speech information; after generating the texture feature map and high-level semantic feature map of the sample image, the texture feature map and the high-level semantic feature map of the sample image can be spliced to obtain the target feature map of the sample image, and the target feature map can be input into the fully connected layer of the neural network model for training to obtain the liveness recognition model, and then call the liveness recognition model to extract the texture features and high-level semantic features of the target image.
[0085] Step S306: predict whether the target object is alive based on the texture features and high-level semantic features of the target image.
[0086] In the technical solution provided in the above step S306 of the present invention, after acquiring the texture features and high-level semantic features of the target image, it is possible to predict whether the target object is alive based on the texture features and high-level semantic features of the target image. A living body recognition model may be used to learn the above-mentioned texture features and semantic features, thereby predicting whether the target object in the target image is alive.
[0087] The embodiment of the present invention also provides another data processing method from the attendance application scenario.
[0088] Figure 4 FIG. 1 is a flow chart of another data processing method according to an embodiment of the present invention. Figure 4 As shown, the method may include the following steps:
[0089] Step S402: the display interface of the attendance system receives an attendance request, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera.
[0090] In the technical solution provided in the above step S402 of the present invention, the attendance system has a display interface, and the embodiment can receive an attendance request on the above display interface, and the attendance request can be triggered by the target object to be checked. The attendance request is used to capture a target image of the target object, and any one of the multiple images of the target object captured by the camera can be determined as the target image to be subjected to live body recognition, that is, the embodiment relies on a picture captured by the camera for recognition to perform silent live body recognition.
[0091] Step S404: the attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request.
[0092] In the technical solution provided in the above step S404 of the present invention, a liveness recognition model is used to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image.
[0093] In this embodiment, after the display interface of the attendance system receives the attendance request, the attendance system can respond to the attendance request and use a pre-trained liveness recognition model to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image.
[0094] Step S406: The attendance system displays the liveness recognition result on the display interface, wherein it predicts whether the target object is alive based on the texture features and high-level semantic features of the target image.
[0095] In the technical solution provided in the above step S406 of the present invention, after the attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request, the attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request.
[0096] In this embodiment, a liveness recognition model may be used to learn the texture features and semantic features received by the attendance system to predict whether the target object in the target image is alive, and then the liveness recognition result of whether the target object is alive is displayed on the display interface of the attendance system, thereby achieving the purpose of facial attendance.
[0097] The embodiment of the present invention also provides another data processing method from a payment application scenario.
[0098] Figure 5 FIG. 1 is a flow chart of another data processing method according to an embodiment of the present invention. Figure 5 As shown, the method may include the following steps:
[0099] Step S502: The payment system displays a liveness authentication interface, and displays a target image to be subjected to liveness recognition in the liveness authentication interface.
[0100] In the technical solution provided in the above step S502 of the present invention, the target image is any one of the multiple images of the target object captured by the camera, and the target object is within the authentication area of the liveness authentication interface.
[0101] In this embodiment, a liveness authentication interface is displayed on the payment system, and the liveness authentication interface is an interface for performing liveness authentication on a target object. This embodiment can capture multiple images of the target object through a camera, and determine any one of the multiple images as a target image to be subjected to liveness identification and display it in the above-mentioned liveness authentication interface.
[0102] Step S504: the payment system outputs a verification instruction, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction.
[0103] In the technical solution provided in the above step S504 of the present invention, after the target image to be subjected to liveness identification is displayed in the liveness authentication interface, the payment system can output a verification instruction, which can be used to instruct the target object displayed in the target image to perform a predetermined action, such as raising the head, lowering the head, turning the head, etc., so that the target object's action meets the standard, which can be voice, text, etc., and is not specifically limited here.
[0104] Step S506: The payment system obtains texture features and high-level semantic features of the target image based on the verification instruction.
[0105] In the technical solution provided in the above step S506 of the present invention, after the payment system outputs the verification instruction, the payment system obtains the texture features and high-level semantic features of the target image based on the verification instruction, wherein the liveness recognition model is used to extract the texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training the target feature map, wherein the target feature map is composed of the texture feature map and the high-level semantic feature map identified from the sample image.
[0106] In this embodiment, the payment system can respond to the verification instruction and use a pre-trained liveness recognition model to extract the texture features and high-level semantic features of the target image. Among them, the liveness recognition model is a pre-trained model for liveness recognition of the target image to be subjected to liveness recognition. Optionally, this embodiment first calculates the texture information of the local details of the sample image to form a feature expression, and then performs multiple pooling operations on the features to continuously increase the receptive field of the convolutional neurons to form texture features at different scales, and generates a texture feature map of the sample image through it; it can also fully extract the high-level semantic information of the above sample image through a cascade structure of multiple residual modules, and generate a high-level speech feature map of the sample image through the high-level speech information; after generating the texture feature map and the high-level semantic feature map of the sample image, the texture feature map and the high-level semantic feature map of the sample image can be spliced to obtain the target feature map of the sample image, and the target feature map can be input into the fully connected layer of the neural network model for training to obtain the liveness recognition model, and then the liveness recognition model is used to extract the texture features and high-level semantic features of the target image.
[0107] Step S508: The payment system displays the liveness recognition result on the liveness authentication interface, wherein it predicts whether the target object is alive based on the texture features and high-level semantic features of the target image.
[0108] In the technical solution provided in the above step S508 of the present invention, after the payment system obtains the texture features and high-level semantic features of the target image, it can use a liveness recognition model to learn the above texture features and semantic features, so as to predict whether the target object in the target image is alive, and then display the liveness recognition result of whether the target object is alive on the liveness authentication interface.
[0109] Step S510: When the payment system confirms that the target object is alive, the payment operation is performed.
[0110] In the technical solution provided in the above step S510 of the present invention, if the liveness authentication interface shows that the target object is live, it is determined that the target object himself operates the payment system, thereby executing the payment operation and achieving the purpose of face-swiping payment.
[0111] In the related technology, criminals can perform face recognition by holding a photo or video clip of the relevant person to achieve the purpose of disguising the identity of the person; the related technology can also use silent liveness recognition to achieve liveness recognition, but in the process of silent liveness recognition, the recognition algorithm used has errors in the recognized liveness features, resulting in inaccurate technical problems in the liveness recognition results.
[0112] However, in this embodiment, a liveness recognition model can be used to judge whether the object in a picture collected by the current camera is a live sample or a non-live attack sample, and non-live attack samples can be rejected, thereby achieving the purpose of filtering non-live attack samples, solving the technical problem of inaccurate liveness recognition results during liveness recognition, achieving the technical effect of improving the accuracy of liveness recognition, and avoiding the security risks brought by non-liveness attack samples to face recognition.
[0113] Example 2
[0114] The technical solution of the present invention is illustrated below in conjunction with preferred implementation modes.
[0115] Terminal devices for face recognition have gradually entered people's lives and become an important means of identifying personal identities, such as face payment and face attendance. However, these terminal devices are facing challenges from non-live attack samples. Criminals can attack the face recognition system by holding photos or video clips of relevant personnel, thereby disguising the identity of the person and passing through the face recognition system. Therefore, non-live attacks are a major security risk in face recognition.
[0116] Liveness recognition can be divided into time-series liveness recognition and silent liveness recognition based on the data form used. In the process of time-series liveness recognition, the camera collects a continuous sequence of images within a period of time to determine whether the object in the sequence is alive; in the process of silent liveness recognition, only one image collected by the camera is used to determine whether the object in the image is alive.
[0117] In the related art, there are two methods for silent liveness recognition. One is to regard the liveness recognition task as a binary classification task, obtain image features through a series of convolutional neural network operations (such as convolution, pooling, and jump connection), and predict whether a given face is a live sample or a non-live attack sample through a full connection. However, the disadvantage of this solution is that the learned features are not generalizable and the ability to discriminate samples is insufficient; the other method is to introduce blood flow, depth map, LBP map estimation and other methods to convert the liveness recognition task into a supervised problem with a target, assist with some convolution operators that are sensitive to local textures, and finally calculate the estimated depth map to determine whether the given image is a non-live attack sample. However, the disadvantage of this method is that it does not fully consider the camera and lighting conditions of different face acquisitions, and does not take into account the attributes of different attack categories and faces.
[0118] In practical applications, compared with sequential liveness recognition, silent liveness recognition has a faster response speed and fewer model parameters because it processes a single image. This embodiment can use silent liveness recognition to determine whether the sample collected by the current camera is live or a non-live attack sample, refuse to recognize a non-live attack sample, and achieve the function of filtering non-live samples, thereby solving the technical problem of inaccurate liveness recognition results and the potential safety hazards brought by attack samples to the face recognition system.
[0119] The above method of this embodiment is further illustrated below.
[0120] Fig. 6A Schematic diagram of a living body recognition system according to an embodiment of the present invention. Fig. 6A As shown in the figure, a sample image is input. After the sample image enters the entire system, module 1 calculates the texture information of the local details of the sample image through CD conv to form a feature expression. The receptive field of the convolutional neuron can be continuously increased through three pooling operations (pool) to form texture features at different scales, thereby obtaining the feature Figure 1 Optionally, the above three pooling operations are performed alternately with three CD conv operations, where CD conv includes one convolutional layer (conv).
[0121] It should be noted that the residual ratio of the above CD Conv can be adjusted, and the number of CD Conv in module 1 can be adjusted.
[0122] In this embodiment, the overall network structure of module 2 is consistent with ResNet-18, and the parameters of ResNet-18 pre-trained by ImageNet can be used to initialize module 2. The module 2 can be composed of 1 convolution layer, 1 pooling layer, a cascade structure of 4 residual modules, and 1 pooling layer, wherein the cascade structure of 4 residual modules can be a cascade structure of 4 residual modules Base Block1, Base Block2, Base Block2, and Base Block2, wherein Base Block1 is composed of 4 convs, and Base Block2 is composed of two convs for upsampling and two convs for downsampling. The high-level semantic information of the given sample image is fully extracted through the above module 2, and a high-level semantic feature map is output, that is, the feature Figure 2 .
[0123] Considering that the local details of the sample image have a significant guiding role in the classification of live attack and live identification, this embodiment can Figure 1 and Features Figure 2The target feature map is obtained by splicing along the channel direction. The spliced target feature map is used as the input of the subsequent fully connected layer (FC) to train a liveness recognition model for predicting non-liveness attack types and liveness recognition.
[0124] Optionally, this embodiment can perform binary classification of liveness identification on the spliced target feature map through one FC, and perform classification of non-liveness attack type through another FC. Optionally, the above-mentioned binary classification of liveness identification in this embodiment is to predict true / false, and the above-mentioned classification of non-liveness attack type can be to predict an 11-dimensional one-hot vector to indicate the category of non-liveness attack.
[0125] Optionally, this embodiment has the feature Figure 2 CD conv operation is performed to obtain depth estimation of living samples and reflection estimation of non-living attack samples (adaptive). In this embodiment, the adaptation corresponds to different loss function calculation methods for different types of attack samples. Optionally, for attack samples of A4 paper and photo type, the reflection estimation loss may not be back-propagated.
[0126] In this embodiment, considering that the face attributes and the lighting of the environment belong to high-level semantic information, this embodiment can directly use the feature Figure 2 The facial attributes are classified by one FC, and the ambient illumination is classified by another FC. For depth map and reflectance map estimation, this embodiment can use MES loss as the loss function, and set its weight to 0.1; for the two-classification of liveness recognition (liveness recognition), non-liveness attack type classification, and illumination classification, this embodiment can use cross entropy loss, and the corresponding weights are set to 1, 0.1, and 0.01 respectively; considering that the same face may have multiple facial attributes, for facial attribute classification, this embodiment can use BCE loss, and its weight can be set to 1. Considering that the reflectance map has a negative effect on the prediction of printed attack samples, when processing printed attack samples (photos, A4 paper), this embodiment can cut off the gradient propagation from the reflectance map. By performing moiré detection on sample images predicted to be alive, false negative samples are further excluded, thereby improving the accuracy of liveness recognition.
[0127] Optionally, this embodiment can capture the facial part of the screen-shot image to obtain an image block containing only skin color in the face, and can calculate the third-order color moment of the image block to use it as a feature for clustering, and divide the non-living samples captured by the screen into two categories: containing moiré and not containing moiré. The real samples and the attack samples containing moiré are used as training sets to train a moiré recognition network, so as to perform moiré detection processing on the sample images predicted to be living.
[0128] Figure 6B FIG. 1 is a schematic diagram of a scene of living body recognition according to an embodiment of the present invention. Figure 6B As shown, a sample image is input into a computing device, and any one of the multiple images can be determined as the above-mentioned sample image. After obtaining the sample image, the texture feature map and the high-level semantic feature map of the sample image can be extracted, and the texture feature map and the high-level semantic feature map can be spliced to obtain a target feature map. This embodiment can input the obtained target feature map into a fully connected layer for training to obtain a liveness recognition model, and use the liveness recognition model to extract the texture features and high-level semantic features of the target image. Based on the texture features and high-level semantic features of the target image, it can be predicted whether the target object of the target image is alive, and then the liveness recognition result of whether the target object is alive is input into the computing device for display.
[0129] Figure 6C FIG. 1 is a schematic diagram of another scenario of living body recognition according to an embodiment of the present invention. Figure 6C As shown, the target image to be subjected to liveness recognition can be added to the display interface, wherein the target image is any one of the multiple images of the target object captured by the camera. Then, in response to the feature extraction operation acting on the display interface, the liveness recognition model is called to extract the texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training the target feature map, which is composed of the texture feature map and the high-level semantic feature map identified from the sample image, and then based on the texture features and high-level semantic features of the target image, it is predicted whether the target object of the target image is alive, and the liveness recognition result of whether the target object is alive is displayed on the display interface.
[0130] Through the above scheme, this embodiment can make full use of the attribute information of living and non-living attack samples, and use different high-level semantic information (ResNet-18) and low-level texture information (CD Conv) for processing through different category labels (for example, face attributes, non-living attack attributes).
[0131] In addition, in this embodiment, liveness recognition is not only regarded as a binary classification problem, but also, while introducing depth estimation and reflection estimation, fully explores the facial attributes, the categories of non-live attack samples, and the lighting information when the attack samples are collected. When predicting different attributes, the underlying texture feature extraction network and the high-level semantic feature extraction network can be designed in a targeted manner, and post-processing methods can be used to further improve the accuracy; this embodiment also uses post-processing methods for moiré detection to further improve the accuracy of liveness recognition, thereby solving the technical problem of inaccurate liveness recognition results when performing liveness recognition.
[0132] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0134] Example 3
[0135] According to an embodiment of the present invention, there is also provided a method for implementing the above Figure 2 A data processing device according to the data processing method shown.
[0136] Figure 7 is a schematic diagram of a data processing device according to an embodiment of the present invention. Figure 7 As shown, the data processing device 70 may include: a first acquiring unit 71 , a first extracting unit 72 and an input unit 73 .
[0137] The first acquisition unit 71 is used to acquire a sample image, wherein the sample image is any one of the multiple images captured by the camera.
[0138] The first extraction unit 72 is used to extract a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are concatenated to obtain a target feature map.
[0139] The input unit 73 is used to input the target feature map into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of the verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0140] It should be noted that the first acquisition unit 71, the first extraction unit 72 and the input unit 73 correspond to steps S202 to S206 of Example 1, respectively. The three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above units, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0141] According to an embodiment of the present invention, there is also provided a method for implementing the above Figure 3 The living body identification method and the living body identification device are shown.
[0142] Figure 8 FIG. 1 is a schematic diagram of a living body recognition device according to an embodiment of the present invention. Figure 8 As shown, the living body recognition device 80 may include: a second acquisition unit 81, a second extraction unit 82 and a prediction unit 83.
[0143] The second acquisition unit 81 is used to acquire a target image to be used for living body recognition, wherein the target image is any one of multiple images of the target object captured by a camera.
[0144] The second extraction unit 82 is used to call the liveness recognition model to extract the texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training the target feature map, and the target feature map is composed of the texture feature map and the high-level semantic feature map identified from the sample image.
[0145] The prediction unit 83 is used to predict whether the target object is alive based on the texture features and high-level semantic features of the target image.
[0146] It should be noted that the second acquisition unit 81, the second extraction unit 82 and the prediction unit 83 correspond to steps S302 to S306 of Example 1, respectively. The three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above units, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0147] According to an embodiment of the present invention, there is also provided a method for implementing the above Figure 4 A data processing device according to the data processing method shown.
[0148] Fig. 9 is a schematic diagram of another data processing device according to an embodiment of the present invention. Fig. 9 As shown, the data processing device 90 may include: a first receiving unit 91 , a second receiving unit 92 and a first display unit 93 .
[0149] The first receiving unit 91 is used to receive an attendance request through a display interface of the attendance system, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera.
[0150] The second receiving unit 92 is used to receive the texture features and high-level semantic features of the target image fed back based on the attendance request through the attendance system, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image.
[0151] The first display unit 93 is used to display the live body recognition result on the display interface through the attendance system, wherein whether the target object is a live body is predicted based on the texture features and high-level semantic features of the target image.
[0152] It should be noted that the first receiving unit 91, the second receiving unit 92 and the first display unit 93 correspond to steps S402 to S406 of Example 1, respectively. The three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above units, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0153] According to an embodiment of the present invention, there is also provided a method for implementing the above Figure 4 A data processing device according to the data processing method shown.
[0154] Fig.10 is a schematic diagram of another data processing device according to an embodiment of the present invention. Fig.10 As shown, the data processing device 100 may include: a second display unit 101 , an output unit 102 , a third acquisition unit 103 , a third display unit 104 and an execution unit 105 .
[0155] The second display unit 101 is used to display the liveness authentication interface through the payment system, and display the target image to be subjected to liveness recognition in the liveness authentication interface, wherein the target image is any one of multiple images of the target object captured by the camera, and the target object is in the authentication area of the liveness authentication interface.
[0156] The output unit 102 is used to output a verification instruction through the payment system, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction.
[0157] The third acquisition unit 103 is used to obtain the texture features and high-level semantic features of the target image based on the verification instruction through the payment system, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image.
[0158] The third display unit 104 is used to display the liveness recognition result on the liveness authentication interface through the payment system, wherein whether the target object is alive is predicted based on the texture features and high-level semantic features of the target image.
[0159] The execution unit 105 is used to execute the payment operation when the target object is confirmed to be alive through the payment system.
[0160] It should be noted that the second display unit 101, the output unit 102, the third acquisition unit 103, the third display unit 104 and the execution unit 105 correspond to steps S502 to S510 of Example 1 respectively, and the five units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above units, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0161] In the liveness recognition device of this embodiment, a silent liveness recognition method is used to extract the texture feature map and the high-level semantic feature map of the sample image and concatenate them to obtain a target feature map, so as to train a liveness recognition model through the target feature map, and then use the liveness recognition model to perform liveness recognition on the image to be subjected to liveness recognition, thereby refusing to recognize non-liveness attack samples, achieving the purpose of filtering non-liveness attack samples, solving the technical problem of inaccurate liveness recognition results when performing liveness recognition, and achieving the technical effect of improving the accuracy of liveness recognition.
[0162] Example 4
[0163] The embodiment of the present invention can provide a living body recognition system, which can include a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0164] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
[0165] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the data processing method of the application: obtaining a sample image, wherein the sample image is any one of the multiple images captured by the camera; extracting the texture feature map and the high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain the target feature map; inputting the target feature map into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract the texture features and high-level semantic features of the verification image, and the texture features and the high-level semantic features are used to perform liveness recognition on the verification image.
[0166] Optionally, Fig.11 is a structural block diagram of a computer terminal according to an embodiment of the present invention. Fig.11 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 112 , a memory 114 , and a transmission device 116 .
[0167] Among them, the transmission device is used to transmit sample pictures; wherein the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the data processing method and device in the embodiment of the present invention, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned data processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0168] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain a sample image, wherein the sample image is any one of the multiple images captured by the camera; extract a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; input the target feature map into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of the verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0169] Optionally, the processor may also execute the program code of the following steps: determining the area image where the target object is located in the sample image; identifying the texture information in the area image; and using a central difference convolutional network model to perform a pooling operation on the texture information in the area image to generate a texture feature map.
[0170] Optionally, the processor may also execute the program code of the following steps: analyzing the sample image using an extraction network cascaded with multiple residual modules in a pre-trained model to extract high-level semantic information from the sample image; and constructing a high-level semantic feature map based on the high-level semantic information.
[0171] Optionally, the processor may also execute program code of the following steps: training facial information in sample images to generate attribute information of the face in the sample images; training light information in sample images to generate illumination information of the face in the sample images; and constructing a high-level semantic feature map based on the attribute information and illumination information of the face in the sample images.
[0172] Optionally, the processor may also execute the program code of the following steps: after extracting the texture feature map and the high-level semantic feature map of the sample image, obtain the depth estimation and reflection estimation of the sample image based on the high-level semantic feature map, wherein, for the estimation function of the depth estimation and the reflection estimation, MES loss is used as the loss function.
[0173] Optionally, the processor may also execute the program code of the following steps: after training the liveness recognition model, obtaining a verification image to be subjected to liveness recognition; using the liveness recognition model to extract underlying texture features and high-level semantic features from the verification image; and based on the texture features and semantic features, predicting whether the target object in the verification image is alive.
[0174] Optionally, the processor may also execute program code of the following steps: when the target object in the verification picture is non-living, determining the type of the verification picture based on texture features and semantic features extracted from the verification picture, wherein the type of the verification picture includes at least one of the following: a flat photo, a photo of a three-dimensional mask, and a display image in a display device.
[0175] As an optional example, the processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtain the target image to be identified as alive, wherein the target image is any one of the multiple images of the target object captured by the camera; call the liveness recognition model to extract the texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training the target feature map, and the target feature map is composed of the texture feature map and the high-level semantic feature map identified from the sample image; based on the texture features and high-level semantic features of the target image, predict whether the target object is alive.
[0176] As an optional example, the processor can also call the information and application stored in the memory through the transmission device to perform the following steps: the display interface of the attendance system receives an attendance request, wherein the attendance request is used to capture a target image of the target object, and the target image is any one of multiple images of the target object captured by the camera; the attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the attendance system displays the liveness recognition result on the display interface, wherein based on the texture features and high-level semantic features of the target image, it predicts whether the target object is alive.
[0177] As an optional example, the processor can also call the information and application stored in the memory through the transmission device to perform the following steps: the payment system displays a liveness authentication interface, and displays the target image to be liveness identified in the liveness authentication interface, wherein the target image is any one of the multiple images of the target object captured by the camera, and the target object is in the authentication area of the liveness authentication interface; the payment system outputs a verification instruction, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; the payment system obtains the texture features and high-level semantic features of the target image based on the verification instruction, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the payment system displays the liveness recognition result on the liveness authentication interface, wherein the target object is predicted to be live based on the texture features and high-level semantic features of the target image; when the payment system confirms that the target object is live, the payment operation is performed.
[0178] According to an embodiment of the present invention, a data processing scheme is provided. A sample image is obtained, wherein the sample image is any one of the multiple images collected by the camera; a texture feature map and a high-level semantic feature map of the sample image are extracted, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; the target feature map is input into the fully connected layer of the neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract the texture features and high-level semantic features of the verification image, and the texture features and the high-level semantic features are used to perform liveness recognition on the verification image. That is, the embodiment uses a silent liveness recognition method to extract the texture feature map and the high-level semantic feature map of the sample image, splice them to obtain the target feature map, and train the liveness recognition model through the target feature map, and then use the liveness recognition model to perform liveness recognition on the liveness recognition image to reject the recognition of non-liveness attack samples, thereby achieving the purpose of filtering non-liveness attack samples, solving the technical problem of inaccurate liveness recognition results when performing liveness recognition, and achieving the technical effect of improving the accuracy of liveness recognition.
[0179] It can be understood by those skilled in the art that Fig.11 The structure shown is for illustration only, and the computer terminal A may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.11 It does not limit the structure of the above-mentioned computer terminal A. For example, the computer terminal A may also include Fig.11 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.11 Different configurations are shown.
[0180] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0181] Example 4
[0182] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0183] Optionally, in this embodiment, a computer-readable storage medium is configured to store program codes for executing the following steps: obtaining a sample image, wherein the sample image is any one of multiple images captured by a camera; extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are concatenated to obtain a target feature map; inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image.
[0184] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: determining a region image where a target object is located in a sample image; identifying texture information in the region image; and using a central difference convolutional network model to perform a pooling operation on the texture information in the region image to generate a texture feature map.
[0185] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: analyzing a sample image using an extraction network cascaded with multiple residual modules in a pre-trained model to extract high-level semantic information from the sample image; and constructing a high-level semantic feature map based on the high-level semantic information.
[0186] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: training facial information in sample images to generate attribute information of the face in the sample images; training light information in sample images to generate illumination information of the face in the sample images; and constructing a high-level semantic feature map based on the attribute information and illumination information of the face in the sample images.
[0187] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: after extracting the texture feature map and the high-level semantic feature map of the sample image, obtaining the depth estimation and reflection estimation of the sample image based on the high-level semantic feature map, wherein, for the estimation function of the depth estimation and the reflection estimation, MES loss is used as the loss function.
[0188] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: after training the liveness recognition model, obtaining a verification image to be subjected to liveness recognition; using the liveness recognition model to extract underlying texture features and high-level semantic features from the verification image; and based on the texture features and semantic features, predicting whether the target object in the verification image is alive.
[0189] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: in a case where the target object in the verification picture is non-living, determining the type of the verification picture based on texture features and semantic features extracted from the verification picture, wherein the type of the verification picture includes at least one of the following: a flat photo, a photo of a three-dimensional mask, and a display image in a display device.
[0190] As an optional example, the computer-readable storage medium is also configured to store program code for executing the following steps: obtaining a target image to be subjected to liveness identification, wherein the target image is any one of multiple images of the target object captured by a camera; calling a liveness identification model to extract texture features and high-level semantic features of the target image, wherein the liveness identification model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; based on the texture features and high-level semantic features of the target image, predicting whether the target object is alive.
[0191] As an optional example, the computer-readable storage medium is also configured to store program code for executing the following steps: the display interface of the attendance system receives an attendance request, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera; the attendance system receives texture features and high-level semantic features of the target image fed back based on the attendance request, wherein a liveness recognition model is used to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the attendance system displays the liveness recognition result on the display interface, wherein based on the texture features and high-level semantic features of the target image, it is predicted whether the target object is alive.
[0192] As an optional example, the computer-readable storage medium is also configured to store program codes for executing the following steps: displaying a liveness authentication interface on the payment system, and displaying a target image to be subjected to liveness authentication in the liveness authentication interface, wherein the target image is any one of multiple images of the target object captured by the camera, and the target object is within the authentication area of the liveness authentication interface; the payment system outputs a verification instruction, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; the payment system obtains texture features and high-level semantic features of the target image based on the verification instruction, wherein a liveness recognition model is used to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; the payment system displays the liveness recognition result on the liveness authentication interface, wherein based on the texture features and high-level semantic features of the target image, it predicts whether the target object is live; and when the payment system confirms that the target object is live, it executes a payment operation.
[0193] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0194] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0195] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0196] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0197] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0198] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0199] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that: include: Obtain a sample image, wherein the sample image is any one of multiple images captured by the camera; Extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are concatenated to obtain a target feature map; Inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image; The method also includes: intercepting the face of the target object in the sample image to obtain an image block containing skin color in the face; performing third-order color moment calculation on the image block, and clustering the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
2. The method according to claim 1, characterized in that Extracting a texture feature map of the sample image includes: Determine a region image where the target object is located in the sample image; Identifying texture information in the region image; A central difference convolutional network model is used to perform a pooling operation on the texture information in the area image to generate the texture feature map.
3. The method according to claim 1, characterized in that Extracting a high-level semantic feature map of the sample image includes: An extraction network in which multiple residual modules in a pre-trained model are cascaded is used to analyze the sample image, and high-level semantic information is extracted from the sample image; Based on the high-level semantic information, the high-level semantic feature map is constructed.
4. The method according to claim 3, characterized in that The high-level semantic information includes at least one of the following: facial information and lighting information of the sample image, wherein the high-level semantic feature map is constructed based on the high-level semantic information, including: Training the facial information in the sample image to generate attribute information of the face in the sample image; Training the light information in the sample image to generate illumination information of the face in the sample image; The high-level semantic feature map is constructed based on the attribute information and lighting information of the face in the sample image.
5. The method according to claim 3, characterized in that: After extracting the texture feature map and the high-level semantic feature map of the sample image, the method further includes: A depth estimation and a reflection estimation of the sample image are obtained based on the high-level semantic feature map, wherein MESloss is used as a loss function for the estimation functions of the depth estimation and the reflection estimation.
6. The method according to any one of claims 1 to 5, characterized in that After obtaining the living body recognition model, the method further includes: Obtain a verification image to be used for liveness recognition; The living body recognition model is used to extract the texture features at the bottom layer and the semantic features at the high layer from the verification image; Based on the texture features and the semantic features, predict whether the target object in the verification image is alive.
7. The method according to claim 6, characterized in that In the case where the target object in the verification picture is non-living, the type of the verification picture is determined based on texture features and semantic features extracted from the verification picture, wherein the type of the verification picture includes at least one of the following: a flat photo, a photo of a three-dimensional mask, and a display image in a display device.
8. A method for identifying a living body, characterized in that: include: Obtaining a target image to be subjected to liveness recognition, wherein the target image is any one of a plurality of images of the target object captured by a camera; Calling a liveness recognition model to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; Predicting whether the target object is alive based on the texture features and high-level semantic features of the target image; The method also includes: intercepting the face of the target object in the sample image to obtain an image block containing skin color in the face; performing third-order color moment calculation on the image block, and clustering the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
9. A data processing method, characterized in that: include: The display interface of the attendance system receives an attendance request, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera; The attendance system receives the texture features and high-level semantic features of the target image fed back based on the attendance request, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; The attendance system displays the liveness recognition result on the display interface, wherein based on the texture features and high-level semantic features of the target image, it predicts whether the target object is alive; The method also includes: intercepting the face of the target object in the sample image to obtain an image block containing skin color in the face; performing third-order color moment calculation on the image block, and clustering the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
10. A data processing method, characterized in that: include: The payment system displays a liveness authentication interface, and displays a target image to be subjected to liveness identification on the liveness authentication interface, wherein the target image is any one of a plurality of images of the target object captured by a camera, and the target object is within an authentication area of the liveness authentication interface; The payment system outputs a verification instruction, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; The payment system obtains the texture features and high-level semantic features of the target image based on the verification instruction, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; The payment system displays the liveness recognition result on the liveness authentication interface, wherein based on the texture features and high-level semantic features of the target image, it is predicted whether the target object is live; When the payment system confirms that the target object is alive, the payment operation is performed; The method also includes: intercepting the face of the target object in the sample image to obtain an image block containing skin color in the face; performing third-order color moment calculation on the image block, and clustering the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
11. A data processing device, characterized in that: include: A first acquisition unit is used to acquire a sample image, wherein the sample image is any one of the multiple images collected by the camera; A first extraction unit is used to extract a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; An input unit, used for inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and high-level semantic features are used to perform liveness recognition on the verification image; The device is also used to: intercept the face of the target object in the sample image to obtain an image block containing skin color in the face; calculate the third-order color moment of the image block, and cluster the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
12. A living body identification device, characterized in that: include: A second acquisition unit is used to acquire a target image to be used for living body recognition, wherein the target image is any one of multiple images of the target object captured by a camera; A second extraction unit is used to call a liveness recognition model to extract texture features and high-level semantic features of the target image, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; A prediction unit, configured to predict whether the target object is alive based on texture features and high-level semantic features of the target image; The device is also used to: intercept the face of the target object in the sample image to obtain an image block containing skin color in the face; calculate the third-order color moment of the image block, and cluster the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
13. A data processing device, characterized in that: include: A first receiving unit is used to receive an attendance request through a display interface of the attendance system, wherein the attendance request is used to capture a target image of a target object, and the target image is any one of multiple images of the target object captured by a camera; A second receiving unit is used to receive the texture features and high-level semantic features of the target image fed back based on the attendance request through the attendance system, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, and the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; A first display unit, used for displaying a liveness recognition result on the display interface through the attendance system, wherein whether the target object is a live object is predicted based on the texture features and high-level semantic features of the target image; The device is also used to: intercept the face of the target object in the sample image to obtain an image block containing skin color in the face; calculate the third-order color moment of the image block, and cluster the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
14. A data processing device, characterized in that: include: A second display unit is used to display a liveness authentication interface through the payment system, and to display a target image to be subjected to liveness recognition in the liveness authentication interface, wherein the target image is any one of a plurality of images of a target object captured by a camera, and the target object is within an authentication area of the liveness authentication interface; an output unit, configured to output a verification instruction through the payment system, wherein the verification instruction is used to instruct the target object displayed in the target image to perform a predetermined action indicated by the verification instruction; A third acquisition unit is used to acquire, through the payment system and based on the verification instruction, the texture features and high-level semantic features of the target image, wherein the texture features and high-level semantic features of the target image are extracted using a liveness recognition model, wherein the liveness recognition model is generated by training a target feature map, wherein the target feature map is composed of a texture feature map and a high-level semantic feature map identified from a sample image; a third display unit, configured to display a liveness recognition result on the liveness authentication interface through the payment system, wherein whether the target object is live is predicted based on texture features and high-level semantic features of the target image; an execution unit, configured to execute a payment operation when the payment system confirms that the target object is alive; The device is also used to: intercept the face of the target object in the sample image to obtain an image block containing skin color in the face; calculate the third-order color moment of the image block, and cluster the obtained calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 10.
16. A processor, characterized in that: The processor is used to run a program, wherein the program executes the method according to any one of claims 1 to 10 when being run by the processor.
17. A data processing system, characterized in that: include: processor; A memory connected to the processor, and used to provide the processor with instructions for processing the following processing steps: obtaining a sample image, wherein the sample image is any one of multiple images collected by a camera; extracting a texture feature map and a high-level semantic feature map of the sample image, wherein the texture feature map and the high-level semantic feature map are spliced to obtain a target feature map; inputting the target feature map into a fully connected layer of a neural network model for training to obtain a liveness recognition model, wherein the liveness recognition model is used to extract texture features and high-level semantic features of a verification image, and the texture features and the high-level semantic features are used to perform liveness recognition on the verification image; The memory is also used to provide the processor with instructions for processing the following processing steps: intercepting the face of the target object in the sample image to obtain an image block containing skin color in the face; performing third-order color moment calculation on the image block, and clustering the calculation results as features to obtain clustering results, wherein the clustering results are used to identify living samples and non-living attack samples containing moiré patterns, and the living samples and the non-living attack samples containing moiré patterns are used to train a moiré recognition network, and the moiré recognition network is used to perform moiré detection processing on the verification image predicted by the living recognition model as living.
Citation Information
Patent Citations
Face living body detection method and device based on convolutional neural network
CN111914758A