Face recognition method and apparatus, and electronic device

Generating face recognition codes through the target code generation model solves the inefficiency problem caused by manual research and development, and achieves more efficient and accurate recognition effects.

WO2025148548A1PCT designated stage expired Publication Date: 2025-07-17CHINA TELECOM BESTPAY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135287
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-11-28
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the prior art, facial recognition codes are manually developed to meet the problem of inefficiency in different demand scenarios.

Method used

The syntax tree is generated through the encoder and decoder in the object code generation model, and code snippets for face key point detection, anti-counterfeiting recognition and live detection are generated. Combined with natural language information, it can dynamically adapt to different demand scenarios.

Benefits of technology

It improves the generation efficiency and recognition accuracy of facial recognition codes, avoids inefficient methods of manual research and development, and generates identification codes that are more suitable for the current environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135287_17072025_PF_FP_ABST
    Figure CN2024135287_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and discloses a face recognition method and apparatus, and an electronic device. The method comprises: acquiring an initial face image and natural language information corresponding to a current face recognition requirement scenario; on the basis of the initial face image and the natural language information, generating a grammar tree by means of an encoder in a target code generation model, and obtaining a target grammar tree; on the basis of the target grammar tree and by means of a decoder in the target code generation model, generating a first code snippet, a second code snippet, and a third code snippet, and obtaining target face recognition code; and, on the basis of the target face recognition code, performing face recognition on a subject to be subjected to face recognition, and obtaining a recognition result. The present application solves the problem in the related art of low face recognition efficiency caused by manual research and development of face recognition code to meet different requirement scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Face recognition method, device and electronic equipment

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 10, 2024, with application number 2024100361395 and application name “Face Recognition Method and Device and Electronic Device,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and more specifically, to a face recognition method and device and an electronic device. Background Art

[0003] Liveness detection technology is a primary method for customer identity verification and transaction validation in the financial technology sector and is currently widely used in various internet financial products. Existing liveness detection demand processes are primarily based on changing scenarios, with model design tailored to specific needs. For example, for the identification of key points in daily testing, feature extraction models based on convolutional mechanisms are often used. To integrate natural language input, companies often adopt Transformer-like models to build multimodal input and output structures. Composite models are pre-trained on large-scale datasets and transfer learning is performed on test sets. However, the existing set of recognition models cannot dynamically adapt to various needs, requiring programmers to develop different detection algorithms based on different scenarios. However, this manual approach significantly reduces the efficiency of facial recognition.

[0004] Currently, no effective solution has been proposed to the problem that face recognition efficiency is relatively low due to manual research and development of face recognition codes to meet different demand scenarios in related technologies. Summary of the Invention

[0005] The main purpose of this application is to provide a face recognition method and device and electronic equipment to solve the problem in related technologies that face recognition codes are manually developed to meet different demand scenarios, resulting in relatively low efficiency of face recognition.

[0006] To achieve the above-mentioned purpose, according to one aspect of the present application, a face recognition method is provided. The method includes: obtaining an initial face image and natural language information corresponding to a current face recognition requirement scenario; generating a syntax tree based on the initial face image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree; generating a face key point detection code based on the target syntax tree by a decoder in the target code generation model to obtain a first code fragment; generating a face anti-counterfeiting recognition code based on the target syntax tree by the decoder to obtain a second code fragment; generating a judgment function code for coordinated liveness detection based on the target syntax tree by the decoder to obtain a third code fragment; determining a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; performing face recognition on an object to be face recognized based on the target face recognition code to obtain a recognition result.

[0007] In some embodiments, generating a syntax tree based on the initial facial image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree includes: preprocessing the initial facial image to obtain a target facial image; shearing the natural language information to obtain processed natural language information; performing target character masking on the processed natural language information to obtain target natural language information; generating a syntax tree based on the target facial image and the natural language information by the encoder to obtain the target syntax tree.

[0008] In some embodiments, a syntax tree is generated based on the initial facial image and the natural language information by an encoder in a target code generation model, and obtaining a target syntax tree includes: performing feature extraction on the initial facial image by a feature extraction module in the target code generation model to obtain first feature information; performing feature extraction on the natural language information by the feature extraction module to obtain second feature information; performing feature fusion on the first feature information and the second feature information by the feature extraction module to obtain a target feature vector; and generating a syntax tree based on the target feature vector by the encoder to obtain the target syntax tree.

[0009] In some embodiments, the feature extraction module performs feature fusion on the first feature information and the second feature information to obtain a target feature vector, including: performing dimensional transformation on the first feature information to obtain transformed first feature information; performing dimensional transformation on the second feature information to obtain transformed second feature information; and splicing the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0010] In some embodiments, before the encoder in the target code generation model generates a syntax tree based on the initial face image and the natural language information to obtain the target syntax tree, the method also includes: obtaining a training sample set, wherein the training sample set includes at least: multiple face recognition samples, sample language information corresponding to each face recognition sample, and real face recognition code corresponding to each face recognition sample; performing code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; initializing the first parameter of the encoder and the second parameter of the decoder based on Gaussian distribution to obtain an initial encoder and an initial decoder; determining an initial code generation model based on the initial encoder and the initial decoder; training the initial code generation model based on the sample syntax tree and the training sample set to obtain the target code generation model.

[0011] In some embodiments, training the initial code generation model based on the sample syntax tree and the training sample set to obtain the target code generation model includes: generating a syntax tree by an initial encoder in the initial code generation model based on the multiple face recognition samples and the sample language information corresponding to each face recognition sample to obtain a predicted syntax tree; generating a face recognition code based on the predicted syntax tree by an initial decoder in the initial code generation model to obtain a predicted face recognition code; iteratively training the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; and determining the target code generation model based on the first encoder and the first decoder.

[0012] In some embodiments, determining the target code generation model based on the first encoder and the first decoder includes: randomly occluding characters in a sample syntax tree to obtain a processed sample syntax tree; predicting the occluded characters in the processed sample syntax tree by the first encoder to obtain a prediction result, and iteratively training the first encoder based on the prediction result to obtain a trained first encoder; performing sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; and determining the target code generation model based on the second encoder and the first decoder.

[0013] To achieve the above-mentioned purpose, according to another aspect of the present application, a face recognition device is provided. The device includes: a first acquisition unit for acquiring an initial face image and natural language information corresponding to a current face recognition requirement scenario; a first generation unit for generating a syntax tree based on the initial face image and the natural language information through an encoder in a target code generation model to obtain a target syntax tree; a second generation unit for generating a face key point detection code based on the target syntax tree through a decoder in the target code generation model to obtain a first code fragment; a third generation unit for generating a face anti-counterfeiting recognition code based on the target syntax tree through the decoder to obtain a second code fragment; a fourth generation unit for generating a judgment function code for coordinated liveness detection based on the target syntax tree through the decoder to obtain a third code fragment; a determination unit for determining a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; and a recognition unit for performing face recognition on an object to be face recognized based on the target face recognition code to obtain a recognition result.

[0014] In some embodiments, the first generation unit includes: a first processing module for preprocessing the initial facial image to obtain a target facial image; a second processing module for performing shearing processing on the natural language information to obtain processed natural language information; a third processing module for performing target character masking processing on the processed natural language information to obtain target natural language information; and a generation module for generating a grammar tree based on the target facial image and the natural language information through the encoder to obtain the target grammar tree.

[0015] In some embodiments, the first generation unit includes: a first extraction module, used to perform feature extraction on the initial face image through the feature extraction module in the target code generation model to obtain first feature information; a second extraction module, used to perform feature extraction on the natural language information through the feature extraction module to obtain second feature information; a fusion module, used to perform feature fusion on the first feature information and the second feature information through the feature extraction module to obtain a target feature vector; a generation module, used to generate a syntax tree based on the target feature vector through the encoder to obtain the target syntax tree.

[0016] In some embodiments, the fusion module includes: a first transformation submodule, used to perform dimensional transformation on the first feature information to obtain the transformed first feature information; a second transformation submodule, used to perform dimensional transformation on the second feature information to obtain the transformed second feature information; and a splicing submodule, used to splice the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0017] In some embodiments, the device also includes: a second acquisition unit, used to acquire a training sample set before generating a syntax tree based on the initial face image and the natural language information through the encoder in the target code generation model to obtain the target syntax tree, wherein the training sample set at least includes: multiple face recognition samples, sample language information corresponding to each face recognition sample, and real face recognition code corresponding to each face recognition sample; a parsing unit, used to perform code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; an initialization unit, used to initialize the first parameter of the encoder and the second parameter of the decoder based on Gaussian distribution to obtain an initial encoder and an initial decoder; a determination unit, used to determine the initial code generation model based on the initial encoder and the initial decoder; a training unit, used to train the initial code generation model based on the sample syntax tree and the training sample set to obtain the target code generation model.

[0018] In some embodiments, the training unit includes: a fourth generation module, which is used to generate a syntax tree based on the multiple face recognition samples and the sample language information corresponding to each face recognition sample through the initial encoder in the initial code generation model to obtain a predicted syntax tree; a fifth generation module, which is used to generate a face recognition code based on the predicted syntax tree through the initial decoder in the initial code generation model to obtain a predicted face recognition code; a first training module, which is used to iteratively train the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; a second determination module, which is used to determine the target code generation model based on the first encoder and the first decoder.

[0019] In some embodiments, the second determination module includes: an occlusion module, which is used to randomly occlude the characters in the sample syntax tree to obtain a processed sample syntax tree; a prediction module, which is used to predict the occluded characters in the processed sample syntax tree through the first encoder to obtain a prediction result, and iteratively train the first encoder based on the prediction result to obtain a trained first encoder; a second training module, which is used to perform sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; a third determination module, which is used to determine the target code generation model based on the second encoder and the first decoder.

[0020] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a computer-readable storage medium is provided, which stores a program, wherein when the program is running, the device where the storage medium is located is controlled to execute any one of the above-mentioned face recognition methods.

[0021] In order to achieve the above-mentioned purpose, according to another aspect of the present application, an electronic device is further provided, which includes one or more processors and a memory, and the memory is used to store the face recognition method implemented by one or more processors as described above.

[0022] The present application adopts the following steps: obtaining an initial face image and natural language information corresponding to the current face recognition requirement scenario; generating a syntax tree based on the initial face image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree; generating a face key point detection code based on the target syntax tree by a decoder in the target code generation model to obtain a first code fragment; generating a face anti-counterfeiting recognition code based on the target syntax tree by the decoder to obtain a second code fragment; generating a judgment function code for coordinated liveness detection based on the target syntax tree by the decoder to obtain a third code fragment; determining a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; performing face recognition on an object to be face recognized based on the target face recognition code to obtain a recognition result, which leads to a problem of relatively low efficiency of face recognition. In this solution, the encoder generates a corresponding target syntax tree based on the initial face image and the natural language information corresponding to the current face recognition requirement scenario, and then the decoder generates a face recognition code based on the target syntax tree, avoiding the need for manual code development, and the target code generation model can create a code that is more suitable for the current face recognition environment requirements, thereby achieving the effect of improving face recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0024] FIG1 is a flow chart of a face recognition method according to an embodiment of the present application;

[0025] FIG2 is a first schematic diagram of a face recognition method according to an embodiment of the present application;

[0026] FIG3 is a second schematic diagram of a face recognition method according to an embodiment of the present application;

[0027] FIG4 is a schematic diagram of a face recognition device according to an embodiment of the present application;

[0028] FIG5 is a schematic diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display and analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information fed back by the aforementioned user or organization.

[0033] The present invention is described below in conjunction with preferred implementation steps. FIG1 is a flow chart of a face recognition method provided according to an embodiment of the present application. As shown in FIG1 , the method includes the following steps:

[0034] Step S101: Obtain an initial face image and natural language information corresponding to the current face recognition requirement scenario.

[0035] In some embodiments, an initial face recognition image and text description information corresponding to the face recognition scene currently required (ie, the natural language information mentioned above) are obtained.

[0036] Step S102 : generating a syntax tree based on the initial face image and natural language information by an encoder in the target code generation model to obtain a target syntax tree.

[0037] In some embodiments, the initial facial image and natural language information are input into a target code generation model, and a syntax tree is generated using an encoder in the target code generation model to obtain the target syntax tree.

[0038] In some embodiments, the initial facial image and natural language information may be first feature processed by transform, and then a grammar tree may be generated by an encoder to obtain the target grammar tree.

[0039] Step S103, generating facial key point detection code based on the target syntax tree through a decoder in the target code generation model to obtain a first code fragment;

[0040] Step S104: generating a face anti-counterfeiting recognition code based on the target syntax tree through a decoder to obtain a second code segment;

[0041] Step S105: generating a judgment function code for coordinated liveness detection based on the target syntax tree through the decoder to obtain a third code segment;

[0042] Step S106, determining a target face recognition code based on the first code segment, the second code segment, and the third code segment;

[0043] In some embodiments, after obtaining the target syntax tree, the target syntax tree is parsed and code generated by a decoder in a target code generation model to obtain the target face recognition code.

[0044] In some embodiments, generating face recognition code based on the target syntax tree through the decoder in the target code generation model includes the following steps: the decoder generates face key point detection code according to the target syntax tree, and then generates a judgment function code fragment for face anti-counterfeiting recognition, and generates a judgment function code for coordinated liveness detection, for example, a face blinking judgment function, a mouth opening judgment function, a head shaking judgment function and a nodding judgment function, etc.

[0045] After obtaining the above code fragments, the first code fragment, the second code fragment and the third code fragment can be spliced ​​together according to the above natural language information, and the complete generated code is obtained through the decoder, that is, the above target face recognition code is obtained.

[0046] The above steps improve the generation efficiency of face recognition codes and the accuracy of recognition codes.

[0047] Step S107 , performing face recognition on the object to be face recognized according to the target face recognition code to obtain a recognition result.

[0048] In some embodiments, face recognition is performed on the object to be face recognized using the target face recognition code, thereby improving information security.

[0049] In summary, the encoder generates a corresponding target syntax tree based on the initial face image and the natural language information corresponding to the current face recognition requirement scenario, and then the decoder generates the face recognition code based on the target syntax tree, avoiding the manual implementation of code development. In addition, the target code generation model can create code that is more suitable for the current face recognition environment requirements, thereby achieving the effect of improving face recognition accuracy.

[0050] In some embodiments, in the face recognition method provided in the embodiments of the present application, a syntax tree is generated based on an initial face image and natural language information by an encoder in a target code generation model, and obtaining a target syntax tree includes: preprocessing the initial face image to obtain a target face image; shearing the natural language information to obtain processed natural language information; performing target character masking processing on the processed natural language information to obtain target natural language information; and generating a syntax tree based on the target face image and natural language information by an encoder to obtain a target syntax tree.

[0051] In some embodiments, when an encoder in a target code generation model generates a syntax tree based on an initial facial image and natural language information, the following steps may be performed: preprocessing the initial facial image, for example, scaling the initial facial image [F, W, H, 3] proportionally, where F is the number of frames, W / H is the length and width of the frame image, and 3 is the number of RGB channels per frame image. For another example, the dimensions of the initial facial image may be changed to facilitate subsequent matching with the natural language information.

[0052] Then, operations such as cutting and masking special characters are performed on the natural language information to convert the natural language information into a sequence of equal length composed of tokens (ie, the target natural language information mentioned above).

[0053] Finally, the encoder generates a syntax tree based on the target face image and natural language information to obtain the target syntax tree.

[0054] In some embodiments, in the face recognition method provided in the embodiments of the present application, a syntax tree is generated based on an initial face image and natural language information by an encoder in a target code generation model, and obtaining a target syntax tree includes: performing feature extraction on the initial face image by a feature extraction module in the target code generation model to obtain first feature information; performing feature extraction on the natural language information by the feature extraction module to obtain second feature information; performing feature fusion on the first feature information and the second feature information by the feature extraction module to obtain a target feature vector; and generating a syntax tree based on the target feature vector by an encoder to obtain a target syntax tree.

[0055] In some embodiments, generating a syntax tree based on an initial facial image and natural language information by an encoder in a target code generation model further includes the following steps: extracting features from the initial facial image and extracting features from the natural language information by a feature extraction module (e.g., a transform module) in the target code generation model to obtain a text feature vector (i.e., the second feature information described above) and a picture feature vector (i.e., the first feature information described above). Then, aggregating heterogeneous information of the picture and text, i.e., performing feature fusion on the first feature information and the second feature information by the feature extraction module to obtain a target feature vector, and finally generating a syntax tree based on the target feature vector by an encoder to obtain a target syntax tree.

[0056] By fusing the text feature information and the image feature information through the above steps, the accuracy of the subsequent generated syntax tree is improved.

[0057] In some embodiments, in the face recognition method provided in the embodiments of the present application, the first feature information and the second feature information are subjected to feature fusion through a feature extraction module to obtain a target feature vector, including: performing a dimensional transformation on the first feature information to obtain the transformed first feature information; performing a dimensional transformation on the second feature information to obtain the transformed second feature information; and performing splicing processing on the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0058] In some embodiments, feature information fusion can be achieved by the following steps: first, the first feature information and the second feature information are dimensional transformed so that the dimensions of the transformed first feature information and the transformed second feature information are the same, and then the transformed first feature information and the transformed second feature information are spliced ​​to obtain the target feature vector.

[0059] In some embodiments, feature information fusion can also be achieved by the following steps: performing a product operation on the first feature information and the second feature information to obtain an embedding matrix (embedding matrix), and performing sum pooling processing on the embedding matrix to obtain a target feature vector.

[0060] In some embodiments, feature information fusion can also be achieved by the following steps: adding the first feature information and the second feature information to obtain an initial feature vector, and processing the initial feature vector through a self-attention mechanism to obtain a target feature vector.

[0061] The above method can better integrate feature information so that the syntax tree can be generated more accurately in the future.

[0062] In some embodiments, in the face recognition method provided in the embodiments of the present application, a face recognition code is generated based on a target syntax tree by a decoder in a target code generation model, and obtaining the target face recognition code includes: generating a face key point detection code based on the target syntax tree by a decoder to obtain a first code fragment; generating a face anti-counterfeiting recognition code based on the target syntax tree by a decoder to obtain a second code fragment; generating a coordinated liveness detection judgment function code based on the target syntax tree by a decoder to obtain a third code fragment; and determining the target face recognition code based on the first code fragment, the second code fragment, and the third code fragment.

[0063] In some embodiments, in the face recognition method provided in the embodiments of the present application, before the encoder in the target code generation model generates a syntax tree based on the initial face image and natural language information to obtain the target syntax tree, the method also includes: obtaining a training sample set, wherein the training sample set includes at least: multiple face recognition samples, sample language information corresponding to each face recognition sample, and real face recognition code corresponding to each face recognition sample; performing code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; initializing the first parameter of the encoder and the second parameter of the decoder based on Gaussian distribution to obtain an initial encoder and an initial decoder; determining an initial code generation model based on the initial encoder and the initial decoder; training the initial code generation model based on the sample syntax tree and the training sample set to obtain a target code generation model.

[0064] In some embodiments, a training sample set is first obtained based on a plurality of face recognition samples, sample language information corresponding to each face recognition sample, and a real face recognition code corresponding to each face recognition sample. For example, a live face detection sample (I 1 ), the corresponding natural language description (I 2 ), code snippet y code .

[0065] Then, according to the above code snippet (i.e. the above real face recognition code), the corresponding sample syntax tree is obtained. And the first parameter φ of the encoder and the second parameter θ of the decoder are initialized according to the regularized Gaussian distribution to obtain the above initial encoder and initial decoder. For example, the encoder Decoder The above-mentioned initial code generation model is determined based on the initial encoder and the initial decoder.

[0066] Finally, the initial code generation model is trained based on the sample syntax tree and the training sample set to obtain the target code generation model.

[0067] In some embodiments, the following steps can be used to obtain the above-mentioned sample syntax tree: use a scanner to read code snippets and merge them into a sequence of identification tokens according to predetermined rules. At the same time, the scanner will remove whitespace, comments, etc. during the merging process. Finally, the entire code will be divided into tokens vectors. Then a parser can be used for syntax analysis. It will convert the array obtained by lexical analysis of the code snippet into a tree-like expression. At the same time, the syntax is verified, and if the syntax is wrong, a syntax error is thrown. When generating the tree, the parser will delete some unnecessary identification tokens (such as incomplete brackets), and finally obtain the above-mentioned sample syntax tree.

[0068] In some embodiments, the living body detection portrait data I 1 Use the torchvision open source library for unified preprocessing and set the corresponding parameters of the composite function so that each frame of the image has a unified channel, size and representation format after processing. Use the transform image preprocessing module to normalize the pixel values ​​of each frame of the image after processing. 2 The preprocessing includes cutting into fixed length, masking into fixed sequence, and then training the model with the processed sample data.

[0069] The above steps improve the accuracy of face recognition code generation by the target code generation model.

[0070] In some embodiments, in the face recognition method provided in the embodiments of the present application, the initial code generation model is trained based on the sample syntax tree and the training sample set to obtain the target code generation model, including: generating a syntax tree based on multiple face recognition samples and sample language information corresponding to each face recognition sample by the initial encoder in the initial code generation model to obtain a predicted syntax tree; generating a face recognition code based on the predicted syntax tree by the initial decoder in the initial code generation model to obtain a predicted face recognition code; iteratively training the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; and determining the target code generation model based on the first encoder and the first decoder.

[0071] In some embodiments, the initial code generation model is trained based on the sample syntax tree and the training sample set to obtain the target code generation model, including the following steps: inputting multiple face recognition samples and the sample language information corresponding to each face recognition sample into the initial encoder, generating a syntax tree through the initial encoder to obtain a predicted syntax tree; and inputting the predicted syntax tree into the initial decoder, generating a predicted face recognition code through the initial decoder, and then iteratively training the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder, and then obtaining the final target code generation model.

[0072] In some embodiments, the input triple data distribution (live detection image, natural language description of task requirements, and parsed code tree): x:={(I 1 , I 2 ,y code )},z:={(y tree )},y code For real code, y tree

[0073] is the parsed code book tree. Then, the encoder is The decoder is The loss function is calculated based on the KL divergence principle for the encoding and decoding parameters θ and φ:

[0074] Among them, q φ (z|x) is when x:={(I 1 , I 2 ,y code )}, generate z:={(y tree )}, p θ (x, z) = p θ (x|z)p θ (z), pθ (x|z) when z:={(y tree )} under the premise of generating y code The probability, p θ (z) is to generate z: = {(y tree )}, the above q can be obtained by predicting the syntax tree, predicting the face recognition code, the sample syntax tree and the real face recognition code φ (z|x), p θ (x, z), etc. Represents q φ The expectation of (z|x). D KL Represents divergence. It can be solved by maximizing the expectation. At the same time, in order to make the second term derivative (the sampling process cannot be derivative, and auxiliary parameters are needed, assuming that the AST parameters in the latent space are generated by a standard Gaussian distribution, the second term of the loss function is transformed into:

[0075] Among them, p θ (z|x) obeys Gaussian distribution μ is the mean, σ 2 is the variance. Through training with the above loss function, the above target code generation model is obtained.

[0076] In some embodiments, in the face recognition method provided in the embodiments of the present application, determining the target code generation model based on the first encoder and the first decoder includes: randomly occluding characters in the sample syntax tree to obtain a processed sample syntax tree; predicting the occluded characters in the processed sample syntax tree by the first encoder to obtain a prediction result, and iteratively training the first encoder based on the prediction result to obtain a trained first encoder; performing sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; and determining the target code generation model based on the second encoder and the first decoder.

[0077] In some embodiments, in order to improve the accuracy of subsequent generation of face recognition codes, the encoder can be trained a second time: the characters in the sample syntax tree are randomly masked, for example, the MLM (Masked Language Model) method is used to randomly mask certain positions in the input code sequence, and then the masked characters in the processed sample syntax tree are predicted by the first encoder. Since the MLM prediction task enables the results obtained by model encoding to also include contextual information of the context, it is conducive to training a deeper and broader encoder.

[0078] Based on the above training, the trained first encoder is trained on the next sentence prediction task (NSP) of the code snippet to obtain the above second encoder. Finally, the target code generation model is determined based on the second encoder and the first decoder.

[0079] In some embodiments, a schematic diagram of generating face recognition code through a target code generation model is shown in Figure 2. A face image and recognition requirements, for example, face key point detection under lighting conditions, are input, and code snippet 1: lighting judgment and code snippet 2: key point detection code are generated through the target code generation model.

[0080] In some embodiments, face recognition code generation can be achieved by using the second schematic diagram of the target code generation model shown in FIG3 , where the face image and recognition requirements are input into the encoder. In the coder, we get the grammar tree (Grammer Tree), and the decoder Generate the corresponding face recognition code based on the syntax tree.

[0081] The face recognition method provided in the embodiment of the present application obtains an initial face image and natural language information corresponding to the current face recognition requirement scenario; generates a syntax tree based on the initial face image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree; generates a face key point detection code based on the target syntax tree by a decoder in the target code generation model to obtain a first code fragment; generates a face anti-counterfeiting recognition code based on the target syntax tree by the decoder to obtain a second code fragment; generates a judgment function code for coordinated liveness detection based on the target syntax tree by the decoder to obtain a third code fragment; determines a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; performs face recognition on an object to be face recognized based on the target face recognition code to obtain a recognition result, thereby solving the problem in the related art of manually developing face recognition codes to meet different requirement scenarios, resulting in relatively low efficiency of face recognition. In this solution, the encoder generates a corresponding target syntax tree based on the initial face image and the natural language information corresponding to the current face recognition requirement scenario, and then the decoder generates the face recognition code based on the target syntax tree, avoiding the manual implementation of code development. The target code generation model can create code that is more suitable for the current face recognition environment requirements, thereby achieving the effect of improving face recognition accuracy.

[0082] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0083] The embodiment of the present application also provides a face recognition device. It should be noted that the face recognition device of the embodiment of the present application can be used to execute the face recognition method provided in the embodiment of the present application. The face recognition device provided in the embodiment of the present application is introduced below.

[0084] Figure 4 is a schematic diagram of a face recognition device according to an embodiment of the present application. As shown in Figure 4, the device includes: a first acquisition unit 401, a first generation unit 402, a second generation unit 403, a third generation unit 404, a fourth generation unit 405, a determination unit 406, and a recognition unit 407.

[0085] The first acquisition unit 401 is used to acquire the initial face image and the natural language information corresponding to the current face recognition requirement scenario;

[0086] The first generating unit 402 is configured to generate a syntax tree based on the initial face image and natural language information by an encoder in a target code generation model to obtain a target syntax tree;

[0087] The second generating unit 403 is configured to generate a facial key point detection code based on the target syntax tree by using a decoder in the target code generation model to obtain a first code segment;

[0088] The third generating unit 404 is configured to generate a face anti-counterfeiting recognition code based on the target syntax tree through a decoder to obtain a second code segment;

[0089] The fourth generating unit 405 is configured to generate a judgment function code for cooperative liveness detection based on the target syntax tree through a decoder to obtain a third code segment;

[0090] A determination unit 406 is configured to determine a target face recognition code based on the first code segment, the second code segment, and the third code segment;

[0091] The recognition unit 407 is configured to perform face recognition on the object to be face recognized according to the target face recognition code to obtain a recognition result.

[0092] The face recognition device provided in the embodiment of the present application obtains the initial face image and the natural language information corresponding to the current face recognition requirement scenario through the first acquisition unit 401; the first generation unit 402 generates a syntax tree based on the initial face image and the natural language information through the encoder in the target code generation model to obtain the target syntax tree; the second generation unit 403 is used to generate a face key point detection code based on the target syntax tree through the decoder in the target code generation model to obtain a first code fragment; the third generation unit 404 generates a face anti-counterfeiting recognition code based on the target syntax tree through the decoder to obtain a second code fragment; the fourth generation unit 405 generates a judgment function code for coordinated liveness detection based on the target syntax tree through the decoder to obtain a third code fragment; the determination unit 406 determines the target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; the recognition unit 407 performs face recognition on the object to be face recognized based on the target face recognition code to obtain a recognition result, which solves the problem in the related art that the face recognition code is manually developed to meet different requirement scenarios, resulting in relatively low efficiency of face recognition. In this solution, the encoder generates a corresponding target syntax tree based on the initial face image and the natural language information corresponding to the current face recognition requirement scenario, and then the decoder generates the face recognition code based on the target syntax tree, avoiding the manual implementation of code development. The target code generation model can create code that is more suitable for the current face recognition environment requirements, thereby achieving the effect of improving face recognition accuracy.

[0093] In some embodiments, in the face recognition device provided in the embodiments of the present application, the first generation unit includes: a first processing module, used to preprocess the initial face image to obtain a target face image; a second processing module, used to perform shearing processing on the natural language information to obtain processed natural language information; a third processing module, used to perform target character masking processing on the processed natural language information to obtain target natural language information; a generation module, used to generate a syntax tree based on the target face image and natural language information through an encoder to obtain a target syntax tree.

[0094] In some embodiments, in the face recognition device provided in the embodiments of the present application, the first generation unit includes: a first extraction module, used to perform feature extraction on the initial face image through the feature extraction module in the target code generation model to obtain first feature information; a second extraction module, used to perform feature extraction on the natural language information through the feature extraction module to obtain second feature information; a fusion module, used to perform feature fusion on the first feature information and the second feature information through the feature extraction module to obtain a target feature vector; a generation module, used to generate a syntax tree based on the target feature vector through an encoder to obtain a target syntax tree.

[0095] In some embodiments, in the face recognition device provided in the embodiments of the present application, the fusion module includes: a first transformation submodule, used to perform dimensional transformation on the first feature information to obtain the transformed first feature information; a second transformation submodule, used to perform dimensional transformation on the second feature information to obtain the transformed second feature information; and a splicing submodule, used to splice the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0096] In some embodiments, in the face recognition device provided in the embodiments of the present application, the device also includes: a second acquisition unit, used to obtain a training sample set before obtaining a target syntax tree by generating a syntax tree based on the initial face image and natural language information through an encoder in a target code generation model, wherein the training sample set includes at least: multiple face recognition samples, sample language information corresponding to each face recognition sample, and a real face recognition code corresponding to each face recognition sample; a parsing unit, used to perform code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; an initialization unit, used to initialize the first parameter of the encoder and the second parameter of the decoder based on a Gaussian distribution to obtain an initial encoder and an initial decoder; a determination unit, used to determine the initial code generation model based on the initial encoder and the initial decoder; a training unit, used to train the initial code generation model based on the sample syntax tree and the training sample set to obtain a target code generation model.

[0097] In some embodiments, in the face recognition device provided in the embodiments of the present application, the training unit includes: a fourth generation module, which is used to generate a syntax tree based on multiple face recognition samples and sample language information corresponding to each face recognition sample through an initial encoder in an initial code generation model to obtain a predicted syntax tree; a fifth generation module, which is used to generate a face recognition code based on the predicted syntax tree through an initial decoder in the initial code generation model to obtain a predicted face recognition code; a first training module, which is used to iteratively train the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; a second determination module, which is used to determine the target code generation model based on the first encoder and the first decoder.

[0098] In some embodiments, in the face recognition device provided in the embodiments of the present application, the second determination module includes: an occlusion module, which is used to randomly occlude the characters in the sample syntax tree to obtain a processed sample syntax tree; a prediction module, which is used to predict the occluded characters in the processed sample syntax tree through the first encoder to obtain a prediction result, and iteratively train the first encoder based on the prediction result to obtain a trained first encoder; a second training module, which is used to perform sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; and a third determination module, which is used to determine the target code generation model based on the second encoder and the first decoder.

[0099] The face recognition device includes a processor and a memory. The above-mentioned first acquisition unit 401, first generation unit 402, second generation unit 403, third generation unit 404, fourth generation unit 405, determination unit 406 and recognition unit 407 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.

[0100] The processor contains a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the accuracy of face recognition can be improved by adjusting the kernel parameters.

[0101] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0102] An embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which implements a face recognition method when executed by a processor.

[0103] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes a face recognition method when it is run.

[0104] As shown in Figure 5, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the processor executes the program, the following steps are implemented: obtaining an initial face image and natural language information corresponding to the current face recognition requirement scenario; generating a syntax tree based on the initial face image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree; generating a face key point detection code based on the target syntax tree by a decoder in the target code generation model to obtain a first code fragment; generating a face anti-counterfeiting recognition code based on the target syntax tree by the decoder to obtain a second code fragment; generating a judgment function code for coordinated liveness detection based on the target syntax tree by the decoder to obtain a third code fragment; determining a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; performing face recognition on the object to be face recognized based on the target face recognition code to obtain a recognition result.

[0105] In some embodiments, generating a syntax tree based on an initial facial image and natural language information by an encoder in a target code generation model to obtain a target syntax tree includes: preprocessing the initial facial image to obtain a target facial image; shearing the natural language information to obtain processed natural language information; performing target character masking on the processed natural language information to obtain target natural language information; generating a syntax tree based on the target facial image and natural language information by an encoder to obtain a target syntax tree.

[0106] In some embodiments, a syntax tree is generated based on an initial facial image and natural language information by an encoder in a target code generation model, and obtaining a target syntax tree includes: performing feature extraction on the initial facial image by a feature extraction module in the target code generation model to obtain first feature information; performing feature extraction on the natural language information by the feature extraction module to obtain second feature information; performing feature fusion on the first feature information and the second feature information by the feature extraction module to obtain a target feature vector; and generating a syntax tree based on the target feature vector by an encoder to obtain a target syntax tree.

[0107] In some embodiments, the feature extraction module performs feature fusion on the first feature information and the second feature information to obtain a target feature vector, including: performing dimensional transformation on the first feature information to obtain the transformed first feature information; performing dimensional transformation on the second feature information to obtain the transformed second feature information; and splicing the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0108] In some embodiments, before the encoder in the target code generation model generates a syntax tree based on the initial face image and natural language information to obtain the target syntax tree, the method also includes: obtaining a training sample set, wherein the training sample set includes at least: multiple face recognition samples, sample language information corresponding to each face recognition sample, and real face recognition code corresponding to each face recognition sample; performing code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; initializing the first parameter of the encoder and the second parameter of the decoder based on Gaussian distribution to obtain an initial encoder and an initial decoder; determining an initial code generation model based on the initial encoder and the initial decoder; training the initial code generation model based on the sample syntax tree and the training sample set to obtain a target code generation model.

[0109] In some embodiments, the initial code generation model is trained based on the sample syntax tree and the training sample set to obtain the target code generation model, including: generating a syntax tree based on multiple face recognition samples and sample language information corresponding to each face recognition sample by an initial encoder in the initial code generation model to obtain a predicted syntax tree; generating a face recognition code based on the predicted syntax tree by an initial decoder in the initial code generation model to obtain a predicted face recognition code; iteratively training the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; and determining the target code generation model based on the first encoder and the first decoder.

[0110] In some embodiments, determining the target code generation model based on the first encoder and the first decoder includes: randomly occluding characters in the sample syntax tree to obtain a processed sample syntax tree; predicting the occluded characters in the processed sample syntax tree by the first encoder to obtain a prediction result, and iteratively training the first encoder based on the prediction result to obtain a trained first encoder; performing sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; and determining the target code generation model based on the second encoder and the first decoder.

[0111] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0112] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialized program having the following method steps: obtaining an initial face image and natural language information corresponding to a current face recognition requirement scenario; generating a syntax tree based on the initial face image and the natural language information by an encoder in a target code generation model to obtain a target syntax tree; generating a face key point detection code based on the target syntax tree by a decoder in the target code generation model to obtain a first code fragment; generating a face anti-counterfeiting recognition code based on the target syntax tree by the decoder to obtain a second code fragment; generating a judgment function code for coordinated liveness detection based on the target syntax tree by the decoder to obtain a third code fragment; determining a target face recognition code based on the first code fragment, the second code fragment, and the third code fragment; performing face recognition on an object to be face recognized based on the target face recognition code to obtain a recognition result.

[0113] In some embodiments, generating a syntax tree based on an initial facial image and natural language information by an encoder in a target code generation model to obtain a target syntax tree includes: preprocessing the initial facial image to obtain a target facial image; shearing the natural language information to obtain processed natural language information; performing target character masking on the processed natural language information to obtain target natural language information; generating a syntax tree based on the target facial image and natural language information by an encoder to obtain a target syntax tree.

[0114] In some embodiments, a syntax tree is generated based on an initial facial image and natural language information by an encoder in a target code generation model, and obtaining a target syntax tree includes: performing feature extraction on the initial facial image by a feature extraction module in the target code generation model to obtain first feature information; performing feature extraction on the natural language information by the feature extraction module to obtain second feature information; performing feature fusion on the first feature information and the second feature information by the feature extraction module to obtain a target feature vector; and generating a syntax tree based on the target feature vector by an encoder to obtain a target syntax tree.

[0115] In some embodiments, the feature extraction module performs feature fusion on the first feature information and the second feature information to obtain a target feature vector, including: performing dimensional transformation on the first feature information to obtain the transformed first feature information; performing dimensional transformation on the second feature information to obtain the transformed second feature information; and splicing the transformed first feature information and the transformed second feature information to obtain the target feature vector.

[0116] In some embodiments, before the encoder in the target code generation model generates a syntax tree based on the initial face image and natural language information to obtain the target syntax tree, the method also includes: obtaining a training sample set, wherein the training sample set includes at least: multiple face recognition samples, sample language information corresponding to each face recognition sample, and real face recognition code corresponding to each face recognition sample; performing code parsing based on the real face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; initializing the first parameter of the encoder and the second parameter of the decoder based on Gaussian distribution to obtain an initial encoder and an initial decoder; determining an initial code generation model based on the initial encoder and the initial decoder; training the initial code generation model based on the sample syntax tree and the training sample set to obtain a target code generation model.

[0117] In some embodiments, the initial code generation model is trained based on the sample syntax tree and the training sample set to obtain the target code generation model, including: generating a syntax tree based on multiple face recognition samples and sample language information corresponding to each face recognition sample by an initial encoder in the initial code generation model to obtain a predicted syntax tree; generating a face recognition code based on the predicted syntax tree by an initial decoder in the initial code generation model to obtain a predicted face recognition code; iteratively training the initial encoder and the initial decoder based on the predicted syntax tree, the predicted face recognition code, the sample syntax tree and the real face recognition code to obtain a first encoder and a first decoder; and determining the target code generation model based on the first encoder and the first decoder.

[0118] In some embodiments, determining the target code generation model based on the first encoder and the first decoder includes: randomly occluding characters in the sample syntax tree to obtain a processed sample syntax tree; predicting the occluded characters in the processed sample syntax tree by the first encoder to obtain a prediction result, and iteratively training the first encoder based on the prediction result to obtain a trained first encoder; performing sentence prediction training on the trained first encoder based on the real face recognition code corresponding to each face recognition sample to obtain a second encoder; and determining the target code generation model based on the second encoder and the first decoder.

[0119] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0120] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0121] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0123] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0124] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0125] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0126] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0127] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A face recognition method, comprising: Obtaining an initial face image and natural language information corresponding to the current face recognition requirement scenario; Generating a syntax tree based on the initial face image and the natural language information through an encoder in a target code generation model to obtain a target syntax tree; Generating face key point detection code based on the target syntax tree through a decoder in the target code generation model to obtain a first code segment; Generating face anti-counterfeiting recognition code based on the target syntax tree through the decoder to obtain a second code segment; Generating a judgment function code for cooperative live detection based on the target syntax tree through the decoder to obtain a third code segment; Determining target face recognition code according to the first code segment, the second code segment, and the third code segment; Performing face recognition on an object to be face recognized according to the target face recognition code to obtain a recognition result.

2. The method according to claim 1, wherein, Generating a syntax tree based on the initial face image and the natural language information through an encoder in a target code generation model to obtain a target syntax tree, including: Preprocessing the initial face image to obtain a target face image; Performing clipping processing on the natural language information to obtain processed natural language information; Performing target character masking processing on the processed natural language information to obtain target natural language information; Generating a syntax tree based on the target face image and the natural language information through the encoder to obtain the target syntax tree.

3. The method according to claim 1, wherein, Generating a syntax tree based on the initial face image and the natural language information through an encoder in a target code generation model to obtain a target syntax tree, including: Performing feature extraction on the initial face image through a feature extraction module in the target code generation model to obtain first feature information; Performing feature extraction on the natural language information through the feature extraction module to obtain second feature information; Performing feature fusion on the first feature information and the second feature information through the feature extraction module to obtain a target feature vector; Generating a syntax tree based on the target feature vector through the encoder to obtain the target syntax tree.

4. The method according to claim 3, wherein Performing feature fusion on the first feature information and the second feature information through the feature extraction module to obtain a target feature vector, including: Performing dimensionality transformation on the first feature information to obtain transformed first feature information; Performing dimensionality transformation on the second feature information to obtain transformed second feature information; Performing splicing processing on the transformed first feature information and the transformed second feature information to obtain the target feature vector.

5. The method according to claim 1, wherein Before generating a syntax tree based on the initial face image and the natural language information through an encoder in a target code generation model to obtain a target syntax tree, the method further includes: Obtaining a training sample set, where the training sample set at least includes: multiple face recognition samples, sample language information corresponding to each face recognition sample, and true face recognition code corresponding to each face recognition sample; Performing code parsing according to the true face recognition code to obtain a sample syntax tree corresponding to each face recognition sample; Initialize the first parameter of the encoder and the second parameter of the decoder according to the Gaussian distribution to obtain an initial encoder and an initial decoder; Determine an initial code generation model based on the initial encoder and the initial decoder; Train the initial code generation model according to the sample syntax tree and the training sample set to obtain the target code generation model.

6. The method according to claim 5, wherein Training the initial code generation model according to the sample syntax tree and the training sample set to obtain the target code generation model includes: Generate a syntax tree through the initial encoder in the initial code generation model based on the multiple face recognition samples and the sample language information corresponding to each face recognition sample to obtain a predicted syntax tree; Generate a face recognition code through the initial decoder in the initial code generation model based on the predicted syntax tree to obtain a predicted face recognition code; Iteratively train the initial encoder and the initial decoder according to the predicted syntax tree, the predicted face recognition code, the sample syntax tree, and the true face recognition code to obtain a first encoder and a first decoder; Determine the target code generation model based on the first encoder and the first decoder.

7. The method according to claim 6, wherein Determining the target code generation model based on the first encoder and the first decoder includes: Randomly occlude the characters in the sample syntax tree to obtain a processed sample syntax tree; Predict the occluded characters in the processed sample syntax tree through the first encoder to obtain a prediction result, and iteratively train the first encoder according to the prediction result to obtain a trained first encoder; Perform statement prediction training on the trained first encoder according to the true face recognition code corresponding to each face recognition sample to obtain a second encoder; Determine the target code generation model based on the second encoder and the first decoder.

8. A face recognition device, comprising: A first acquisition unit, configured to acquire an initial face image and natural language information corresponding to a current face recognition requirement scenario; A first generation unit, configured to generate a syntax tree through the encoder in the target code generation model based on the initial face image and the natural language information to obtain a target syntax tree; A second generation unit, configured to generate a face key point detection code through the decoder in the target code generation model based on the target syntax tree to obtain a first code segment; A third generation unit, configured to generate a face anti-counterfeiting recognition code through the decoder based on the target syntax tree to obtain a second code segment; A fourth generation unit, configured to generate a judgment function code for cooperative live detection through the decoder based on the target syntax tree to obtain a third code segment; A determination unit, configured to determine a target face recognition code according to the first code segment, the second code segment, and the third code segment; An identification unit, configured to perform face recognition on an object to be face recognized according to the target face recognition code to obtain an identification result.

9. A computer-readable storage medium, including a stored program, wherein, Control the storage medium to execute the face recognition method according to any one of claims 1 to 7 when the program is running.

10. An electronic device includes one or more processors and a memory, where the memory is used to store one or more programs, and wherein, When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the face recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fine-grained code automatic generation method and system based on multi-view code features

    CN113342318A

  • Dual Bayesian Encoding-Decoding Techniques For Text-to-Code Transform

    CN115964029A

  • Living body detection method and device, storage medium and electronic equipment

    CN116798129A

  • Face recognition method and device and electronic equipment

    CN117539452A