Facial recognition intention identification methods, devices, electronic devices and storage media

By acquiring two-dimensional and three-dimensional images to generate multimodal human body segmentation feature maps, the problem of mistakenly scanning other people's faces in facial recognition payment has been solved, improving the security and accuracy of facial recognition payment.

CN116110093BActive Publication Date: 2026-05-26ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2022-12-01
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In public places, the collection of multiple images can lead to mistaken identity theft, which reduces the security of facial recognition payment and necessitates improving the accuracy of facial recognition intent identification.

Method used

By acquiring two-dimensional and three-dimensional images of the target object, a multimodal human body segmentation feature map is generated. The mask map is used to distinguish the face region from other regions. The two-dimensional and three-dimensional modal feature maps are combined to confirm the face recognition intention result.

Benefits of technology

It improves the accuracy of facial recognition intention identification, reduces the chance of mistakenly scanning someone else's face, and enhances the security of facial recognition payment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110093B_ABST
    Figure CN116110093B_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, electronic device, and storage medium for facial recognition intent identification. The method includes: acquiring a two-dimensional image and a three-dimensional image of a target detection object for a facial recognition transaction; generating a multimodal human body segmentation feature map based on the two-dimensional image and the three-dimensional image; generating a mask image corresponding to the two-dimensional image based on the two-dimensional image and a first face region of the target detection object in the two-dimensional image to distinguish the first face region from other regions in the two-dimensional image; obtaining a two-dimensional modality fusion feature map based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map; obtaining a three-dimensional modality fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map; and confirming the facial recognition intent result of the target detection object for the facial recognition transaction based on the two-dimensional modality fusion feature map and the three-dimensional modality fusion feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, electronic device and storage medium for facial recognition intention identification. Background Technology

[0002] With the development of computer and internet technology, people's payment methods have also changed. The use of facial recognition as an identity verification method has brought great convenience to users and is widely loved by users.

[0003] However, in public places, such as when using offline facial recognition payment devices, there are often scenarios where users queue to pay. Because these are open public spaces, if multiple people are present in the image captured by the facial recognition device, there's a possibility that user A's facial recognition might mistakenly register user B's face. If the facial recognition payment system recognizes user B's face in the scene when user A initiates payment, user B could suffer financial loss, thus reducing the security of facial recognition payments. Therefore, in offline facial recognition services, facial recognition intent detection is a crucial step in ensuring facial recognition security and improving the user experience. Based on this, there is an urgent need to propose a more secure facial recognition intent detection method. Summary of the Invention

[0004] The main purpose of this specification is to provide a method, device, electronic device, and storage medium for facial recognition intent identification, aiming to improve the accuracy of facial recognition intent prediction. The technical solution is as follows:

[0005] Firstly, the embodiments of this specification provide a method for facial recognition intention identification, including:

[0006] Obtain 2D and 3D images of the target object for facial recognition transactions;

[0007] Based on the two-dimensional image and the three-dimensional image, a multimodal human body segmentation feature map is generated;

[0008] Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask image corresponding to the two-dimensional image is generated. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region.

[0009] Based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained;

[0010] Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained;

[0011] Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the facial recognition intention of the target detection object for the facial recognition transaction is confirmed.

[0012] Secondly, embodiments of this specification provide a training method for a facial recognition intention identification model, including:

[0013] Collect sample two-dimensional images and sample three-dimensional images of the face recognition transaction as training datasets, and label the sample three-dimensional images and the intention classification labels of the sample detection objects in the sample three-dimensional images and the sample human body regions of the sample detection objects to generate the first label of the training dataset;

[0014] Randomly sample from the training dataset to generate training batch data and the second annotation labels corresponding to the training batch data;

[0015] The training batch data is used as the input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value.

[0016] The multimodal segmentation loss value corresponding to the human target segmentation probability map and the face recognition intention classification probability value are calculated using the face recognition intention recognition loss function;

[0017] The initialization parameters of the face recognition intention model are adjusted based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and the trained face recognition intention model is generated based on the target parameters.

[0018] Thirdly, embodiments of this specification provide a facial recognition intention identification device, including:

[0019] The acquisition module is used to acquire two-dimensional and three-dimensional images of the target detection object for facial recognition transactions;

[0020] The human body segmentation module is used to generate a multimodal human body segmentation feature map based on the two-dimensional image and the three-dimensional image;

[0021] A two-dimensional modal feature generation module is used to generate a mask image corresponding to the two-dimensional image based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region.

[0022] A two-dimensional modal feature fusion module is used to obtain a two-dimensional modal fusion feature map based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map;

[0023] A three-dimensional modal feature generation module is used to obtain a three-dimensional modal fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map;

[0024] The face recognition intention determination module is used to confirm the face recognition intention result of the target detection object for the face recognition transaction based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map.

[0025] Fourthly, embodiments of this specification provide a facial recognition intention identification device, comprising:

[0026] The sample acquisition module is used to acquire sample two-dimensional images and sample three-dimensional images of the sample detection object for the transaction as training datasets, and to annotate the sample three-dimensional images and the intention classification labels of the sample detection object and the sample human body regions of the sample detection object in the sample three-dimensional images, thereby generating the first annotation label of the training dataset;

[0027] The sampling module is used to randomly sample from the training dataset to generate training batch data and the second annotation label corresponding to the batch data;

[0028] The prediction module is used to take the training batch data as input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value.

[0029] The loss calculation module is used to calculate the multimodal segmentation loss value corresponding to the human target segmentation probability map and the face recognition intention loss value corresponding to the face recognition intention classification probability value using the face recognition intention recognition loss function;

[0030] The training module is used to adjust the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generate the trained face recognition intention model based on the target parameters.

[0031] Fifthly, embodiments of this specification provide an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described above.

[0032] Sixthly, embodiments of this specification provide a storage medium storing a computer program, wherein the facial recognition intention recognition program, when executed by a processor, implements the steps of the method described above.

[0033] In a seventh aspect, embodiments of this specification provide a computer program product, comprising: a computer program that, when executed by a processor of an electronic device, enables the processor to at least implement the methods described in the first to fourth aspects.

[0034] In the embodiments of this specification, two-dimensional and three-dimensional images of the target detection object for facial recognition transactions are acquired. Based on the two-dimensional and three-dimensional images, a multimodal human body segmentation feature map is generated. Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask map corresponding to the two-dimensional image is generated to distinguish the first face region from other regions in the two-dimensional image. Based on the two-dimensional image, the mask map, and the multimodal human body segmentation feature map, a two-dimensional modality fusion feature map is obtained. Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modality fusion feature map is obtained. Based on the two-dimensional and three-dimensional modality fusion feature maps, the facial recognition intention of the target detection object for the facial recognition transaction is confirmed. By fusing the features extracted from the human target segmentation task with 2D and 3D modal features respectively, the network focuses more on learning the behavior of people in the image. Then, by fusing and learning the 2D and 3D multimodal features, accurate prediction of facial recognition intention is achieved. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is an example schematic diagram of a facial recognition intention recognition method provided in the embodiments of this specification;

[0037] Figure 2 This is a flowchart illustrating a face recognition intention identification method provided in the embodiments of this specification;

[0038] Figure 3 This is a schematic diagram of the overall process of a face recognition intention identification method provided in the embodiments of this specification;

[0039] Figure 4 This is a detailed flowchart illustrating a face recognition intention identification method provided in the embodiments of this specification;

[0040] Figure 5 This is a detailed flowchart illustrating a face recognition intention identification method provided in the embodiments of this specification;

[0041] Figure 6 This is a flowchart illustrating a training method for a facial recognition intention identification model provided in the embodiments of this specification;

[0042] Figure 7This is a schematic diagram of the structure of a facial recognition intention recognition model provided in the embodiments of this specification;

[0043] Figure 8 This is a schematic diagram of the structure of a facial recognition intention identification device provided in the embodiments of this specification;

[0044] Figure 9 This is a schematic diagram of the structure of a facial recognition intention identification device provided in the embodiments of this specification;

[0045] Figure 10 This is a structural schematic diagram of a facial recognition intention identification device provided in the embodiments of this specification. Detailed Implementation

[0046] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0047] Furthermore, it should be noted that the acquisition, storage, use, and processing of data involved in the embodiments of this specification, such as face recognition and face scan intention recognition, all comply with the relevant provisions of national laws and regulations.

[0048] The facial recognition intention recognition device can be a terminal device such as a mobile phone, computer, tablet, smart wearable device, or in-vehicle device, or it can be a module in a terminal device used to implement the facial recognition intention recognition method. The facial recognition intention recognition device can acquire two-dimensional and three-dimensional images of the target detection object for the facial recognition transaction. Based on the two-dimensional image and the three-dimensional image, a multimodal human body segmentation feature map is generated. Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask map corresponding to the two-dimensional image is generated to distinguish the first face region from other regions in the two-dimensional image. Based on the two-dimensional image, the mask map, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained. Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained. Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the facial recognition intention result of the target detection object for the facial recognition transaction is confirmed.

[0049] Correspondingly, the facial recognition intention recognition device can also train a facial recognition intention recognition model. The facial recognition intention recognition model can collect sample 2D and 3D images of the target object for facial recognition transactions as training datasets, and label the 3D images with intention classification labels and the target object's human body region, generating the first label of the training dataset. Random sampling is then performed from the training dataset to generate training batch data and corresponding second label. This training batch data is used as input to the facial recognition intention recognition model to obtain a sample human body target segmentation probability map and a sample facial recognition intention classification probability value. A facial recognition intention recognition loss function is used to calculate the multimodal segmentation loss value corresponding to the human body target segmentation probability map and the facial recognition intention loss value corresponding to the facial recognition intention classification probability value. Based on the multimodal segmentation loss value and the facial recognition intention loss value, the initialization parameters of the facial recognition intention recognition model are adjusted to obtain target parameters. Finally, the trained facial recognition intention recognition model is generated based on the target parameters.

[0050] It should be noted that the face recognition device used for face recognition intention identification and the face recognition intention identification device used for face recognition intention identification model training can be the same device or different devices. Preferably, the face recognition intention identification method and the face recognition intention identification model training method are implemented on different face recognition intention identification devices.

[0051] Please see also Figure 1 This document provides an example schematic diagram of a face recognition intention identification method. After acquiring two-dimensional and three-dimensional images of the target detection object for face recognition transactions, the face recognition intention identification device generates a multimodal human body segmentation feature map based on the two-dimensional and three-dimensional images. Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask image corresponding to the two-dimensional image is generated. Based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained. Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained. Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the face recognition intention result of the target detection object for face recognition transactions is confirmed.

[0052] The facial recognition intention recognition method provided in this specification will be described in detail below with reference to specific embodiments.

[0053] Please see Figure 2 This is a flowchart illustrating a facial recognition intention identification method provided in an embodiment of this specification. Figure 2 As shown, the method described in the embodiments of this specification may include the following steps S102-S112.

[0054] S102, acquire the two-dimensional and three-dimensional images of the target detection object for face recognition transactions;

[0055] Specifically, facial recognition image acquisition devices collect two-dimensional and three-dimensional images of the target object. For example, offline IoT devices can collect two-dimensional and three-dimensional images of the target object for facial recognition transactions. These transactions can include facial recognition payment, facial recognition access control, and facial recognition attendance. Two-dimensional images are planar images, lacking three-dimensional information and a sense of depth, while three-dimensional images carry three-dimensional information and can directly or indirectly represent a three-dimensional image. A typical facial recognition 3D image is a depth image, which combines 2D and depth information to represent 3D. The following examples mainly use depth images as illustrations. Depth images, also known as distance images, are images that use the actual distance (depth) from the image acquisition device to various points in the scene as pixel values. They directly reflect the geometry of the visible surfaces of objects.

[0056] Both 2D and 3D images must include at least the face of the target object. It's understandable that, since facial recognition is typically performed using public devices, such as for queuing at stores, the 2D and 3D images captured in this scenario also include other people and the background besides the target object. Therefore, human detection can be used to identify all human bodies in the 2D image and then select the target object from among them. For example, the distance of all human bodies from the camera can be calculated, and the object closest to the camera can be selected as the target object; alternatively, based on the position of each human body, the human body located in the center of the image can be identified as the target object. Alternatively, facial recognition prompts can be used to guide the user to place their head or face within a preset frame, and the user within the preset frame can be selected as the target object. It should be noted that the method for determining the target object is not limited.

[0057] S104, Based on the two-dimensional image and the three-dimensional image, generate a multimodal human body segmentation feature map;

[0058] Understandably, since two-dimensional and three-dimensional images also include other people and backgrounds besides the target object, and the process of facial recognition mainly focuses on human behavior, while the image contains both human body regions and background regions, and the background regions vary greatly, the human body segmentation task makes the network pay more attention to the human body region and reduce the interference of the background region, thus improving the accuracy of learning.

[0059] Specifically, a pre-trained feature encoding network can be used to extract features from two-dimensional and three-dimensional images. The extracted multimodal human segmentation feature map represents pixels in the image that belong to humans and those that do not. Commonly used recognition networks such as ResNet, VGG, and ShuffleNetV2 can be used for the feature encoding network.

[0060] S106, Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask image corresponding to the two-dimensional image is generated. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region.

[0061] In this approach, the region where the target object is located in the 2D image is referred to as the first face region. The first face region contains the appearance feature information of the target object, at least facial information, and may also include torso and limb information if necessary. Since a 2D image may contain the faces of multiple users and is subject to background interference, this scheme specifically targets the face of the target object for facial recognition payment intent identification. This is achieved by detecting the first face region in the 2D image. Correspondingly, all other regions in the 2D image are designated as other regions. By identifying the first face region in the 2D image and generating a corresponding mask image, the face portion can be prioritized for analysis during training, reducing processing time and increasing accuracy.

[0062] S108, Based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained;

[0063] Specifically, a 2D image mask is obtained based on the 2D image and the mask image, distinguishing the first face region from other regions. Features of the 2D image mask are then extracted using a pre-trained neural network model. These features can include facial features, torso features, and limb features. The mask image is created by performing a masking operation on each pixel in the 2D image. A mask kernel operator recalculates the values ​​of each pixel in the image, characterizing the influence of neighboring pixels on the new pixel value. Simultaneously, a weighted average is applied to the pixels based on the weighting factors in the mask operator, thereby highlighting the key face region. The multimodal human segmentation feature map is then further fused with the 2D image fusion features to generate a 2D modal feature map fused with human segmentation, further enhancing the facial region features of the target object in the 2D image.

[0064] S110, Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained;

[0065] It is understandable that when performing facial recognition on two-dimensional images, the recognition results are affected by external factors such as lighting, posture, and expression, which affects the accuracy of facial recognition and thus the accuracy of facial payment intention recognition. Three-dimensional images contain richer information, such as optical flow information. Therefore, it is necessary to make full use of the two-dimensional and three-dimensional information of the face for multimodal recognition and improve the recognition accuracy based on multimodal visual information.

[0066] Specifically, features can be extracted from 3D images to obtain 3D image feature maps, which are then fused with multimodal human body segmentation maps to obtain 3D modal fusion feature maps, thereby enhancing the facial region features of the target object in the 3D image.

[0067] S112, based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, confirm the target detection object's willingness to use facial recognition for the facial recognition transaction.

[0068] Specifically, by leveraging the complementarity of two types of features—two-dimensional modal fusion feature maps and three-dimensional modal fusion feature maps—the corresponding target detection object is represented more accurately. This allows for improved accuracy in determining facial recognition intent based on these two types of feature maps. In some feasible implementations, machine learning models can be used to identify facial recognition intent. For example, the two-dimensional and three-dimensional modal fusion feature maps can be input into a pre-learned supervised facial recognition intent recognition model. This model performs high-dimensional mapping of the target detection object's feature information and outputs the mapping result. Then, the mapping result is converted into a facial recognition intent probability value, which is used to determine whether the target detection object intends to use facial recognition.

[0069] For example, if the probability value is greater than the preset probability threshold, it can be considered that the target object has the willingness to use facial recognition for the transaction; if the probability value is less than or equal to the preset probability threshold, it can be considered that the target object does not have the willingness to use facial recognition for the transaction and may be the facial recognition transaction triggered by someone else.

[0070] Understandably, once the facial recognition intent result of the target detection object is obtained, the facial recognition intent recognition device, or the device corresponding to the facial recognition intent recognition device, can be controlled to execute the facial recognition transaction based on the facial recognition intent result. For example, if the facial recognition transaction is facial recognition payment, the process involves the user facing the camera on the screen of the facial recognition payment device. The facial recognition payment system automatically acquires the user's facial information and performs facial recognition intent recognition. If the user has the intent to use facial recognition, the system identifies the payment account associated with the user and completes the payment. If the user does not have the intent to use facial recognition, the facial recognition payment request is blocked.

[0071] In the embodiments of this specification, two-dimensional and three-dimensional images of the target detection object for facial recognition transactions are acquired. Based on the two-dimensional and three-dimensional images, a multimodal human body segmentation feature map is generated. Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask map corresponding to the two-dimensional image is generated. A two-dimensional modality fusion feature map is obtained based on the two-dimensional image, the mask map, and the multimodal human body segmentation feature map. Simultaneously, a three-dimensional modality fusion feature map is obtained based on the three-dimensional image and the multimodal human body segmentation feature map. Based on the two-dimensional and three-dimensional modality fusion feature maps, the facial recognition intent of the target detection object is confirmed. By introducing the human target segmentation task into intent recognition, the human body region in the acquired multimodal images is segmented, and non-human background regions are removed. This allows the network to focus more on learning the payment behavior of human targets in the multimodal images, thereby improving the accuracy of facial recognition intent determination for the target detection object.

[0072] Please see Figure 3 This diagram illustrates a detailed process for a facial recognition intention identification method provided in this specification. Figure 3 As shown, the method described in the embodiments of this specification may include the following steps S202-S212.

[0073] S202, stitch the two-dimensional image and the three-dimensional image together;

[0074] S204, extract the features of the stitched two-dimensional image and the three-dimensional image to generate a multimodal human body segmentation feature map;

[0075] Specifically, a multimodal image obtained by directly stitching together two-dimensional and three-dimensional images is used as a multi-channel input. Channel-by-channel fusion is then performed, and the fused feature representation is obtained through an encoding network. Compared to single-modal images, multimodal images help extract features from different views, providing complementary information and contributing to better data representation and network discrimination capabilities. Therefore, utilizing multimodal images can reduce information uncertainty and improve the accuracy of human body segmentation.

[0076] S206, based on the face region of the target detection object in the two-dimensional image, determine the first filling region of the mask image corresponding to the target detection object, and the second filling region outside the first filling region;

[0077] S208, assign a first fill value to the first filled region and a second fill value to the second filled region to generate the mask image with the same resolution as the two-dimensional image.

[0078] In some feasible implementations, a mask image is generated based on the position of the selected target detection object in the acquired image (face region position). The mask image represents the position of the target detection object in the two-dimensional image, and the mask image needs to include at least or all of the face region of the target detection object. The resolution of the mask image is consistent with that of the acquired two-dimensional image. Specifically, firstly, the width w = x2 - x1 and the height h = y2 - y1 of the face bounding box are calculated based on the position (x1, y1, x2, y2) of the target detection object in the two-dimensional image. Then, the radius R = max(w / 2, h / 2) of the circular attention region is determined. A circular region with radius R centered at the center of the face bounding box ((x1+x2) / 2, (y1+y2) / 2) is filled with 1s, and other background regions (second filling regions) are filled with 0s.

[0079] Optionally, mask images of shapes such as ellipses and squares can be generated based on the position of the face bounding box of the target object in the two-dimensional image. The specific shape can be selected according to actual needs.

[0080] S210, Extract the features of the two-dimensional image;

[0081] S212, fuse the features of the two-dimensional image and the mask image to obtain the first fused feature;

[0082] Specifically, features of a 2D image can be extracted using a convolutional network. These features, along with a mask image corresponding to the target object, are then input into a pre-built attention mechanism network, outputting a first fused feature. Commonly used convolutional networks include ResNet, VGG, and ShuffleNetV2. This first fused feature enhances image contrast, thus focusing attention on the target object with the intention to perform facial recognition.

[0083] S214, according to the channel dimension, the first fusion feature and the multimodal human body segmentation feature map are connected to obtain a two-dimensional modal fusion feature map.

[0084] Specifically, according to the channel dimension, the first fusion feature and the multimodal human segmentation feature are connected, and the connected features are input into a convolutional network for fusion processing to obtain a two-dimensional modal fusion feature map.

[0085] Assuming the multimodal human segmentation feature map extracted by the human target segmentation network is 512-dimensional, and the two-dimensional image feature extracted by the two-dimensional modal feature extraction network is 256-dimensional (provided that w and h of the multimodal human segmentation feature map and the first fusion feature map are consistent), the two are concatenated to obtain 256+512 channels. Based on the 768 channels, a convolutional kernel is used to concatenate the two features.

[0086] Please see Figure 4 This diagram illustrates a detailed process for a facial recognition intention identification method provided in this specification. Figure 4 As shown, the method described in the embodiments of this specification may include the following steps S216-S218.

[0087] S216, Extract the features of the three-dimensional image;

[0088] Specifically, pre-trained convolutional networks are used to extract 3D image features. Commonly used convolutional networks include ResNet, VGG, and ShuffleNetV2.

[0089] Optionally, before extracting features from the 3D image, the pixel depth values ​​in the 3D image can be normalized. For example, pixels in the 3D image that are more than a preset threshold away from the image acquisition device can be identified, and these pixels can be uniformly generalized, such as unifying the values ​​representing these pixels to a single specified value. The aim is to reduce the differences between these pixels and their contribution to model training and use, thereby allowing the model to focus more on pixels within the preset threshold, concentrating computational power on more valuable pixels, improving efficiency, and reducing interference.

[0090] S218, according to the channel latitude, connect the features of the three-dimensional image and the multimodal human body segmentation feature map to obtain a three-dimensional modal fusion feature map.

[0091] Specifically, the extracted features of the 3D image and the multimodal human body segmentation feature map are concatenated according to the channel dimension, which is the same as the feature concatenation method of the 2D image and the multimodal human body segmentation feature map mentioned above, and will not be elaborated here.

[0092] Please see Figure 5 This diagram illustrates a detailed process for a facial recognition intention identification method provided in this specification. Figure 5 As shown, the method described in the embodiments of this specification may include the following steps S220-S224.

[0093] S220, Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, a multimodal fusion feature is obtained;

[0094] Optionally, after obtaining the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map are connected according to the channel dimension, that is, they are added pixel by pixel to obtain the multimodal fusion feature.

[0095] S222, Based on the multimodal fusion features, obtain the probability value of face recognition willingness;

[0096] Specifically, the multimodal fusion features are input into a pre-trained convolutional network for processing to obtain the probability value of facial recognition payment intention, including two possible cases: intention to pay securely and intention to pay insecurely.

[0097] S224, Based on the face recognition willingness probability value and the preset probability threshold, confirm the face recognition willingness result of the target detection object for the face recognition transaction.

[0098] Specifically, the probability value of the willingness to pay via facial recognition is compared with a preset probability threshold. If the probability value is higher than the preset threshold, the target object is considered to have the willingness to pay via facial recognition. Similarly, if the probability value is lower than or equal to the preset threshold, the target object is considered not to have the willingness to pay via facial recognition.

[0099] In the embodiments of this specification, a multimodal human body segmentation feature map is obtained by extracting features from the stitched 2D and 3D images. The extracted multimodal human body segmentation feature map focuses more on the human body region in the image, enabling the network to focus more on learning the behavior of human targets in the image and achieve accurate prediction of face recognition intention. Furthermore, a mask image is generated based on the first face region of the target detection object, and then features are extracted from the 2D image. By adding new channels to the features of the 2D image, the features of the mask image are represented by the newly added channels, thus obtaining the first fused feature. This allows for more targeted learning of the target detection object region in the first fused feature. Simultaneously, features are extracted from the 3D image, and then the features of the 3D image and the 2D image are linked with the multimodal human body segmentation feature map to obtain a 2D modal fusion feature map and a 3D modal fusion feature map. Further, based on the 2D modal fusion feature map and the 3D modal fusion feature map, the multimodal fusion feature is obtained to predict the probability value of face recognition intention, thus obtaining the face recognition intention result. By leveraging the segmentation information provided by multimodal human body segmentation feature maps, attention learning mechanisms for 2D and 3D modalities are offered, enabling multimodal task consistency learning and improving the accuracy of facial recognition intent identification for target detection objects.

[0100] Please see Figure 6 This is a schematic diagram illustrating the training process of a facial recognition intention recognition model provided in an embodiment of this specification. Figure 6 As shown, the method described in the embodiments of this specification may include the following steps S302-S310.

[0101] S302, Collect sample two-dimensional images and sample three-dimensional images of the sample detection objects for face recognition transactions as training datasets, and label the sample three-dimensional images and the intention classification labels of the sample detection objects and the sample human body regions of the sample detection objects in the sample three-dimensional images, and generate the first label of the training dataset;

[0102] Specifically, first set the initial structural parameters of the facial recognition intention identification model, such as the model's network structure parameters and loss function. The model's network structure can adopt the existing network's number and type of intermediate layers (fully connected, dropout layers, normalized layers, convolutional layers, etc.), number of neurons per layer, and activation function. The loss function can be set according to actual needs, and this specification does not impose specific limitations on the embodiments.

[0103] Next, a training dataset is created. Specifically, sample 2D and 3D images of the objects to be detected for facial recognition transactions are collected. The objects to be detected can be different people. 2D and 3D images of the objects to be detected are collected using IoT devices. Regardless of whether multiple people are present in the collected images, one person is selected and labeled as a positive sample (indicating safety). For example, if only one person is present in the collected images, user A is selected and labeled as a positive sample (indicating safety). If multiple people are present in the collected images, one person is selected and labeled as a positive sample, and the remaining users are selected as negative samples. For example, user A is selected and user B is selected and labeled as a negative sample (indicating insecurity). Simultaneously, the human body regions of the objects to be detected in the collected 2D and 3D images are labeled.

[0104] The training dataset includes: a face recognition image I captured by a camera. The first annotation labels include: the position (x1, y1, x2, y2) of the face bounding box of the face recognition user selected from image I for comparison and identification; and a label {0, 1} indicating whether the selected user is willing to participate, where 1 represents willingness and 0 represents involuntary participation. Each pixel in the labeled image I is binary classified into human region and non-human region, where pixel category 1 represents human region and 0 represents non-human region.

[0105] S304, Randomly sample from the training dataset to generate training batch data and the second annotation label corresponding to the training batch data;

[0106] Random sampling involves randomly selecting data from a given set of data, with each sample having an equal probability of being selected. Training batch data, or training data, is obtained from the training dataset through random sampling, and the corresponding labeled data is then acquired.

[0107] S306, The training batch data is used as the input of the face recognition intention recognition model to obtain the sample human body target segmentation probability map and the sample face recognition intention classification probability value.

[0108] Specifically, the training batch data is input into the facial recognition intention recognition model, and the network in the model obtains the sample human target segmentation probability map and the sample facial recognition intention classification probability value. The sample human target segmentation probability map and the sample facial recognition intention classification probability value are obtained through different networks.

[0109] S308, The face recognition intention recognition loss function is used to calculate the multimodal segmentation loss value corresponding to the human target segmentation probability map and the face recognition intention classification probability value.

[0110] Specifically, a pre-defined loss function for facial recognition intent is used to calculate the multimodal segmentation loss value corresponding to the human target segmentation probability map. Then, a pre-defined loss function for facial recognition intent is used to calculate the facial recognition intent classification probability value corresponding to the facial recognition intent loss value. By calculating the difference between the predicted human target segmentation probability map and the facial recognition intent classification probability value and the true value, the next step of training is guided in the correct direction.

[0111] S310, adjust the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generate the trained face recognition intention model based on the target parameters.

[0112] Specifically, the loss function is calculated by outputting the network with the labels corresponding to the training batch. Based on the calculated multimodal segmentation loss value and face recognition intention loss value, the network is trained using gradient descent to obtain the final model target parameters. Then, the trained face recognition intention recognition model is generated based on the target parameters.

[0113] Optionally, in one embodiment, the method includes:

[0114] The sample 2D images and sample 3D images in the training batch data are input into the multimodal human target segmentation network, and the sample human target segmentation probability map is output.

[0115] Based on the third face region where the sample detection object is located in the sample two-dimensional image and the sample two-dimensional image, a corresponding sample mask image is generated;

[0116] The sample 2D image and the sample mask map are input into a 2D feature extraction network to obtain sample 2D features. The sample 2D features and the sample human target segmentation probability map are then input into a 2D feature fusion network to output a sample 2D fusion feature map.

[0117] Based on the pixel depth value of the fourth face region where the sample detection object is located in the sample 3D image, the pixel value of the sample 3D image is normalized to obtain a normalized sample 3D image.

[0118] The normalized sample 3D image is input into a 3D feature extraction network to obtain sample 3D features. The sample 3D features and the sample human target segmentation probability map are input into a 3D feature fusion network to output a sample 3D fusion feature map.

[0119] The sample's two-dimensional modality fusion feature map and the sample's three-dimensional modality fusion feature map are input into a multimodal fusion intention prediction network, which outputs the sample's face recognition intention classification probability value.

[0120] Specifically, refer to Figure 7 , Figure 7 This is a schematic diagram of a face recognition intention recognition model provided in an embodiment of this specification. The face recognition intention recognition model includes a multimodal human target segmentation network, a two-dimensional feature extraction network, a two-dimensional feature fusion network, a three-dimensional feature extraction network, a three-dimensional feature fusion network, and a multimodal fusion intention prediction network. The multimodal human target segmentation network includes an encoding network module and a decoding network module, used to perform human segmentation tasks on 2D and 3D images. The encoding network module extracts human segmentation feature maps from the 2D and 3D images, and inputs the human segmentation feature maps into the decoding module to obtain a human target segmentation probability map. The 2D feature extraction network is used to extract features from the 2D image and mask image to obtain 2D features, and then the 2D features and the human target segmentation probability map are fused through the 2D feature fusion network. The 3D feature extraction network is used to extract features from the normalized depth image to obtain 3D features, and the 3D feature fusion network is used to fuse the 3D features and the human target segmentation probability map. The multimodal fusion intention prediction network includes a multimodal fusion convolutional network module and a face recognition intention result prediction convolutional network module. The multimodal fusion convolutional network module can be composed of one or more convolutional network modules. The output of the multimodal fusion convolutional network module is a multimodal fusion feature, and then the face recognition intention result prediction network module, which is composed of one or more convolutional network modules, outputs the face recognition intention classification probability value.

[0121] After obtaining the sample 3D image, the pixel values ​​of the sample 3D image can be normalized. Specifically, features of the 3D image are extracted. First, based on the face region selection box of the target detection object in the 2D image, the face region of the target detection object in the 3D image is determined. For example, the position coordinates of the face region selection box in the 2D image are extracted, and then based on the position coordinates, the face region selection box is mapped to the 3D image, thereby obtaining the fourth face region of the target detection object in the 3D image.

[0122] Then, multiple depth values ​​are calculated for all pixels in the fourth face region. Finally, the depth value of the target object's face region in the 3D image is determined based on the average of these multiple depth values. For example, the average of multiple depth values ​​is used as the depth value of the face region. In other words, the average of the depth measurement results in the face region is taken as the depth value of the face.

[0123] Furthermore, when filtering pixels that are more than a preset threshold away from the image acquisition device based on the depth value of the fourth face region, for the convenience of data processing, the depth value of the fourth face region is used as a benchmark. After processing, the multiple depth values ​​of all pixels are mapped to the required specified range to better reflect the differences between pixels and facilitate subsequent calculations. The processed multiple depth values ​​can be filtered to filter pixels that are more than a preset threshold away from the image acquisition device.

[0124] Specifically, the ratio between the multiple depth values ​​of all pixels and the depth value of the face region is calculated first. Since the depth value of the face region is obtained based on the average of the multiple depth values ​​of all pixels, the depth value of the face region is obtained.

[0125] Then, based on the baseline value and ratio of the depth values ​​of the face region, the multiple depth values ​​of all pixels are processed to be within the range of the baseline value. The baseline value of the depth values ​​of the face region can be set according to actual needs. For example, setting the baseline value of the depth values ​​of the face region to 127 means that, assuming multiple depth values ​​are all positive, the data distribution of the multiple depth values ​​of all processed pixels will mostly be between 0 and 254.

[0126] Optionally, in one embodiment, the face recognition intention recognition loss function includes a multimodal segmentation loss function and a face recognition intention loss function; the step of using the face recognition intention recognition loss function to calculate the multimodal segmentation loss value corresponding to the sample human target segmentation probability map and the face recognition intention loss value corresponding to the sample face recognition intention classification probability value includes:

[0127] Based on the sample human target segmentation probability map and the corresponding human target segmentation label in the second annotation label, the multimodal segmentation loss value is calculated using the multimodal segmentation loss function;

[0128] Based on the sample face recognition willingness classification probability value and the corresponding willingness classification label in the second annotation label, the face recognition willingness loss value is calculated using the face recognition willingness loss function.

[0129] Specifically, the multimodal segmentation loss function and the face recognition intention loss function can adopt the cross-entropy loss function, respectively calculating the difference between the sample human target segmentation probability map and the real human target segmentation label, and the difference between the sample face recognition intention classification probability value and the real intention classification label, to obtain the multimodal segmentation loss value and the face recognition intention loss value.

[0130] Optionally, in one embodiment, adjusting the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generating the trained face recognition intention model based on the target parameters, includes:

[0131] The total loss value is calculated based on the face recognition intention loss value, the balance coefficient corresponding to the face recognition intention loss value, and the multimodal segmentation loss value;

[0132] Based on the total loss value, the initialization parameters of the face recognition intention model are iteratively adjusted using gradient descent. Then, the process proceeds to the step of using the training batch data as input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value, until the total loss value meets the preset convergence condition, and the trained target parameters are obtained. Based on the target weights, the trained face recognition intention model is generated.

[0133] Specifically, since the facial recognition intention identification model has two parts, and this solution focuses more on the intention prediction part, the total loss function for network training is L = L2 + λL1, where L1 is the multimodal segmentation loss function, L2 is the facial recognition intention loss function, and λ is the balancing coefficient. In practical applications, λ < 1 to ensure that the overall network training is more inclined towards learning the main task of intention prediction. Gradient descent is used to train the network without interrupting the optimization of the loss function.

[0134] In the embodiments of this specification, sample two-dimensional images and sample three-dimensional images of the sample detection object for face recognition transactions are collected as training datasets. The intention classification labels of the sample three-dimensional images and the sample human body regions of the sample detection objects in the sample three-dimensional images are labeled to generate the first label of the training dataset. Then, random sampling is performed from the training dataset to generate training batch data and the corresponding second label of the training batch data. The training batch data is used as input to the face recognition intention recognition model to obtain the sample human body target segmentation probability map and the sample face recognition intention classification probability value. The face recognition intention recognition loss function is used to calculate the multimodal segmentation loss value corresponding to the human body target segmentation probability map and the face recognition intention loss value corresponding to the face recognition intention classification probability value. Based on the multimodal segmentation loss value and the face recognition intention loss value, the initialization parameters of the face recognition intention recognition model are adjusted to obtain the target parameters. The trained face recognition intention recognition model is generated based on the target parameters. By introducing the human target segmentation task into the end-to-end learning network, the human body region in the acquired image is segmented and the non-human background region is removed, so that the network focuses more on learning the payment behavior of human targets in the image. Then, the two-dimensional and three-dimensional multimodal features are fused and trained to obtain the face recognition intention recognition model, so as to achieve accurate prediction of face recognition intention.

[0135] The following will be combined with the appendix Figure 8-9 This specification provides a detailed description of the facial recognition intention identification device provided in the embodiments. It should be noted that the appendix... Figure 8 The facial recognition device described herein is used to perform the functions described in this manual. Figures 2-5 The methods shown in the embodiments are illustrated for ease of explanation, showing only the parts related to the embodiments of this specification. For specific technical details not disclosed, please refer to this specification. Figures 2-5 The example shown.

[0136] Please see Figure 8 This diagram illustrates the structure of a facial recognition intention identification device provided in an exemplary embodiment of this specification. The facial recognition intention identification device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes an acquisition module 11, a human body segmentation module 12, a two-dimensional modal feature generation module 13, a two-dimensional modal feature fusion module 14, a three-dimensional modal feature generation module 15, and a facial recognition intention determination module 16.

[0137] The acquisition module 11 is used to acquire two-dimensional and three-dimensional images of the target detection object for face recognition transactions;

[0138] The human body segmentation module 12 is used to generate a multimodal human body segmentation feature map based on the two-dimensional image and the three-dimensional image;

[0139] The two-dimensional modal feature generation module 13 is used to generate a mask image corresponding to the two-dimensional image based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region.

[0140] The two-dimensional modal feature fusion module 14 is used to obtain a two-dimensional modal fusion feature map based on the two-dimensional image, the mask map, and the multimodal human body segmentation feature map;

[0141] The three-dimensional modal feature generation module 15 is used to obtain a three-dimensional modal fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map;

[0142] The face recognition intention determination module 16 is used to confirm the face recognition intention result of the target detection object for the face recognition transaction based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map.

[0143] Optionally, the human body segmentation module 12 is specifically used to stitch the two-dimensional image and the three-dimensional image;

[0144] Features are extracted from the stitched 2D and 3D images to generate a multimodal human body segmentation feature map.

[0145] Optionally, the two-dimensional modal feature generation module 13 is specifically used to determine the first filling region of the mask image corresponding to the target detection object and the second filling region outside the first filling region based on the face region of the target detection object in the two-dimensional image;

[0146] A first fill value is assigned to the first filled region, and a second fill value is assigned to the second filled region to generate a mask map with a resolution consistent with that of the two-dimensional image.

[0147] Optionally, the two-dimensional modal feature fusion module 14 is specifically used to extract features from the two-dimensional image;

[0148] The features of the two-dimensional image and the mask image are fused to obtain the first fused feature;

[0149] According to the channel dimension, the first fusion feature and the multimodal human body segmentation feature map are connected to obtain a two-dimensional modal fusion feature map.

[0150] Optionally, the three-dimensional modal feature generation module 15 is specifically used to extract features from the three-dimensional image;

[0151] According to the channel latitude, the features of the three-dimensional image and the multimodal human body segmentation feature map are connected to obtain a three-dimensional modal fusion feature map.

[0152] Optionally, the face recognition intention determination module 16 is specifically used to obtain multimodal fusion features based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map;

[0153] Based on the multimodal fusion features, the probability value of face recognition willingness is obtained;

[0154] Based on the face recognition willingness probability value and the preset probability threshold, the face recognition willingness result of the target detection object for the face recognition transaction is confirmed.

[0155] Optionally, the face recognition intention determination module 16 is specifically used to connect the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map according to the channel dimension to obtain multimodal fusion features.

[0156] Further, refer to the appendix Figure 9 The facial recognition device shown is equipped with... Figure 9 The facial recognition device described herein is used to perform the functions described in this manual. Figure 6-7 The methods shown in the embodiments are illustrated for ease of explanation, showing only the parts related to the embodiments of this specification. For specific technical details not disclosed, please refer to this specification. Figure 6-7 The example shown.

[0157] Please see Figure 9 This diagram illustrates the structure of a facial recognition intention identification device provided in an exemplary embodiment of this specification. This facial recognition intention identification device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 2 includes a sample acquisition module 21, a sampling module 22, a prediction module 23, a loss calculation module 24, and a training module 25.

[0158] The sample acquisition module 21 is used to acquire sample two-dimensional images and sample three-dimensional images of the sample detection object for the transaction as training datasets, and to label the sample three-dimensional images and the intention classification labels of the sample detection object and the sample human body regions of the sample detection object in the sample three-dimensional images, thereby generating the first label of the training dataset;

[0159] Sampling module 22 is used to randomly sample from the training dataset to generate training batch data and the second annotation label corresponding to the batch data;

[0160] The prediction module 23 is used to take the training batch data as input to the face recognition intention recognition model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value.

[0161] Loss calculation module 24 is used to calculate the multimodal segmentation loss value corresponding to the human target segmentation probability map and the face recognition intention loss value corresponding to the face recognition intention classification probability value using the face recognition intention recognition loss function;

[0162] Training module 25 is used to adjust the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generate the trained face recognition intention model based on the target parameters.

[0163] Optionally, the prediction module 23 is specifically used to input the sample two-dimensional images and sample three-dimensional images in the training batch data into the multimodal human target segmentation network and output the sample human target segmentation probability map;

[0164] Based on the third face region where the sample detection object is located in the sample two-dimensional image and the sample two-dimensional image, a corresponding sample mask image is generated;

[0165] The sample 2D image and the sample mask map are input into a 2D feature extraction network to obtain sample 2D features. The sample 2D features and the sample human target segmentation probability map are then input into a 2D feature fusion network to output a sample 2D fusion feature map.

[0166] Based on the pixel depth value of the fourth face region where the sample detection object is located in the sample 3D image, the pixel value of the sample 3D image is normalized to obtain a normalized sample 3D image.

[0167] The normalized sample 3D image is input into a 3D feature extraction network to obtain sample 3D features. The sample 3D features and the sample human target segmentation probability map are input into a 3D feature fusion network to output a sample 3D fusion feature map.

[0168] The sample's two-dimensional modality fusion feature map and the sample's three-dimensional modality fusion feature map are input into a multimodal fusion intention prediction network, which outputs the sample's face recognition intention classification probability value.

[0169] Optionally, the loss calculation module 24 is specifically used to calculate the multimodal segmentation loss value based on the sample human target segmentation probability map and the corresponding human target segmentation label in the second annotation label, using a multimodal segmentation loss function;

[0170] Based on the sample face recognition willingness classification probability value and the corresponding willingness classification label in the second annotation label, the face recognition willingness loss value is calculated using the face recognition willingness loss function.

[0171] Optionally, the training module 25 is specifically used to calculate the total loss value based on the face recognition intention loss value, the balance coefficient corresponding to the face recognition intention loss value, and the multimodal segmentation loss value;

[0172] Based on the total loss value, the initialization parameters of the face recognition intention model are iteratively adjusted using gradient descent. Then, the process proceeds to the step of using the training batch data as input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value, until the total loss value meets the preset convergence condition, and the trained target parameters are obtained. Based on the target weights, the trained face recognition intention model is generated.

[0173] It should be noted that the facial recognition intention recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the facial recognition intention recognition method and the training method of the facial recognition intention recognition model. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the facial recognition intention recognition device and the facial recognition intention recognition method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0174] The embodiment numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0175] This specification also provides a storage medium storing a computer program, which, when executed by a processor, implements the above-described functionality. Figures 2-7 The facial recognition intention recognition method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 2-7 The specific details of the illustrated embodiments will not be elaborated here.

[0176] Please refer to Figure 10 This diagram illustrates the structure of a facial recognition intention identification device provided in an exemplary embodiment of this specification. The facial recognition intention identification device in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 can be connected via the bus 150.

[0177] The processor 110 may include one or more processing cores. The processor 110 connects to various parts of the facial recognition device via various interfaces and lines, and executes various functions of the terminal 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user page, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately using a communication chip.

[0178] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems.

[0179] The memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly.

[0180] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0181] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 can be a touch display screen.

[0182] The touch display screen can be designed as a full-screen, curved screen, or irregularly shaped screen. It can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; however, this specification does not limit the specific design of the embodiments.

[0183] In addition, those skilled in the art will understand that the structure of the facial recognition intention device shown in the above figures does not constitute a limitation on the facial recognition intention device. The facial recognition intention device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the facial recognition intention device may also include radio frequency circuits, input units, sensors, audio circuits, WiFi modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0184] exist Figure 10 In the facial recognition device shown, the processor 110 can be used to call the facial recognition application stored in the memory 120 and specifically perform the following operations:

[0185] Obtain 2D and 3D images of the target object for facial recognition transactions;

[0186] Based on the two-dimensional image and the three-dimensional image, a multimodal human body segmentation feature map is generated;

[0187] Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask image corresponding to the two-dimensional image is generated. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region.

[0188] Based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained;

[0189] Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained;

[0190] Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the facial recognition intention of the target detection object for the facial recognition transaction is confirmed.

[0191] In one embodiment, when the processor 110 generates a multimodal human segmentation feature map based on the two-dimensional and three-dimensional images, it specifically performs the following operations:

[0192] The two-dimensional image and the three-dimensional image are stitched together;

[0193] Features are extracted from the stitched 2D and 3D images to generate a multimodal human body segmentation feature map.

[0194] In one embodiment, when the processor 110 generates a mask image corresponding to the two-dimensional image based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, and the mask image is used to distinguish the first face region from other regions in the two-dimensional image, the processor 110 specifically performs the following operations:

[0195] Based on the face region of the target detection object in the two-dimensional image, a first filling region of the mask image corresponding to the target detection object and a second filling region outside the first filling region are determined.

[0196] A first fill value is assigned to the first filled region, and a second fill value is assigned to the second filled region to generate a mask map with a resolution consistent with that of the two-dimensional image.

[0197] In one embodiment, when the processor 110 executes the process of obtaining a two-dimensional modal fusion feature map based on the two-dimensional image, the mask image, and the multimodal human segmentation feature map, it specifically performs the following operations:

[0198] Extract the features of the two-dimensional image;

[0199] The features of the two-dimensional image and the mask image are fused to obtain the first fused feature;

[0200] According to the channel dimension, the first fusion feature and the multimodal human body segmentation feature map are connected to obtain a two-dimensional modal fusion feature map.

[0201] In one embodiment, when the processor 110 executes the process of obtaining a three-dimensional modal fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map, it specifically performs the following operations:

[0202] Extract the features of the three-dimensional image;

[0203] According to the channel latitude, the features of the three-dimensional image and the multimodal human body segmentation feature map are connected to obtain a three-dimensional modal fusion feature map.

[0204] In one embodiment, when the processor 110 executes the step of confirming the face-scanning intention of the target detection object for the face-scanning transaction based on the two-dimensional modality fusion feature map and the three-dimensional modality fusion feature map, it specifically performs the following operations:

[0205] Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, multimodal fusion features are obtained;

[0206] Based on the multimodal fusion features, the probability value of face recognition willingness is obtained;

[0207] Based on the face recognition willingness probability value and the preset probability threshold, the face recognition willingness result of the target detection object for the face recognition transaction is confirmed.

[0208] In one embodiment, when the processor 110 executes the process of obtaining multimodal fusion features based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, it specifically performs the following operations:

[0209] The two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map are connected according to the channel latitude to obtain multimodal fusion features.

[0210] The processor 110 can also be used to call the training application of the face recognition intention model stored in the memory 120, and specifically perform the following operations:

[0211] Collect sample two-dimensional images and sample three-dimensional images of the face recognition transaction as training datasets, and label the sample three-dimensional images and the intention classification labels of the sample detection objects in the sample three-dimensional images and the sample human body regions of the sample detection objects to generate the first label of the training dataset;

[0212] Randomly sample from the training dataset to generate training batch data and the second annotation labels corresponding to the training batch data;

[0213] The training batch data is used as the input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value.

[0214] The multimodal segmentation loss value corresponding to the human target segmentation probability map and the face recognition intention classification probability value are calculated using the face recognition intention recognition loss function;

[0215] The initialization parameters of the face recognition intention model are adjusted based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and the trained face recognition intention model is generated based on the target parameters.

[0216] In one embodiment, when the processor 110 executes the step of using the training batch data as input to the face recognition intention recognition model to obtain a sample human target segmentation probability map and a sample face recognition intention classification probability value, it specifically performs the following operations:

[0217] The sample 2D images and sample 3D images in the training batch data are input into the multimodal human target segmentation network, and the sample human target segmentation probability map is output.

[0218] Based on the third face region where the sample detection object is located in the sample two-dimensional image and the sample two-dimensional image, a corresponding sample mask image is generated;

[0219] The sample 2D image and the sample mask map are input into a 2D feature extraction network to obtain sample 2D features. The sample 2D features and the sample human target segmentation probability map are then input into a 2D feature fusion network to output a sample 2D fusion feature map.

[0220] Based on the pixel depth value of the fourth face region where the sample detection object is located in the sample 3D image, the pixel value of the sample 3D image is normalized to obtain a normalized sample 3D image.

[0221] The normalized sample 3D image is input into a 3D feature extraction network to obtain sample 3D features. The sample 3D features and the sample human target segmentation probability map are input into a 3D feature fusion network to output a sample 3D fusion feature map.

[0222] The sample's two-dimensional modality fusion feature map and the sample's three-dimensional modality fusion feature map are input into a multimodal fusion intention prediction network, which outputs the sample's face recognition intention classification probability value.

[0223] In one embodiment, when the processor 110 calculates the multimodal segmentation loss value corresponding to the sample human target segmentation probability map and the face recognition intention classification probability value corresponding to the sample face recognition intention loss function, it specifically performs the following operations:

[0224] Based on the sample human target segmentation probability map and the corresponding human target segmentation label in the second annotation label, the multimodal segmentation loss value is calculated using the multimodal segmentation loss function;

[0225] Based on the sample face recognition willingness classification probability value and the corresponding willingness classification label in the second annotation label, the face recognition willingness loss value is calculated using the face recognition willingness loss function.

[0226] In one embodiment, when the processor 110 performs the following operations to adjust the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generates the trained face recognition intention model based on the target parameters:

[0227] The total loss value is calculated based on the face recognition intention loss value, the balance coefficient corresponding to the face recognition intention loss value, and the multimodal segmentation loss value;

[0228] Based on the total loss value, the initialization parameters of the face recognition intention model are iteratively adjusted using gradient descent. Then, the process proceeds to the step of using the training batch data as input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value, until the total loss value meets the preset convergence condition, and the trained target parameters are obtained. Based on the target weights, the trained face recognition intention model is generated.

[0229] In the embodiments of this specification, two-dimensional and three-dimensional images of the target detection object for facial recognition transactions are acquired. Based on the two-dimensional and three-dimensional images, a multimodal human body segmentation feature map is generated. Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask map corresponding to the two-dimensional image is generated. A two-dimensional modality fusion feature map is obtained based on the two-dimensional image, the mask map, and the multimodal human body segmentation feature map. Simultaneously, a three-dimensional modality fusion feature map is obtained based on the three-dimensional image and the multimodal human body segmentation feature map. Based on the two-dimensional and three-dimensional modality fusion feature maps, the facial recognition intent of the target detection object is confirmed. By introducing the human target segmentation task into intent recognition, the human body region in the acquired multimodal images is segmented, and non-human background regions are removed. This allows the network to focus more on learning the payment behavior of human targets in the multimodal images, thereby improving the accuracy of facial recognition intent determination for the target detection object. By extracting features from the stitched 2D and 3D images, a multimodal human segmentation feature map is obtained. This extracted feature map focuses more on the human body region in the image, enabling the network to learn the behavior of the target in the image and accurately predict the face recognition intention. Furthermore, a mask image is generated based on the first face region of the target object, and then features are extracted from the 2D image. By adding new channels to the features of the 2D image, the features of the mask image are represented by these new channels, resulting in the first fused feature. This allows for more targeted learning of the target object region within the first fused feature. Simultaneously, features are extracted from the 3D image, and then the features of the 3D and 2D images are linked with the multimodal human segmentation feature map to obtain a 2D modal fusion feature map and a 3D modal fusion feature map. Further, based on the 2D and 3D modal fusion feature maps, a multimodal fusion feature is obtained to predict the probability value of face recognition intention, thus obtaining the face recognition intention result. By leveraging the segmentation information provided by multimodal human body segmentation feature maps, attention learning mechanisms for 2D and 3D modalities are offered, enabling multimodal task consistency learning and improving the accuracy of facial recognition intent identification for target detection objects.

[0230] Furthermore, by collecting sample 2D and 3D images of the target objects for face recognition transactions as training datasets, and labeling the intention classification labels and sample human body regions of the target objects in the sample 3D images, the first label of the training dataset is generated. Then, random sampling is performed from the training dataset to generate training batch data and corresponding second label. The training batch data is used as input to the face recognition intention recognition model to obtain the sample human body target segmentation probability map and the sample face recognition intention classification probability value. The face recognition intention recognition loss function is used to calculate the multimodal segmentation loss value corresponding to the human body target segmentation probability map and the face recognition intention loss value corresponding to the face recognition intention classification probability value. Based on the multimodal segmentation loss value and the face recognition intention loss value, the initialization parameters of the face recognition intention recognition model are adjusted to obtain the target parameters. The trained face recognition intention recognition model is generated based on the target parameters. By introducing the human target segmentation task into the end-to-end learning network, the human body region in the acquired image is segmented and the non-human background region is removed, so that the network focuses more on learning the payment behavior of human targets in the image. Then, the two-dimensional and three-dimensional multimodal features are fused and trained to obtain the face recognition intention recognition model, so as to achieve accurate prediction of face recognition intention.

[0231] Additionally, embodiments of this specification provide a computer program product comprising a computer program that, when executed by a processor of an electronic device, enables the processor to at least perform the functions described above. Figures 2 to 7 The method provided in the illustrated embodiment.

[0232] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0233] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.

Claims

1. A method for facial recognition intent identification, comprising: Obtain 2D and 3D images of the target object for facial recognition transactions; The two-dimensional image and the three-dimensional image are stitched together; Features are extracted from the stitched 2D and 3D images to generate a multimodal human body segmentation feature map; Based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, a mask image corresponding to the two-dimensional image is generated. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region. Based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map, a two-dimensional modal fusion feature map is obtained; Based on the three-dimensional image and the multimodal human body segmentation feature map, a three-dimensional modal fusion feature map is obtained; Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, the facial recognition intention of the target detection object for the facial recognition transaction is confirmed.

2. The method as described in claim 1, wherein generating a mask image corresponding to the two-dimensional image based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image, the mask image being used to distinguish the first face region from other regions in the two-dimensional image besides the first face region, includes: Based on the face region of the target detection object in the two-dimensional image, determine the first filling region of the mask image corresponding to the target detection object, and the second filling region outside the first filling region; A first fill value is assigned to the first filled region, and a second fill value is assigned to the second filled region to generate a mask map with a resolution consistent with that of the two-dimensional image.

3. The method as described in claim 2, wherein obtaining the two-dimensional modal fusion feature map based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map includes: Extract the features of the two-dimensional image; The features of the two-dimensional image and the mask image are fused to obtain the first fused feature; According to the channel dimension, the first fusion feature and the multimodal human body segmentation feature map are connected to obtain a two-dimensional modal fusion feature map.

4. The method as described in claim 1, wherein obtaining the three-dimensional modal fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map includes: Extract the features of the three-dimensional image; According to the channel dimension, the features of the three-dimensional image and the multimodal human body segmentation feature map are connected to obtain a three-dimensional modal fusion feature map.

5. The method as described in claim 1, wherein confirming the facial recognition intent of the target detection object for the facial recognition transaction based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map includes: Based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map, multimodal fusion features are obtained; Based on the multimodal fusion features, the probability value of face recognition willingness is obtained; Based on the face recognition willingness probability value and the preset probability threshold, the face recognition willingness result of the target detection object for the face recognition transaction is confirmed.

6. The method of claim 5, wherein obtaining multimodal fusion features based on the two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map includes: The two-dimensional modal fusion feature map and the three-dimensional modal fusion feature map are connected according to the channel dimension to obtain multimodal fusion features.

7. A training method for a facial recognition intention identification model, comprising: Collect sample two-dimensional images and sample three-dimensional images of the face recognition transaction as training datasets, and label the sample three-dimensional images and the intention classification labels of the sample detection objects in the sample three-dimensional images and the sample human body regions of the sample detection objects to generate the first label of the training dataset; Randomly sample from the training dataset to generate training batch data and the second annotation labels corresponding to the training batch data; The training batch data is used as the input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value. The multimodal segmentation loss value corresponding to the human target segmentation probability map is calculated based on the multimodal segmentation loss function, and the face recognition intention loss value corresponding to the face recognition intention classification probability value is calculated based on the face recognition intention loss function. The initialization parameters of the face recognition intention model are adjusted based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and the trained face recognition intention model is generated based on the target parameters.

8. The method as described in claim 7, wherein using the training batch data as input to the face recognition intention recognition model to obtain a sample human target segmentation probability map and a sample face recognition intention classification probability value includes: The sample 2D images and sample 3D images in the training batch data are input into the multimodal human target segmentation network, and the sample human target segmentation probability map is output. Based on the third face region where the sample detection object is located in the sample two-dimensional image and the sample two-dimensional image, a corresponding sample mask image is generated; The sample 2D image and the sample mask map are input into a 2D feature extraction network to obtain sample 2D features. The sample 2D features and the sample human target segmentation probability map are then input into a 2D feature fusion network to output a sample 2D fusion feature map. Based on the pixel depth value of the fourth face region where the sample detection object is located in the sample 3D image, the pixel value of the sample 3D image is normalized to obtain a normalized sample 3D image. The normalized sample 3D image is input into a 3D feature extraction network to obtain sample 3D features. The sample 3D features and the sample human target segmentation probability map are input into a 3D feature fusion network to output a sample 3D fusion feature map. The sample's two-dimensional modality fusion feature map and the sample's three-dimensional modality fusion feature map are input into a multimodal fusion intention prediction network, which outputs the sample's face recognition intention classification probability value.

9. The method as described in claim 8, wherein calculating the multimodal segmentation loss value corresponding to the human target segmentation probability map based on the multimodal segmentation loss function, and calculating the face recognition intention loss value corresponding to the face recognition intention classification probability value based on the face recognition intention loss function, comprises: Based on the sample human target segmentation probability map and the corresponding human target segmentation label in the second annotation label, the multimodal segmentation loss value is calculated using the multimodal segmentation loss function; Based on the sample face recognition willingness classification probability value and the corresponding willingness classification label in the second annotation label, the face recognition willingness loss value is calculated using the face recognition willingness loss function.

10. The method of claim 9, wherein adjusting the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generating the trained face recognition intention model based on the target parameters, comprises: The total loss value is calculated based on the face recognition intention loss value, the balance coefficient corresponding to the face recognition intention loss value, and the multimodal segmentation loss value; Based on the total loss value, the initialization parameters of the face recognition intention model are iteratively adjusted using gradient descent. Then, the process proceeds to the step of using the training batch data as input to the face recognition intention model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value, until the total loss value meets the preset convergence condition, and the trained target parameters are obtained. Based on the target weights, the trained face recognition intention model is generated.

11. A facial recognition intention identification device, the device comprising: The acquisition module is used to acquire two-dimensional and three-dimensional images of the target detection object for facial recognition transactions; A human body segmentation module is used to stitch the two-dimensional image and the three-dimensional image; Features are extracted from the stitched 2D and 3D images to generate a multimodal human body segmentation feature map; A two-dimensional modal feature generation module is used to generate a mask image corresponding to the two-dimensional image based on the two-dimensional image and the first face region of the target detection object in the two-dimensional image. The mask image is used to distinguish the first face region from other regions in the two-dimensional image besides the first face region. A two-dimensional modal feature fusion module is used to obtain a two-dimensional modal fusion feature map based on the two-dimensional image, the mask image, and the multimodal human body segmentation feature map; A three-dimensional modal feature generation module is used to obtain a three-dimensional modal fusion feature map based on the three-dimensional image and the multimodal human body segmentation feature map; The face recognition intention determination module is used to confirm the face recognition intention result of the target detection object for the face recognition transaction based on the two-dimensional modality fusion feature map and the three-dimensional modality fusion feature map.

12. A facial recognition intention identification device, the device comprising: The sample acquisition module is used to acquire sample two-dimensional images and sample three-dimensional images of the sample detection object for the transaction as training datasets, and to annotate the sample three-dimensional images and the intention classification labels of the sample detection object and the sample human body regions of the sample detection object in the sample three-dimensional images, thereby generating the first annotation label of the training dataset; The sampling module is used to randomly sample from the training dataset to generate training batch data and the second annotation label corresponding to the batch data; The prediction module is used to take the training batch data as input to the face recognition intention recognition model to obtain the sample human target segmentation probability map and the sample face recognition intention classification probability value. The loss calculation module is used to calculate the multimodal segmentation loss value corresponding to the human target segmentation probability map according to the multimodal segmentation loss function, and to calculate the face recognition intention loss value corresponding to the face recognition intention classification probability value according to the face recognition intention loss function. The training module is used to adjust the initialization parameters of the face recognition intention model based on the multimodal segmentation loss value and the face recognition intention loss value to obtain target parameters, and generate the trained face recognition intention model based on the target parameters.

13. An electronic device, comprising: Processor and memory; The memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1 to 10.

14. A storage medium storing a computer program that, when executed by a processor, implements the steps of the method as claimed in any one of claims 1 to 10.

15. A computer program product comprising: A computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method as described in any one of claims 1 to 10.