Face recognition method, system and device and electronic equipment
By using the main and secondary cameras working together and employing image depth information and spatial location feature fusion technology, the problem of facial recognition accuracy in densely populated scenes has been solved, achieving higher recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
In crowded scenarios, existing facial recognition models struggle to effectively identify faces due to occlusion or small exposed areas, leading to decreased recognition accuracy.
The system employs a main camera and multiple secondary cameras working in concert. The main camera acquires image depth information to determine spatial coordinates, while the secondary cameras capture facial images. Text features are generated using multimodal facial attribute labels, and feature fusion is performed by combining the spatial location information from the secondary cameras. Finally, the images are recognized in a face recognition classifier.
It improves the accuracy of face recognition in densely populated scenes, avoids the problems of face occlusion and small exposed area when using a single camera, and optimizes the accuracy of multi-view recognition.
Smart Images

Figure CN121747162A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of facial recognition, and more particularly to facial recognition methods, systems, devices, and electronic devices. Background Technology
[0002] In crowded settings such as train stations, theaters, school buildings, and large-scale performance venues, security cameras pre-positioned in these locations often struggle to capture all facial information due to the high density of people. This results in numerous issues, such as faces being obscured by others or a small exposed area of faces, making it difficult for existing facial recognition models to perform accurate facial recognition in such scenarios. Summary of the Invention
[0003] This application provides a face recognition method, apparatus, electronic device, and readable storage medium to enable face recognition in densely populated environments.
[0004] This application provides a face recognition method, which is applied to an electronic device, and the method includes:
[0005] Based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment, the spatial coordinate information of the target face in the target scene is determined; the target scene is equipped with a main camera and multiple secondary cameras; the face images of the target face captured by each of the secondary cameras deployed in the target scene at the first moment, which are located in the spatial coordinate information, are obtained.
[0006] Determine at least one face attribute label corresponding to the face image of the target face captured by the main camera at the first moment, and use the at least one face attribute label to generate target text features to describe the target face;
[0007] Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, the target visual features corresponding to each secondary camera are determined. The target visual features refer to features that contain the discrimination information of the target face.
[0008] Using the spatial location information of each secondary camera as prior knowledge for feature fusion, the target visual features corresponding to each secondary camera are fused to obtain the first fused feature;
[0009] The first fusion feature is fused with the facial feature to obtain the second fusion feature; the facial feature is obtained by inputting the facial image of the target face captured by the main camera at the first moment into the facial feature network; the second fusion feature is input into the facial recognition classifier to obtain the facial recognition result of the target face.
[0010] This application also provides a face recognition system, which includes: a main camera and multiple secondary cameras deployed in the same scene, and an electronic device for performing the above method;
[0011] The electronic device is the main camera, one of the secondary cameras, or other devices independent of the main camera and the secondary cameras;
[0012] The main camera is used to collect image depth information of the target face to determine the spatial coordinate information of the target face in the target scene; and to collect facial images of the target face.
[0013] The secondary camera is used to capture facial images.
[0014] This application embodiment also provides a face recognition device, which is applied to an electronic device, the device comprising:
[0015] An image processing module is used to determine the spatial coordinate information of the target face in the target scene based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment; the target scene is equipped with a main camera and multiple secondary cameras; and to obtain the face image of the target face captured by each of the secondary cameras deployed in the target scene at the first moment, which is located under the spatial coordinate information.
[0016] The feature processing module is configured to determine at least one facial attribute label corresponding to the facial image of the target face captured by the main camera at a first moment, and generate target text features describing the target face using the at least one facial attribute label; and,
[0017] Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, the target visual features corresponding to each secondary camera are determined. The target visual features refer to features that contain the discrimination information of the target face.
[0018] The fusion module is used to use the spatial position information of each secondary camera as prior knowledge for feature fusion, to fuse the target visual features corresponding to each secondary camera to obtain a first fused feature; and to fuse the first fused feature with the face feature to obtain a second fused feature; wherein the face feature is obtained by inputting the face image of the target face captured by the main camera at the first moment into the face feature network;
[0019] The face recognition module is used to input the second fused feature into the face recognition classifier to obtain the face recognition result of the target face.
[0020] This application also provides an electronic device, including: a processor; and a computer-readable storage medium storing computer program instructions, which are executed by the processor to perform the steps of the method described above.
[0021] This application also provides a machine-readable storage medium storing computer program instructions that, when executed, enable the implementation of the steps described above.
[0022] As can be seen from the above technical solutions, in this embodiment,
[0023] By combining facial images of the same person captured simultaneously by a main camera and multiple secondary cameras, final facial recognition can be achieved. This enables facial recognition in crowded scenes, avoiding problems such as faces being obscured by other faces or small exposed areas that occur when using facial images captured by a single camera, thus improving the accuracy of facial recognition in crowded scenes.
[0024] Furthermore, in this embodiment, when fusing the target visual features corresponding to multiple secondary cameras, the spatial position information of the secondary cameras corresponding to each target visual feature is also fused. This allows the spatial relationship between the secondary cameras that capture the face images of the target face from different perspectives to be used as prior knowledge for feature fusion. The spatial position information of each secondary camera is used as prior knowledge for feature fusion to set attention weights for the target visual features corresponding to each secondary camera, and the target visual features corresponding to each secondary camera are fused based on the attention weights of the target visual features corresponding to each secondary camera, thereby improving the effect of feature fusion. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0026] Figure 1 A scene structure diagram provided for embodiments of this application;
[0027] Figure 2 A flowchart illustrating the method provided in this application embodiment;
[0028] Figure 3 Example diagrams of the methods provided in the embodiments of this application;
[0029] Figure 4 This is a structural diagram of the device provided in the embodiments of this application;
[0030] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0032] To avoid problems such as numerous faces being obscured by others or limited face exposure in crowded settings like train stations, theaters, school buildings, and large performance venues, this embodiment deploys multiple cameras in these scenarios. One of these cameras serves as the main camera, while the others act as secondary cameras (also called auxiliary cameras). Compared to the secondary cameras, the main camera adds the following function: sensing object depth information. Optionally, this embodiment does not specifically limit the deployment location of the main camera. For example, the main camera can be deployed in the center or entrance of the scenario. Taking a train station as an example, the main camera could be deployed in the middle of the entrance or exit of the train station.
[0033] This embodiment establishes a security monitoring network for the above scenario based on one main camera and multiple secondary cameras, as detailed below. Figure 1 As shown, by using a main camera and multiple secondary cameras in a security monitoring network, the complementary information of features from different perspectives can be fully utilized, avoiding interference from inconsistent information between different perspectives and improving the accuracy of face recognition in the above scenarios. An example is given below:
[0034] See Figure 2 , Figure 2 This is a flowchart illustrating a method provided in an embodiment of this application. This process can be applied to electronic devices. The electronic device here can be, for example, any of the aforementioned cameras, or a device independent of any of the aforementioned cameras, such as a server, cloud platform, etc. This embodiment is not specifically limited to any particular type of device.
[0035] like Figure 2 As shown, the process may include the following steps:
[0036] Step 201: Based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment, determine the spatial coordinate information of the target face in the target scene, and obtain the face image of the target face captured by each secondary camera deployed in the target scene at the first moment under the spatial coordinate information.
[0037] The target scenario here refers to any of the scenarios mentioned above, and this embodiment is not specifically limited to any one of them.
[0038] As described above, the main camera can perceive depth information. Based on this, the main camera captures the image depth information of the target face at the first moment. As an example, this embodiment can determine the spatial coordinate information of the target face in the target scene based on the image depth information of the target face captured by the main camera at the first moment. The spatial coordinate information here refers to the coordinate information in the space corresponding to the target scene, which has a corresponding spatial coordinate system. This spatial coordinate system can be, for example, a coordinate system with the main camera as the origin, the vertical direction of the main camera as the vertical coordinate axis (y-axis), and the horizontal direction as the horizontal coordinate axis (x-axis).
[0039] In this embodiment, the image captured by the main camera at the first moment may include multiple faces, and the target face mentioned above is one of the selected faces. This embodiment can use moving target detection technology and multi-target detection technology to detect faces in the image to determine the target face.
[0040] As an example, in step 201 above, obtaining the face image of the target person's face captured by each of the deployed secondary cameras in the target scene at the first moment, which is located under the aforementioned spatial coordinate information, can be implemented in many ways, such as:
[0041] For each camera, based on the relative position between the secondary camera and the main camera in the pre-set spatial coordinate system, and combined with the image depth information, a target area is selected from the image captured by the secondary camera at the first moment, and the face in the target area is identified as the face image of the target face captured by the secondary camera at the first moment in the aforementioned spatial coordinate information.
[0042] For example, a mapping relationship between each secondary camera and the main camera can be established in advance based on the relative positions between the secondary camera and the main camera in a pre-defined spatial coordinate system. Under this premise, for each secondary camera, the region of the target face in the image captured by the main camera at the first moment is determined by combining the image depth information. The region is then mapped to the image captured by the secondary camera at the first moment according to the mapping relationship to obtain the mapped region (i.e., the target region). The face in the target region is then identified as the face image of the target face captured by the secondary camera at the first moment under the aforementioned spatial coordinate information.
[0043] Step 202: Determine at least one face attribute label corresponding to the face image of the target face captured by the main camera at the first moment, and use the at least one face attribute label to generate target text features to describe the target face.
[0044] As an example, determining at least one facial attribute label corresponding to the facial image of the target face captured by the main camera at the first moment in step 202 may include: inputting the facial image of the target face captured by the main camera at the first moment into a trained facial scoring attribute network to obtain at least one facial attribute label, such as pose, age, gender, whether or not a mask is being worn, etc. Here, "first moment" refers to any moment in general, and this embodiment is not specifically limited to any particular moment.
[0045] In this embodiment, to facilitate the generation of target text features for describing the target face, a dictionary enhancement library is introduced. This library is used to enrich the original face attribute labels and reduce the discrepancy between the face attribute labels and the text description. For example, if the face attribute labels include:
[0046] Attitude label: 30-degree offset;
[0047] Age tag: 30 years old;
[0048] Gender label: Male;
[0049] Mask label: Wear a mask.
[0050] The aforementioned dictionary enhancement library includes definitions for age, such as defining 30 years old as middle-aged. Similarly, the same dictionary enhancement library includes definitions for posture, such as defining an offset angle greater than 10 degrees as a profile view.
[0051] Based on this, using a dictionary enhancement library, the text description used to describe the target face can be obtained as: "This is the profile of a middle-aged man wearing a mask."
[0052] Based on the above description, the above-mentioned generation of target text features for describing a target face using at least one face attribute label may include: generating a specific text description (also known as a prompt) for describing the target face based on the constructed dictionary enhancement library and combined with the above-mentioned at least one face attribute label; inputting the specific text description Prompt into the trained text encoder to obtain the above-mentioned target text features.
[0053] In this embodiment, a prompt is generally used as input to a Large Language Model (LLM) to guide the model to output the desired result. For example, the prompt might be designed as "Please polish the given word set into a grammatically correct statement. The given word set is {computer, artificial intelligence, patent, large language model, prompt, this year, growth}". When this prompt is input into the LLM and inferred, the possible output might be "The design of prompts for large language models is a branch of computer science and artificial intelligence, and related patents have shown a gradual growth trend in recent years."
[0054] In this embodiment, the text encoder described above can be a large language model, such as a pre-trained model (CLIP: Contrastive Language-Image Pre-training) text encoder, which combines pre-training of text and images and uses contrastive learning for training, and can achieve excellent performance in a variety of natural language processing (NLP) and computer vision (CV) tasks.
[0055] Step 203: Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, determine the target visual features corresponding to each secondary camera. The target visual features refer to features that contain the discrimination information of the target face.
[0056] As an example, step 203 can be implemented as follows: The face images of the target face captured by each secondary camera at the first moment, located under the spatial coordinate information, are input into a trained visual encoder to obtain reference visual features corresponding to each secondary camera; then, inconsistencies in the reference visual features corresponding to each secondary camera are filtered to obtain target visual features corresponding to each secondary camera; the target visual features corresponding to each secondary camera retain key ID discrimination information from the face images of the target face captured by each secondary camera, indicating the discrimination information of the target face.
[0057] In this embodiment, the visual encoder may be, for example, a CLIP visual encoder.
[0058] As an example, the above-mentioned filtering of non-consistent information in the reference visual features corresponding to each secondary camera to obtain target visual features may include: using principal component analysis to decouple each reference visual feature to obtain decoupled features; for each decoupled feature, calculating the similarity between the feature and the target text feature; based on the similarity between each decoupled feature and the target text feature, obtaining the target visual features corresponding to each secondary camera from each decoupled feature; wherein the similarity between any target visual feature and the target text feature is greater than or equal to a set similarity threshold.
[0059] In this embodiment, the decoupling operation described above is used to remove redundant information, such as environmental interference. The decoupled features obtained after the decoupling operation can then be free of redundant information, such as environmental interference, and at this point, the decoupled features are mutually orthogonal.
[0060] In addition, in this embodiment, the target visual features are obtained from each decoupled feature based on the similarity between each decoupled feature and the target text features. The purpose is to filter out inconsistent information and ultimately retain the key discrimination information of the target face collected by each secondary camera.
[0061] As can be seen, this embodiment uses fine-grained facial information corresponding to multiple modalities (i.e., multiple facial attribute labels) to filter out the inconsistencies in the facial images captured by each secondary camera, and finally obtains features that retain the key discrimination information of the target face captured by each secondary camera, so that no interference information is introduced when supplementing the facial features of the target face captured by the main camera, thus optimizing the multi-view recognition accuracy.
[0062] Step 204: Using the spatial location information of each secondary camera as prior knowledge for feature fusion, the target visual features corresponding to each secondary camera are fused to obtain the first fused feature. The first fused feature is then fused with the face features to obtain the second fused feature. The face features are obtained by inputting the face image of the target face captured by the main camera at the first moment into the face feature network. The second fused feature is then input into the face recognition classifier to obtain the face recognition result of the target face.
[0063] As an example, in step 204, using the spatial position information of each secondary camera as prior knowledge for feature fusion to fuse the target visual features corresponding to each secondary camera to obtain the first fused feature may include: inputting each target visual feature as a word embedding and the position information of the secondary camera corresponding to each target visual feature as a position embedding into the trained Transformer model, so that the Transformer model can use the spatial position information of each secondary camera as prior knowledge for feature fusion to set attention weights for the target visual features corresponding to each secondary camera, and fuse the target visual features corresponding to each secondary camera based on the attention weights of the target visual features corresponding to each secondary camera to obtain the aforementioned first fused feature.
[0064] As an example, the more feature data or feature attributes contained in the target visual features corresponding to each secondary camera, the greater the attention weight of that target visual feature. Conversely, the less feature data or feature attributes contained in the target visual features corresponding to each secondary camera, the smaller the attention weight of that target visual feature.
[0065] As can be seen, when this embodiment uses the Transformer model to fuse the target visual features corresponding to multiple secondary cameras, it takes the spatial position information of each secondary camera as the position_embedding input. This allows the spatial relationship between the secondary cameras that capture the face images of the target face from different perspectives to serve as prior knowledge for feature fusion, so as to set attention weights for the target visual features corresponding to each secondary camera. For example, if the target visual features corresponding to a secondary camera contain a variety of feature types from a certain perspective, then the target visual features corresponding to that secondary camera can be configured with a larger attention weight, and vice versa. This greatly improves the effect of feature fusion.
[0066] As an example, the face recognition result here may include identifying the identity information corresponding to the face, but this example is not specifically limited.
[0067] This concludes the process. Figure 2 The process is shown below.
[0068] pass Figure 2As can be seen from the process shown, in this embodiment, facial images of the same face captured by a main camera and multiple secondary cameras at the same time are combined to perform the final facial recognition. This enables facial recognition in crowded scenes and avoids problems such as the face being obscured by the face in front or the small exposed area of the face when using facial images captured by a single camera for facial recognition, thereby improving the accuracy of facial recognition in crowded scenes.
[0069] Furthermore, in this embodiment, fine-grained facial information corresponding to multiple modalities (i.e., multiple facial attribute labels) is used to filter out the inconsistencies in the facial images captured by each secondary camera, ultimately obtaining features that retain the key discrimination information of the target face captured by each secondary camera. This ensures that no interference information is introduced when supplementing the facial features of the target face captured by the main camera, thereby optimizing the multi-view recognition accuracy.
[0070] Furthermore, in this embodiment, when fusing the target visual features corresponding to multiple secondary cameras, the spatial position information of the secondary cameras corresponding to each target visual feature is also fused. This allows the spatial relationship between the secondary cameras that capture the face images of the target face from different perspectives to be used as prior knowledge for feature fusion. The spatial position information of each secondary camera is used as prior knowledge for feature fusion to set corresponding attention weights for the target visual features corresponding to each secondary camera, and the target visual features corresponding to each secondary camera are fused based on the attention weights of the target visual features corresponding to each secondary camera, thereby improving the effect of feature fusion.
[0071] For ease of understanding, Figure 3 Examples are shown Figure 2 The diagram shows a specific example of the process.
[0072] like Figure 3 As shown, in this embodiment, the target visual features corresponding to the facial images of the target face captured by multiple secondary cameras, as well as the spatial position information of the secondary cameras corresponding to each target visual feature, are fused together as the first fusion feature. This first fusion feature is then fused with the facial features obtained by inputting the facial image of the target face captured by the main camera at the same time into the facial feature network. Facial recognition is performed based on the fused result, rather than relying solely on the facial features obtained by inputting the facial image of the target face captured by the main camera into the facial feature network. This obviously avoids the problems of the face being obscured by the face in front and the small exposed area of the face when performing facial recognition by relying on the facial image captured by a single camera, thus improving the accuracy of facial recognition in densely populated scenes.
[0073] The methods provided in the embodiments of this application have been described above. The systems and apparatus provided in the embodiments of this application are described below:
[0074] This application provides a face recognition system, which includes: a main camera and multiple secondary cameras deployed in the same scene, and an electronic device for performing the above method;
[0075] The electronic device here is either the main camera, one of the secondary cameras, or other devices independent of the main camera and the secondary cameras. This embodiment does not specifically limit the scope.
[0076] The main camera is used to collect image depth information of the target face to determine the spatial coordinate information of the target face in the target scene; and to collect facial images of the target face.
[0077] The secondary camera is used to capture facial images.
[0078] The apparatus provided in the embodiments of this application is described below:
[0079] See Figure 4 , Figure 4 This is a structural diagram of a device provided in an embodiment of this application. The device is applied to electronic devices, such as… Figure 4 As shown, the device may include:
[0080] An image processing module is used to determine the spatial coordinate information of the target face in the target scene based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment; the target scene is equipped with a main camera and multiple secondary cameras; and to obtain the face image of the target face captured by each of the secondary cameras deployed in the target scene at the first moment, which is located under the spatial coordinate information.
[0081] The feature processing module is configured to determine at least one facial attribute label corresponding to the facial image of the target face captured by the main camera at a first moment, and generate target text features describing the target face using the at least one facial attribute label; and,
[0082] Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, the target visual features corresponding to each secondary camera are determined. The target visual features refer to features that contain the discrimination information of the target face.
[0083] The fusion module is used to use the spatial position information of each secondary camera as prior knowledge for feature fusion, to fuse the target visual features corresponding to each secondary camera to obtain a first fused feature; and to fuse the first fused feature with the face feature to obtain a second fused feature; wherein the face feature is obtained by inputting the face image of the target face captured by the main camera at the first moment into the face feature network;
[0084] The face recognition module is used to input the second fused feature into the face recognition classifier to obtain the face recognition result of the target face.
[0085] As an example, determining at least one face attribute label corresponding to the face image of the target face captured by the main camera at the first moment includes: inputting the face image of the target face captured by the main camera at the first moment into a trained face scoring attribute network to obtain at least one face attribute label.
[0086] As one embodiment, determining the target visual features corresponding to each secondary camera based on the face image of the target face captured by each secondary camera at a first moment under the spatial coordinate information includes:
[0087] The face images of the target face captured by each secondary camera at the first moment under the spatial coordinate information are input into the trained visual encoder to obtain the reference visual features corresponding to each secondary camera.
[0088] Inconsistent information in the reference visual features corresponding to each secondary camera is filtered to obtain the target visual features corresponding to each secondary camera; the target visual features corresponding to each secondary camera retain the key ID discrimination information in the face image of the target face collected by each secondary camera to indicate the discrimination information of the target face.
[0089] As one embodiment, generating target text features to describe the target face using the at least one face attribute label includes:
[0090] Based on the constructed dictionary enhancement library and combined with the at least one face attribute label, a dedicated text description prompt word is generated to describe the target face; the dedicated text description prompt word is input into the trained text encoder to obtain the target text feature.
[0091] As an example, filtering out inconsistencies in the reference visual features corresponding to each secondary camera to obtain the target visual features corresponding to each secondary camera includes: using principal component analysis to decouple each reference visual feature to obtain decoupled features; the decoupling operation is used to remove redundant information; the decoupled features are mutually orthogonal; for each decoupled feature, the similarity between the feature and the target text feature is calculated; based on the similarity between the decoupled features and the target text feature, the target visual features corresponding to each secondary camera are obtained from the decoupled features; the similarity between the target visual feature corresponding to any secondary camera and the target text feature is greater than or equal to a set similarity threshold.
[0092] As one embodiment, the step of using the spatial position information of each secondary camera as prior knowledge for feature fusion to fuse the target visual features corresponding to each secondary camera to obtain the first fused feature includes: embedding the target visual features corresponding to each secondary camera as word embeddings and embedding the position information of each secondary camera corresponding to each target visual feature as position embeddings, and inputting them together into a trained Transformer model. The Transformer model uses the spatial position information of each secondary camera as prior knowledge for feature fusion to set attention weights for the target visual features corresponding to each secondary camera, and fuses the target visual features corresponding to each secondary camera based on the attention weights of the target visual features corresponding to each secondary camera to obtain the first fused feature; wherein, the more features contained in the target visual features corresponding to each secondary camera, the greater the attention weight of the target visual feature.
[0093] As an example, obtaining the face image of the target face captured by each secondary camera deployed in the target scene at the first moment under the spatial coordinate information includes: for each secondary camera, based on the relative position between the secondary camera and the main camera under the pre-set spatial coordinate axis, and combined with the image depth information, selecting a target area from the image captured by the secondary camera at the first moment, and determining the face in the target area as the face image of the target face captured by the secondary camera at the first moment under the spatial coordinate information.
[0094] This concludes the process. Figure 4 Structural description of the device shown.
[0095] Please see Figure 5 , Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. Figure 5 As shown, the hardware structure may include: a processor and a computer-readable storage medium, the computer-readable storage medium storing computer program instructions executable by the processor; the processor is used to execute the computer program instructions to implement the method disclosed in the above example of this application.
[0096] Based on the same concept as the above-described method, this application also provides a computer-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above-described examples of this application.
[0097] For example, the aforementioned computer-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, computer-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0098] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A face recognition method, characterized in that, This method is applied to electronic devices, and the method includes: Based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment, the spatial coordinate information of the target face in the target scene is determined; the target scene is equipped with a main camera and multiple secondary cameras; the face images of the target face captured by each of the secondary cameras deployed in the target scene at the first moment, which are located in the spatial coordinate information, are obtained. Determine at least one face attribute label corresponding to the face image of the target face captured by the main camera at the first moment, and use the at least one face attribute label to generate target text features to describe the target face; Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, the target visual features corresponding to each secondary camera are determined. The target visual features refer to features that contain the discrimination information of the target face. Using the spatial location information of each secondary camera as prior knowledge for feature fusion, the target visual features corresponding to each secondary camera are fused to obtain the first fused feature; The first fusion feature is fused with the facial feature to obtain the second fusion feature; the facial feature is obtained by inputting the facial image of the target face captured by the main camera at the first moment into the facial feature network; the second fusion feature is input into the facial recognition classifier to obtain the facial recognition result of the target face.
2. The method according to claim 1, characterized in that, The determination of the target visual features corresponding to each secondary camera based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information includes: The face images of the target face captured by each secondary camera at the first moment under the spatial coordinate information are input into the trained visual encoder to obtain the reference visual features corresponding to each secondary camera. Inconsistent information in the reference visual features corresponding to each secondary camera is filtered to obtain the target visual features corresponding to each secondary camera; the target visual features corresponding to each secondary camera retain the key ID discrimination information in the face image of the target face collected by each secondary camera to indicate the discrimination information of the target face.
3. The method according to claim 1, characterized in that, The step of generating target text features to describe the target face using the at least one face attribute label includes: Based on the constructed dictionary enhancement library and combined with the at least one face attribute label, a dedicated text description prompt word is generated to describe the target face; the dedicated text description prompt word is input into the trained text encoder to obtain the target text feature.
4. The method according to claim 2, characterized in that, The step of filtering out inconsistencies in the reference visual features corresponding to each secondary camera to obtain the target visual features corresponding to each secondary camera includes: Principal component analysis is used to decouple each reference visual feature to obtain decoupled features; the decoupling operation is used to remove redundant information; the decoupled features are mutually orthogonal. For each decoupled feature, calculate the similarity between that feature and the target text feature; Based on the similarity between each decoupled feature and the target text feature, the target visual features corresponding to each secondary camera are obtained from each decoupled feature; the similarity between the target visual feature corresponding to any secondary camera and the target text feature is greater than or equal to a set similarity threshold.
5. The method according to claim 1, characterized in that, The step of using the spatial location information of each secondary camera as prior knowledge for feature fusion to fuse the target visual features corresponding to each secondary camera to obtain the first fused feature includes: The target visual features corresponding to each secondary camera are used as word embeddings, and the position information of the secondary cameras corresponding to each target visual feature is used as position embeddings. These are input together into the trained Transformer model. The Transformer model uses the spatial position information of each secondary camera as feature fusion prior knowledge to set attention weights for the target visual features corresponding to each secondary camera. Based on the attention weights of the target visual features corresponding to each secondary camera, the target visual features corresponding to each secondary camera are fused to obtain the first fused feature. The more features contained in the target visual features corresponding to each secondary camera, the greater the attention weight of that target visual feature.
6. The method according to claim 1, characterized in that, The facial images of the target face captured at the first moment by each of the secondary cameras deployed in the target scene, located in the spatial coordinate information, include: For each secondary camera, based on the relative position between the secondary camera and the main camera in the pre-defined spatial coordinate system, and in conjunction with the image depth information, a target region is selected from the image captured by the secondary camera at the first moment, and the face in the target region is identified as the face image of the target face captured by the secondary camera at the first moment in the spatial coordinate information.
7. A face recognition system, characterized in that, The system includes: a main camera and multiple secondary cameras deployed in the same scene, and an electronic device for performing the method as claimed in any one of claims 1 to 6; The electronic device is the main camera, one of the secondary cameras, or other devices independent of the main camera and the secondary cameras; The main camera is used to collect image depth information of the target face to determine the spatial coordinate information of the target face in the target scene; and to collect facial images of the target face. The secondary camera is used to capture facial images.
8. A face recognition device, characterized in that, This device is used in electronic devices, and the device includes: An image processing module is used to determine the spatial coordinate information of the target face in the target scene based on the image depth information of the target face captured by the main camera deployed in the target scene at the first moment; the target scene is equipped with a main camera and multiple secondary cameras; and to obtain the face image of the target face captured by each of the secondary cameras deployed in the target scene at the first moment, which is located under the spatial coordinate information. The feature processing module is configured to determine at least one facial attribute label corresponding to the facial image of the target face captured by the main camera at a first moment, and generate target text features describing the target face using the at least one facial attribute label; and, Based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information, the target visual features corresponding to each secondary camera are determined. The target visual features refer to features that contain the discrimination information of the target face. The fusion module is used to use the spatial position information of each secondary camera as prior knowledge for feature fusion, to fuse the target visual features corresponding to each secondary camera to obtain a first fused feature; and to fuse the first fused feature with the face feature to obtain a second fused feature; wherein the face feature is obtained by inputting the face image of the target face captured by the main camera at the first moment into the face feature network; The face recognition module is used to input the second fused feature into the face recognition classifier to obtain the face recognition result of the target face.
9. The apparatus according to claim 8, characterized in that, The determination of the target visual features corresponding to each secondary camera based on the face image of the target face captured by each secondary camera at the first moment under the spatial coordinate information includes: The face images of the target face captured by each secondary camera at the first moment under the spatial coordinate information are input into the trained visual encoder to obtain the reference visual features corresponding to each secondary camera. Inconsistent information in the reference visual features corresponding to each secondary camera is filtered to obtain the target visual features corresponding to each secondary camera; the target visual features corresponding to each secondary camera retain the key ID discrimination information in the face images of the target face captured by each secondary camera, to indicate the discrimination information of the target face; and / or, The step of generating target text features to describe the target face using the at least one face attribute label includes: Based on the constructed dictionary enhancement library and combined with the at least one facial attribute label, a dedicated text description prompt for describing the target face is generated; the dedicated text description prompt is input into a trained text encoder to obtain the target text features; and / or, The step of filtering out inconsistencies in the reference visual features corresponding to each secondary camera to obtain the target visual features corresponding to each secondary camera includes: using principal component analysis to decouple each reference visual feature to obtain decoupled features; the decoupling operation is used to remove redundant information; the decoupled features are mutually orthogonal; for each decoupled feature, the similarity between the feature and the target text feature is calculated; based on the similarity between the decoupled features and the target text feature, the target visual features corresponding to each secondary camera are obtained from the decoupled features; the similarity between the target visual feature corresponding to any secondary camera and the target text feature is greater than or equal to a set similarity threshold; and / or, The step of using the spatial position information of each secondary camera as prior knowledge for feature fusion to fuse the target visual features corresponding to each secondary camera to obtain the first fused feature includes: embedding the target visual features corresponding to each secondary camera as word embeddings and embedding the position information of each secondary camera corresponding to each target visual feature as position embeddings, both of which are input into a trained Transformer model. The Transformer model then uses the spatial position information of each secondary camera as prior knowledge for feature fusion to set attention weights for the target visual features corresponding to each secondary camera, and fuses the target visual features corresponding to each secondary camera based on these attention weights to obtain the first fused feature. The more features contained in the target visual features corresponding to each secondary camera, the greater the attention weight of that target visual feature; and / or, The process of obtaining the face image of the target face captured by each of the secondary cameras deployed in the target scene at the first moment under the spatial coordinate information includes: for each secondary camera, based on the relative position between the secondary camera and the main camera under the pre-set spatial coordinate axis and combined with the image depth information, selecting a target area from the image captured by the secondary camera at the first moment, and determining the face in the target area as the face image of the target face captured by the secondary camera at the first moment under the spatial coordinate information.
10. An electronic device, characterized in that, include: processor; as well as A computer-readable storage medium storing computer program instructions that are executed by the processor to perform the steps of the method as claimed in any one of claims 1 to 6.