Facial recognition method, apparatus, device, and medium
The facial recognition model, trained using near-infrared and visible light sample images, solves the problem of low facial recognition accuracy under environmental factors and achieves efficient recognition under different light source conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL BANK OF CHINA
- Filing Date
- 2022-10-18
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, facial recognition methods are greatly affected by environmental factors, especially visible light recognition, which is easily affected by changes in lighting, while near-infrared recognition lacks texture information, resulting in reduced accuracy.
A face recognition model is trained using both near-infrared and visible light sample images. By feature fusion and loss function optimization, the robustness and feature extraction capabilities of the model are improved.
It improves the accuracy of facial recognition, avoids the problems of low accuracy and model overfitting in single-light source recognition, and ensures efficient recognition in different environments.
Smart Images

Figure CN115439916B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of facial recognition technology, and in particular to a facial recognition method, device, equipment and medium. Background Technology
[0002] Facial recognition technology is widely used in real life, and common facial recognition methods can be achieved using visible light or near-infrared light.
[0003] In existing technologies, facial recognition using visible light is easily affected by environmental factors (such as lighting), leading to a decrease in the accuracy of facial recognition. While facial recognition using near-infrared light is unaffected by environmental factors, the facial images acquired using near-infrared light lack texture information, which also results in a decrease in the accuracy of facial recognition. Summary of the Invention
[0004] This invention provides a facial recognition method, apparatus, device, and medium to improve the accuracy of facial recognition.
[0005] According to one aspect of the present invention, a facial recognition method is provided, comprising:
[0006] Acquire a facial image to be identified; wherein, the facial image to be identified includes a near-infrared image to be identified and / or a visible light image to be identified;
[0007] The face image to be recognized is input into the trained face recognition model to obtain the face recognition result;
[0008] The facial recognition model is trained using both near-infrared and visible light sample images of the training subjects.
[0009] According to another aspect of the present invention, a facial recognition device is provided, comprising:
[0010] A facial image acquisition module is used to acquire a facial image to be identified; wherein, the facial image to be identified includes a near-infrared image to be identified and / or a visible light image to be identified;
[0011] The facial recognition result acquisition module is used to input the facial image to be recognized into the trained facial recognition model to obtain the facial recognition result.
[0012] The facial recognition model is trained using both near-infrared and visible light sample images of the training subjects.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] One or more processors;
[0015] Memory, used to store one or more programs;
[0016] When one or more programs are executed by one or more processors, the one or more processors are able to execute any of the facial recognition methods provided in the embodiments of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement any of the facial recognition methods provided in the embodiments of the present invention.
[0018] This invention provides a facial recognition scheme by acquiring a facial image to be recognized, wherein the facial image to be recognized includes a near-infrared image and / or a visible light image; the facial image to be recognized is input into a trained facial recognition model to obtain a facial recognition result; wherein the facial recognition model is trained based on near-infrared sample images and visible light sample images of the sample training object. This scheme performs facial recognition by using a facial recognition model trained with both near-infrared and visible light sample images. It is convenient to operate and can simultaneously perform facial recognition on both near-infrared and visible light images, avoiding the low accuracy that can occur when performing facial recognition based on a single facial image due to environmental factors or lack of detailed texture information, thus improving the accuracy of facial recognition. Furthermore, using both near-infrared and visible light sample images to train the facial recognition model avoids overfitting when the number of visible light sample images is small, ensuring the feature extraction capability of the facial recognition model.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a facial recognition method provided in Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart of a facial recognition method provided in Embodiment 2 of the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a facial recognition device provided in Embodiment 3 of the present invention;
[0024] Figure 4 This is a schematic diagram of the structure of an electronic device that implements a facial recognition method, as provided in Embodiment 4 of the present invention. Detailed Implementation
[0025] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0026] Example 1
[0027] Figure 1 This is a flowchart of a facial recognition method provided in Embodiment 1 of the present invention. This embodiment can be applied to the situation of facial recognition of facial images to be recognized. The method can be executed by a facial recognition device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries facial recognition function.
[0028] See Figure 1 The facial recognition method shown includes:
[0029] S110. Obtain the facial image to be recognized.
[0030] The facial image to be recognized refers to the image for which facial recognition is required. This embodiment of the invention does not limit the type of facial image to be recognized; it can be set by a technician as needed. For example, the facial image to be recognized can be a human face image or an animal face image, etc. Specifically, the facial image to be recognized includes a near-infrared image to be recognized and / or a visible light image to be recognized. The near-infrared image to be recognized refers to a facial image to be recognized acquired using near-infrared light (wavelength 780–1100 nm). The visible light image to be recognized refers to a facial image to be recognized acquired using visible light (wavelength 380–780 nm).
[0031] Specifically, acquire the near-infrared image and / or the visible light image to be identified.
[0032] S120. Input the face image to be recognized into the trained face recognition model to obtain the face recognition result.
[0033] The facial recognition model can be a model that performs facial recognition tasks based on images of the face to be recognized. Specifically, the facial recognition model is trained using both near-infrared and visible light sample images of the training subject.
[0034] Here, the training sample refers to the object that provides the training samples. This embodiment of the invention does not limit the type of training sample, and it can be set by a technician as needed. For example, the training sample can be a person or an animal, etc.
[0035] Near-infrared sample images refer to sample images acquired using near-infrared light. Visible light sample images refer to sample images acquired using visible light. This embodiment of the invention does not limit the method of acquiring near-infrared and visible light sample images. For example, near-infrared and visible light sample images can be obtained by a technician from an image database as needed. Here, an image database refers to a database that can be used to store facial images.
[0036] Among them, the facial recognition result refers to the result of identifying the object to which the facial image to be identified belongs.
[0037] Specifically, the facial recognition model outputs the predicted probability of belonging to different candidate objects based on the input facial image to be recognized, and selects the target predicted object from the candidate predicted objects based on the predicted probability as the facial recognition result.
[0038] This invention provides a facial recognition scheme by acquiring a facial image to be recognized, wherein the facial image to be recognized includes a near-infrared image and / or a visible light image; the facial image to be recognized is input into a trained facial recognition model to obtain a facial recognition result; wherein the facial recognition model is trained based on near-infrared sample images and visible light sample images of the sample training object. This scheme performs facial recognition by using a facial recognition model trained with both near-infrared and visible light sample images. It is convenient to operate and can simultaneously perform facial recognition on both near-infrared and visible light images, avoiding the low accuracy that can occur when performing facial recognition based on a single facial image due to environmental factors or lack of detailed texture information, thus improving the accuracy of facial recognition. Furthermore, using both near-infrared and visible light sample images to train the facial recognition model avoids overfitting when the number of visible light sample images is small, ensuring the feature extraction capability of the facial recognition model.
[0039] Example 2
[0040] Figure 2This is a flowchart of a facial recognition method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment improves the training mechanism of the facial recognition model by adding the following steps: Near-infrared sample images and visible light sample images are fused in batches to obtain fused sample images; the fused sample images are input into a pre-constructed facial recognition model to obtain fused sample features; the facial recognition model is trained based on the fused sample features, the near-infrared labels of the near-infrared sample images, and the visible light labels of the visible light sample images. It should be noted that parts not detailed in this embodiment can be found in the descriptions of other embodiments.
[0041] See Figure 2 The facial recognition method shown includes:
[0042] S210. The near-infrared sample image and the visible light sample image are fused according to the batch dimension to obtain the fused sample image.
[0043] The fused sample image refers to the image obtained by merging near-infrared and visible light sample images of the same training object according to batch dimensions. Specifically, the near-infrared sample image can be upgraded based on the number of channels in the visible light sample image to update the near-infrared sample image; then, the visible light sample image and the updated near-infrared sample image are fused to obtain the fused sample image. The number of channels refers to the number of color spaces that can be used to jointly form the image channels.
[0044] For example, a visible light sample image has RGB (three primary colors) channels, meaning it has 3 channels; a near-infrared sample image has 1 channel. If the shape of the visible light sample image of any training object is (b1, h, w, 3), and the shape of its near-infrared sample image is (b2, h, w, 1), then (b2, h, w, 1) is copied three times to obtain an updated near-infrared sample image with the shape (b2, h, w, 3). The visible light sample image and the updated near-infrared sample image are then fused according to batch dimensions to obtain a fused sample image with the shape (b1 + b2, h, w, 3). Here, b1 refers to the batch size of the visible light sample images, or the number of visible light sample images in a batch; b2 refers to the batch size of the near-infrared sample images, or the number of near-infrared sample images in a batch; h is the height of the sample image; and w is the width of the sample image.
[0045] Understandably, by upscaling the near-infrared sample images, and then fusing the upscaled near-infrared sample images with the visible light sample images to obtain fused sample images, the near-infrared sample images and the visible light sample images can be used as a set of data to be input into the face recognition model, thereby improving the input efficiency of the face recognition model and reducing the use of resources.
[0046] Specifically, based on the batch dimension, feature fusion is performed on near-infrared sample images and visible light sample images to obtain fused sample images.
[0047] S220. Input the fused sample image into the pre-built facial recognition model to obtain the fused sample features.
[0048] Here, fusion sample features refer to the feature vectors that can be used to represent fusion sample images. Specifically, the features extracted by the face recognition model from the input fusion sample image are called fusion sample features.
[0049] Continuing the previous example, a fused sample image with shape (b1+b2,h,w,3) is input into the face recognition model. The face recognition model extracts features from the fused sample image and outputs the fused sample features as (b1+b2,c). Here, c represents the features extracted from the height h, width w, and batch 3 of the fused sample image.
[0050] S230. The face recognition model is trained based on the features of the fused samples, the near-infrared labels of the near-infrared sample images, and the visible light labels of the visible light sample images.
[0051] Near-infrared tags refer to the real object corresponding to any near-infrared sample image. Visible light tags refer to the real object corresponding to any visible light sample image.
[0052] For example, the fused sample features can be split according to the batch dimension to obtain the near-infrared sample features of the near-infrared sample image and the visible light sample features of the visible light sample image; the feature distance between the near-infrared sample features and the visible light sample features can be determined; and the face recognition model can be trained based on the feature distance, the near-infrared sample features, the visible light sample features, the near-infrared label, and the visible light label.
[0053] Near-infrared sample features refer to feature vectors that can be used to represent near-infrared sample images. Visible light sample features refer to feature vectors that can be used to represent visible light sample images.
[0054] Continuing from the previous example, the output fused sample features (b1+b2,c) are split according to the batch dimension to obtain visible light sample features (b1,c) and near-infrared sample features (b2,c).
[0055] The feature distance is used to quantify the feature differences between near-infrared and visible light sample features. This invention does not limit the method for determining the feature distance; it can be set by a technician based on experience or needs.
[0056] For example, the maximum mean discrepancy (MMD) between near-infrared sample features and visible light sample features can be determined, and this maximum mean discrepancy can be used as the feature distance. The expression for this maximum mean discrepancy is:
[0057]
[0058] Where X represents the set of visible light sample features, Y represents the set of near-infrared sample features, n represents the number of visible light sample features, and m represents the number of near-infrared sample features. i Let y represent the feature of the i-th visible light sample. j Let φ represent the feature of the j-th near-infrared sample, φ represent the mapping function, and H represent the feature distance, which is the measure by which visible light sample features and near-infrared sample features are mapped to the reproducing kernel Hilbert space (RKHS) by φ().
[0059] Expanding the square in the above expression for the maximum mean difference, we get the expanded expression for the maximum mean difference as follows:
[0060]
[0061] To facilitate calculation, a kernel function can be introduced to transform the expanded expression of the maximum mean difference. The expression of the kernel function is as follows:
[0062] κ(x,y)=φ(x)φ(y);
[0063] Where x represents a variable that follows a distribution of visible light sample features X, and y represents a variable that follows a distribution of near-infrared sample features Y.
[0064] The expanded expression for the maximum mean difference after transformation is:
[0065]
[0066] It should be noted that the embodiments of the present invention do not limit the type of kernel function introduced, and it can be selected by those skilled in the art based on experience or needs. For example, a Gaussian kernel function can be introduced for transformation. The expression for the Gaussian kernel function is:
[0067]
[0068] Where e represents the exponent and σ represents the parameter.
[0069] Specifically, the feature distance between near-infrared sample features and visible light sample features is determined based on the expression for the maximum mean difference. Optionally, the smaller the feature distance between near-infrared sample features and visible light sample features, the closer the near-infrared sample features and visible light sample features are; the larger the feature distance between near-infrared sample features and visible light sample features, the less close the near-infrared sample features and visible light sample features are.
[0070] Understandably, by introducing the maximum mean difference to determine the feature distance, the calculation is simple and the result of determining the feature distance is more accurate.
[0071] In this embodiment of the invention, by splitting the fused sample features into near-infrared sample features and visible light sample features according to the batch dimension, and determining the feature distance between the near-infrared sample features and visible light sample features, the distribution of visible light sample features is transferred to the distribution of near-infrared sample features, so that the face recognition model has high face recognition accuracy under the condition of being robust to the environment.
[0072] In one alternative embodiment, the face recognition model is trained based on the determined feature distance, near-infrared sample features, visible light sample features, near-infrared labels, and visible light labels. The target loss can be determined by introducing different loss functions, and the network parameters of the face recognition model can be adjusted according to the determined target loss.
[0073] For example, the target loss can be determined based on feature distance, near-infrared sample features, visible light sample features, near-infrared labels, and visible light labels; the face recognition model can then be trained based on the target loss.
[0074] In one optional embodiment, near-infrared prediction results for near-infrared sample features and visible light prediction results for visible light sample features can be determined. Near-infrared loss is determined based on the near-infrared prediction results and near-infrared labels; visible light loss is determined based on the visible light prediction results and visible light labels; and target loss is determined based on at least one of near-infrared loss, visible light loss, and feature distance. Accordingly, the parameters of the face recognition model are adjusted based on the determined target loss until the model training cutoff condition is met. The model training cutoff condition can be a preset number of training samples, a preset number of training iterations, or the target loss tending to converge. The preset number and preset number of iterations can be set or adjusted by technicians according to needs or experience.
[0075] Understandably, determining the target loss based on feature distance, near-infrared sample features, visible light sample features, near-infrared tags, and visible light tags improves the diversity and richness of target loss determination, making the method of determining target loss more scalable.
[0076] Near-infrared loss is used to quantify the difference between near-infrared prediction results and near-infrared tags. Visible light loss is used to quantify the difference between visible light prediction results and visible light tags.
[0077] The present invention does not impose any limitations on the selection of the loss function, which can be selected by those skilled in the art based on experience or needs. For example, the loss function can be the cross-entropy loss function.
[0078] For example, the visible light loss can be calculated using the following formula:
[0079]
[0080] Where l1 represents the visible light loss, and n represents the size of the visible light sample set, i.e., the number of visible light sample images; represents the loss of the i-th visible light sample image in the visible light sample set; N represents the number of predictable categories; This represents the visible light label (value 0 or 1). If the true label category of the i-th visible light sample image is equal to c, then... Take 1; otherwise, Set to 0; This represents the predicted probability that the i-th visible light image belongs to category c.
[0081] For example, the near-infrared loss can be calculated using the following formula:
[0082]
[0083] Where l2 represents the near-infrared loss, and m represents the size of the near-infrared sample set, i.e., the number of near-infrared sample images; represents the loss of the i-th near-infrared sample image in the near-infrared sample set; M represents the number of predictable categories; Let c represent the near-infrared label (value 0 or 1). If the true label category of the i-th near-infrared sample image is equal to c, then... Take 1; otherwise, Set to 0; This represents the predicted probability that the i-th near-infrared image belongs to category c.
[0084] Accordingly, the target loss can be determined based on at least one of near-infrared loss, visible light loss, and characteristic distance. The corresponding expression for the target loss is:
[0085] loss = α × l1 + β × l2 + γ × d;
[0086] Wherein, l1 represents visible light loss; l2 represents near-infrared loss; d represents feature distance; α, β and γ are loss weights, and their specific values can be empirical or experimental values. The specific values of α, β and γ can be the same or different. This embodiment of the invention does not impose any limitations on this, only requiring that the sum of α, β and γ is 1.
[0087] Understandably, by determining the target loss based on at least one of near-infrared loss, visible light loss, and characteristic distance, the inaccuracy of the determined target loss, which occurs when determining the target loss based on a single data point, is avoided, thus improving the accuracy of the target loss determination.
[0088] S240. Obtain the facial image to be recognized.
[0089] S250. Input the face image to be recognized into the trained face recognition model to obtain the face recognition result.
[0090] It should be noted that the device used to train the facial recognition model and the device used to use the facial recognition model can be the same or different, and this invention does not impose any limitations on this.
[0091] In this embodiment of the invention, near-infrared sample images and visible light sample images are fused in batches to obtain fused sample images. These fused sample images are then input into a pre-built facial recognition model to obtain fused sample features. The facial recognition model is trained based on these fused sample features, near-infrared labels from the near-infrared sample images, and visible light labels from the visible light sample images. The training method for the facial recognition model is described in detail below. This scheme, by introducing fused sample features, near-infrared labels, and visible light labels, enables the training of a pre-built facial recognition model, providing data support for the training process and avoiding inaccuracies that can occur when training a facial recognition model based on single sample features, thus improving the accuracy of facial recognition model training. Simultaneously, fusing features from near-infrared and visible light sample images to obtain fused sample images allows them to share the facial recognition model, making transfer learning easier and preventing the facial recognition model from failing to converge due to significant differences in image feature distribution between the near-infrared and visible light sample images.
[0092] Example 3
[0093] Figure 3This is a schematic diagram of the structure of an electronic device that implements a facial recognition method according to Embodiment 3 of the present invention. This embodiment can be applied to the situation of facial recognition of facial images to be recognized. The method can be executed by a facial recognition device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries facial recognition function.
[0094] like Figure 3 As shown, the device includes: a facial image acquisition module 310 and a facial recognition result acquisition module 320. Among them,
[0095] The facial image acquisition module 310 is used to acquire a facial image to be identified; wherein, the facial image to be identified includes a near-infrared image to be identified and / or a visible light image to be identified;
[0096] The facial recognition result acquisition module 320 is used to input the facial image to be recognized into the trained facial recognition model to obtain the facial recognition result.
[0097] The facial recognition model is trained using both near-infrared and visible light sample images of the training subjects.
[0098] This invention provides a facial recognition scheme. A facial image acquisition module acquires a facial image to be recognized, including a near-infrared image and / or a visible light image. A facial recognition result acquisition module inputs the facial image to a trained facial recognition model to obtain the facial recognition result. The facial recognition model is trained using both near-infrared and visible light sample images of the target object. This scheme, by using a facial recognition model trained with both near-infrared and visible light sample images, offers convenient operation and allows for simultaneous facial recognition of both images. This avoids the inaccuracies often seen with a single facial image, which can be affected by environmental factors or lack of detailed texture information, thus improving accuracy. Furthermore, using both near-infrared and visible light sample images to train the facial recognition model prevents overfitting when the number of visible light sample images is limited, ensuring the model's feature extraction capabilities.
[0099] Optionally, the device further includes:
[0100] The fusion sample image acquisition module is used to perform feature fusion of near-infrared sample images and visible light sample images according to the batch dimension to obtain fusion sample images;
[0101] The fusion sample feature acquisition module is used to input the fusion sample image into a pre-built face recognition model to obtain the fusion sample features;
[0102] The face recognition model training module is used to train the face recognition model based on the features of fused samples, the near-infrared labels of near-infrared sample images, and the visible light labels of visible light sample images.
[0103] Optional, a facial recognition model training module, including:
[0104] The sample feature acquisition unit is used to split the fused sample features according to the batch dimension to obtain the near-infrared sample features of the near-infrared sample image and the visible light sample features of the visible light sample image.
[0105] The feature distance determination unit is used to determine the feature distance between near-infrared sample features and visible light sample features;
[0106] The face recognition model training unit is used to train the face recognition model based on feature distance, near-infrared sample features, visible light sample features, near-infrared labels, and visible light labels.
[0107] Optional, facial recognition model training units include:
[0108] The target loss determination subunit is used to determine the target loss based on feature distance, near-infrared sample features, visible light sample features, near-infrared tags, and visible light tags.
[0109] The model training subunit is used to train the face recognition model based on the target loss.
[0110] Optional, the target loss determination subunit is specifically used for:
[0111] Near-infrared prediction results for near-infrared sample characteristics and visible light prediction results for visible light sample characteristics;
[0112] The near-infrared loss is determined based on the near-infrared prediction results and the near-infrared tag.
[0113] Based on the visible light prediction results and visible light labels, determine the visible light loss;
[0114] The target loss is determined based on at least one of near-infrared loss, visible light loss, and characteristic distance.
[0115] Optional, feature distance determination unit, specifically used for:
[0116] The maximum mean difference between near-infrared sample features and visible light sample features is determined, and the maximum mean difference is used as the feature distance.
[0117] Optionally, the fusion sample image acquisition module includes:
[0118] The near-infrared sample image update unit is used to perform dimensionality upscaling on the near-infrared sample image based on the number of channels of the visible light sample image, so as to update the near-infrared sample image.
[0119] The fusion sample image acquisition unit performs feature fusion on the visible light sample image and the updated near-infrared sample image to obtain the fusion sample image.
[0120] The facial recognition device provided in the embodiments of the present invention can execute the facial recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each facial recognition method.
[0121] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision and disclosure of near-infrared images to be identified, visible light images to be identified, near-infrared sample images and visible light sample images, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0122] Example 4
[0123] Figure 4 This is a schematic diagram of the structure of an electronic device implementing a facial recognition method according to Embodiment 4 of the present invention. The electronic device 410 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0124] like Figure 4As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412 or a random access memory (RAM) 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.
[0125] Multiple components in electronic device 410 are connected to I / O interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of displays, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as facial recognition methods.
[0127] In some embodiments, the facial recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the facial recognition method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to perform the facial recognition method by any other suitable means (e.g., by means of firmware).
[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A facial recognition method, characterized in that, include: Acquire a facial image to be identified; wherein, the facial image to be identified includes a near-infrared image to be identified and / or a visible light image to be identified; The face image to be recognized is input into the trained face recognition model, which outputs the prediction probability of belonging to different candidate prediction objects, and selects the target prediction object from the candidate prediction objects according to the prediction probability as the face recognition result. The facial recognition model is trained using near-infrared and visible light sample images of the training subjects. The facial recognition model was trained in the following manner: The near-infrared sample image is upscaled based on the number of channels in the visible light sample image to update the near-infrared sample image; wherein, the number of channels is the number of color spaces used to jointly form image channels; The visible light sample image and the updated near-infrared sample image are fused to obtain a fused sample image; the fused sample image is then input into a pre-built face recognition model to obtain fused sample features; wherein, the fused sample features are feature vectors used to represent the fused sample image; The fused sample features are split according to the batch dimension to obtain the near-infrared sample features of the near-infrared sample image and the visible light sample features of the visible light sample image; The maximum mean difference between the near-infrared sample features and the visible light sample features is determined, and the maximum mean difference is used as the feature distance; wherein, the feature distance is used to quantify the feature difference between the near-infrared sample features and the visible light sample features; The target loss is determined based on the feature distance, the near-infrared sample features, the visible light sample features, the near-infrared tag, and the visible light tag. The facial recognition model is trained based on the target loss; wherein the near-infrared label is the real object corresponding to any near-infrared sample image, and the visible light label is the real object corresponding to any visible light sample image. The step of determining the target loss based on the feature distance, the near-infrared sample features, the visible light sample features, the near-infrared tag, and the visible light tag includes: Determine the near-infrared prediction result of the near-infrared sample features and the visible light prediction probability of the visible light sample features; Based on the near-infrared prediction results and the near-infrared tag, the near-infrared loss is determined; Based on the visible light prediction results and the visible light label, the visible light loss is determined; The target loss is determined based on at least one of the near-infrared loss, the visible light loss, and the characteristic distance.
2. A facial recognition device, characterized in that, include: A facial image acquisition module is used to acquire a facial image to be identified; wherein, the facial image to be identified includes a near-infrared image to be identified and / or a visible light image to be identified; The facial recognition result acquisition module is used to input the facial image to be recognized into the trained facial recognition model, output the prediction probability of belonging to different candidate prediction objects, and select the target prediction object from the candidate prediction objects according to the prediction probability as the facial recognition result. The facial recognition model is trained using near-infrared and visible light sample images of the training subjects. The device further includes: The fusion sample image acquisition module includes a near-infrared sample image update unit and a fusion sample image acquisition unit; The near-infrared sample image update unit is used to perform dimensionality-upgrading processing on the near-infrared sample image according to the number of channels of the visible light sample image, so as to update the near-infrared sample image; wherein, the number of channels is the number of color spaces used to jointly form image channels; The fused sample image acquisition unit is used to perform feature fusion of the visible light sample image and the updated near-infrared sample image to obtain a fused sample image. The fusion sample feature acquisition module is used to input the fusion sample image into a pre-built face recognition model to obtain fusion sample features; wherein, the fusion sample features are feature vectors used to represent the fusion sample image; The facial recognition model training module includes a sample feature acquisition unit, a feature distance determination unit, and a facial recognition model training unit. The sample feature acquisition unit is used to split the fused sample features according to the batch dimension to obtain the near-infrared sample features of the near-infrared sample image and the visible light sample features of the visible light sample image. The feature distance determination unit is used to determine the maximum mean difference between the near-infrared sample features and the visible light sample features, and to use the maximum mean difference as the feature distance; wherein, the feature distance is used to quantify the feature difference between the near-infrared sample features and the visible light sample features; The facial recognition model training unit includes: The target loss determination subunit is used to determine the target loss based on the feature distance, the near-infrared sample features, the visible light sample features, the near-infrared tag, and the visible light tag. The model training subunit is used to train the face recognition model based on the target loss; wherein, the near-infrared label is the real object corresponding to any near-infrared sample image, and the visible light label is the real object corresponding to any visible light sample image; Specifically, the target loss determination subunit is used for: Determine the near-infrared prediction result of the near-infrared sample features and the visible light prediction probability of the visible light sample features; Based on the near-infrared prediction results and the near-infrared tag, the near-infrared loss is determined; Based on the visible light prediction results and the visible light label, the visible light loss is determined; The target loss is determined based on at least one of the near-infrared loss, the visible light loss, and the characteristic distance.
3. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a facial recognition method as described in claim 1.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a facial recognition method as described in claim 1.