A method and apparatus for identity recognition
By generating 3D point parameters using generative adversarial networks and convolutional neural networks, and combining near-infrared and 3D information for image rotation and translation, the problem of low accuracy and insufficient security caused by image differences in identity recognition is solved, achieving higher recognition accuracy and security.
Patent Information
- Application Number
- CN202111189093.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-10-12
AI Technical Summary
Existing identity recognition technologies suffer from low accuracy and insufficient security due to image differences, and fake substitutes are easily identified.
Generative adversarial networks are used to generate 3D point parameters, combined with convolutional neural networks for image rotation and translation, and multi-dimensional image morphology for identity recognition. Near-infrared parameters and 3D parameters are used to improve recognition accuracy and security.
It improves the accuracy and security of identity recognition, effectively avoids non-human attacks, simplifies the calculation process, and improves recognition efficiency.
Smart Images

Figure CN113903065B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and more specifically, to a method and apparatus for identity recognition. Background Technology
[0002] With the development of science and technology, the application of identity recognition technology is becoming increasingly widespread, especially facial recognition, which is used in many important fields. In the process of identity recognition, user devices need to pre-record a sample image. Then, when recognition is required, a camera captures a picture of the target and compares the captured image with the sample image to complete the identification. However, in actual operation, it is impossible to guarantee that the image captured by the camera will perfectly match the pre-recorded sample image. When the difference between the captured image and the pre-recorded sample image is too large, the user device cannot perform the identification, resulting in low accuracy and efficiency in identifying the target. Furthermore, in current identity recognition technology, many spoofed substitutes for the target can be successfully identified by user devices, which poses a certain impact on the security of identity recognition.
[0003] Therefore, improving the accuracy and security of identity recognition technology is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides a method and apparatus for identity recognition, which can improve the accuracy and security of identity recognition technology.
[0005] In a first aspect, a method for identity recognition is provided, comprising: acquiring a first near-infrared parameter on a first three-dimensional point of a target to be identified, the first near-infrared parameter being used to indicate information about near-infrared light reflected by the target at the position of the first three-dimensional point; acquiring a first three-dimensional parameter on the first three-dimensional point of the target to be identified, the first three-dimensional parameter being used to indicate three-dimensional position information of the first three-dimensional point on the target to be identified; and recognizing a first image of the target to be identified, the first image including a plurality of the first three-dimensional points.
[0006] The identity recognition method provided in this application can combine heterogeneous images—that is, different types of image forms of the same target—for the same target to be identified. This allows for verification of the target's authenticity from multiple dimensions, resulting in higher security and reducing the likelihood of misidentifying fake substitutes as the real target. Especially when the target is a live object such as a face, the method provided in this application can effectively prevent attacks from non-live objects such as paper or masks, making identity recognition more secure and efficient.
[0007] In one possible implementation, the method further includes: acquiring a second near-infrared parameter on a second three-dimensional point of the target to be identified, the second three-dimensional point being generated by a generative adversarial network, the second near-infrared parameter being used to indicate information about the target to be identified reflecting near-infrared light at the location of the second three-dimensional point; and identifying a first image of the target to be identified, including: identifying a second image of the target to be identified, the second image including a plurality of the first three-dimensional points and a plurality of the second three-dimensional points.
[0008] In another possible implementation, the method further includes: obtaining a second three-dimensional parameter on the second three-dimensional point of the target to be identified, the second three-dimensional point being generated by a generative adversarial network, the second three-dimensional parameter being used to indicate the position information of the second three-dimensional point on the target to be identified; and identifying a first image of the target to be identified, including: identifying a second image of the target to be identified, the second image including a plurality of the first three-dimensional points and a plurality of the second three-dimensional points.
[0009] In the identity recognition method provided in this application embodiment, the user equipment can utilize a generative adversarial network to complete the three-dimensional points in the first image. This allows the user equipment to recognize the identity of the target based on parameters from more three-dimensional points, thereby improving the accuracy of identity recognition. Specifically, when generating parameters for the second three-dimensional points, only the second near-infrared parameters, only the second three-dimensional parameters, or both can be generated simultaneously. Especially when the second three-dimensional points include both second near-infrared parameters and second three-dimensional parameters, the user equipment can improve both the accuracy and security of identity recognition.
[0010] In another possible implementation, the second three-dimensional point is generated by the generative adversarial network based on the first three-dimensional point.
[0011] In another possible implementation, the second three-dimensional point is generated by a generative adversarial network that performs symmetric processing on the first three-dimensional point.
[0012] The second 3D points can be randomly generated on the first image and judged entirely by the discriminator in the generative adversarial network (GAN). Alternatively, they can be generated systematically based on the parameters of the first 3D points, and then these second 3D points are combined with the first 3D points to form an image, which is then judged by the discriminator. For example, the parameters of the first 3D points can be symmetrically processed first, and then the GAN can generate the second 3D points based on the symmetrically processed 3D points.
[0013] In another possible implementation, the method further includes: projecting the first three-dimensional point onto a plane to form a first pixel, the first pixel including a first parameter and a second parameter, the first parameter being formed by the first near-infrared parameter and the second parameter being formed by the first three-dimensional parameter; and identifying a first image of the target to be identified, including: identifying a third image of the target to be identified, the third image including a plurality of the first pixels.
[0014] By projecting the parameters of three-dimensional points onto a plane, the calculation process in the identity recognition process can be simplified, the amount of computation required for identity recognition by user equipment can be reduced, and user equipment can quickly perform recognition without losing the parameters of three-dimensional points, thereby improving the efficiency of identity recognition.
[0015] In another possible implementation, the method further includes: obtaining a third parameter on a second pixel of the target to be identified, the second pixel being generated by a generative adversarial network, the third parameter being used to indicate information about the target to be identified reflecting near-infrared light at the location of the second pixel; and identifying a third image of the target to be identified, including: identifying a fourth image of the target to be identified, the fourth image including a plurality of the first pixels and a plurality of the second pixels.
[0016] In another possible implementation, the method further includes: obtaining a fourth parameter on the second pixel of the target to be identified, the fourth parameter being used to indicate the position information of the second pixel on the target to be identified; and identifying a third image of the target to be identified, including: identifying a fourth image of the target to be identified, the fourth image including a plurality of the first pixels and a plurality of the second pixels.
[0017] In another possible implementation, the method further includes: projecting the second three-dimensional point onto a plane to form a third pixel, the third pixel including a fifth parameter and a sixth parameter, the fifth parameter being formed by the second near-infrared parameter and the sixth parameter being formed by the second three-dimensional parameter; and identifying a second image of the target to be identified, including: identifying a fifth image of the target to be identified, the fourth image including a plurality of the first pixel and a plurality of the third pixel.
[0018] In another possible implementation, obtaining the first three-dimensional parameter of the first three-dimensional point includes: obtaining the second three-dimensional parameter of the first three-dimensional point, the second three-dimensional parameter being used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified; generating a first rotation angle and / or a first translation amount using a convolutional neural network; and calculating the second three-dimensional parameter according to the first rotation angle and / or the first translation amount to obtain the first three-dimensional parameter.
[0019] In the method provided in this application embodiment, a convolutional neural network is used to generate rotation angles and translations for the directly acquired three-dimensional parameters. This allows the parameters at the first three-dimensional point to be rotated and translated according to the rotation angles and translations. Thus, even when the directly acquired image and the pre-recorded image are not completely overlapping, the rotation and translation of the three-dimensional points adjust the image to an angle and position similar to the pre-recorded image, thereby improving the accuracy and efficiency of subsequent identity recognition.
[0020] In another possible implementation, obtaining a first near-infrared parameter at a first three-dimensional point includes: obtaining a second near-infrared parameter at the first three-dimensional point, the second near-infrared parameter being used to indicate information about the near-infrared light reflected by the target to be identified at the position of the first three-dimensional point; generating a first rotation angle and / or a first translation amount using a convolutional neural network; and calculating the second near-infrared parameter according to the first rotation angle and / or the first translation amount to obtain the first near-infrared parameter.
[0021] In another possible implementation, obtaining the second three-dimensional parameter of the first three-dimensional point includes: obtaining a third three-dimensional parameter on the first three-dimensional point, the third three-dimensional parameter being acquired by a camera; calculating a first normal vector of the third three-dimensional parameter on the first three-dimensional point based on the normal vector of the adjacent surface of the first three-dimensional point; and calculating the normal vectors of the components of the first normal vector on the X-axis, Y-axis, and Z-axis respectively, wherein the second three-dimensional parameter includes normal vectors on the X-axis, Y-axis, and Z-axis.
[0022] The identity recognition method provided in this application can combine different image forms of the same target to identify images, verifying the authenticity of the target from multiple dimensions, thus enhancing security and preventing the misidentification of fake substitutes as the real target. Furthermore, when the target is not directly facing the camera but at an angle, rotation and translation can be used to allow the user device to recognize the corrected image, improving accuracy. Moreover, missing parts of the original image can be filled in using a model trained by a neural network, enabling the user device to extract more feature points for comparison with images in the database during identity recognition, further improving accuracy and efficiency.
[0023] In a second aspect, a method for training a neural network model is provided, comprising: acquiring first training data and first template data, wherein the first training data includes second three-dimensional parameters of each three-dimensional point on a target to be identified at different angles, and the first template data is the second three-dimensional parameters of each three-dimensional point on the target to be identified at a first angle, wherein the first angle is the angle at which the target to be identified can be identified; inputting the first training data into a first model to obtain a second rotation angle and / or a second translation amount; rotating and / or translating the first training data according to the second rotation angle and / or the second translation amount to obtain second training data; calculating a first loss function value based on the difference between the second training data and the first template data; and adjusting the parameters of the first model based on the first loss function value.
[0024] In one possible implementation, the method further includes: when the value of the first loss function is less than or equal to a first threshold, outputting the second rotation angle and / or the second translation amount.
[0025] In another possible implementation, inputting the first training data into the first model to obtain the second rotation angle and / or the second translation amount includes: inputting the first training data and the first template data into the first model to obtain the second rotation angle and / or the second translation amount.
[0026] The first model trained using the method for training a neural network model provided in this application can output rotation angles and translations based on the parameters of the input 3D points. This allows multiple 3D points to rotate and translate according to their corresponding output rotation angles and translations, forming a relatively ideal image. This enables the user equipment to quickly perform identification during the identity verification process.
[0027] Thirdly, a method for training a neural network model is provided, comprising: acquiring a first training image and a first template image, wherein the first training image includes multiple incomplete images of a target to be identified, and the first template image is a complete image of the target to be identified; inputting the first training image into a second model to obtain a first complete image; calculating a second loss function value based on the difference between the first complete image and the first template image; and adjusting the parameters of the second model based on the second loss function value.
[0028] In one possible implementation, the method further includes: outputting the first complete image when the value of the second loss function is less than or equal to a second threshold.
[0029] The second model trained using the method for training neural network models provided in this application can output a complete image based on an incomplete input image, through generation by an internal generator and judgment by a discriminator. This enables the user device to extract more feature values of parameter points during the process of identifying the target in the image, thereby improving the accuracy and efficiency of identity recognition.
[0030] Fourthly, an identity recognition device is provided, comprising: an acquisition unit, configured to acquire a first near-infrared parameter on a first three-dimensional point of a target to be identified, the first near-infrared parameter being used to indicate information about near-infrared light reflected by the target at the position of the first three-dimensional point; and an acquisition unit, configured to acquire a first three-dimensional parameter on the first three-dimensional point of the target to be identified, the first three-dimensional parameter being used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified; and a processing unit, configured to recognize a first image of the target to be identified, the first image including a plurality of the first three-dimensional points.
[0031] In one possible implementation, the acquisition unit is used to acquire a second near-infrared parameter on a second three-dimensional point of the target to be identified, the second three-dimensional point being generated by a generative adversarial network, and the second near-infrared parameter being used to indicate information about the target to be identified reflecting near-infrared light at the location of the second three-dimensional point; the processing unit is used to identify a second image of the target to be identified, the second image including a plurality of the first three-dimensional points and a plurality of the second three-dimensional points.
[0032] In another possible implementation, the acquisition unit is used to acquire a second three-dimensional parameter on the second three-dimensional point of the target to be identified, the second three-dimensional point being generated by a generative adversarial network, and the second three-dimensional parameter being used to indicate the position information of the second three-dimensional point on the target to be identified; the processing unit is used to identify a second image of the target to be identified, the second image including a plurality of the first three-dimensional points and a plurality of the second three-dimensional points.
[0033] In another possible implementation, the second three-dimensional point is generated by the generative adversarial network based on the first three-dimensional point.
[0034] In another possible implementation, the second three-dimensional point is generated by a generative adversarial network that performs symmetric processing on the first three-dimensional point.
[0035] In another possible implementation, the processing unit is used to project the first three-dimensional point onto a plane to form a first pixel, the first pixel including a first parameter and a second parameter, the first parameter being formed by the first near-infrared parameter and the second parameter being formed by the first three-dimensional parameter, and to identify a third image of the target to be identified, the third image including a plurality of the first pixels.
[0036] In another possible implementation, the acquisition unit is used to acquire a third parameter on a second pixel of the target to be identified, the second pixel being generated by a generative adversarial network, and the third parameter being used to indicate information about the target to be identified reflecting near-infrared light at the location of the second pixel; the processing unit is used to identify a fourth image of the target to be identified, the fourth image including a plurality of the first pixels and a plurality of the second pixels.
[0037] In another possible implementation, the acquisition unit is used to acquire a fourth parameter on the second pixel of the target to be identified, the fourth parameter being used to indicate the position information of the second pixel on the target to be identified; the processing unit is used to identify a fourth image of the target to be identified, the fourth image including a plurality of the first pixels and a plurality of the second pixels.
[0038] In another possible implementation, the processing unit is used to project the second three-dimensional point onto a plane to form a third pixel point, the third pixel point including a fifth parameter and a sixth parameter, the fifth parameter being formed by the second near-infrared parameter and the sixth parameter being formed by the second three-dimensional parameter, to identify a fifth image of the target to be identified, the fourth image including a plurality of the first pixel points and a plurality of the third pixel points.
[0039] In another possible implementation, the acquisition unit is used to acquire a second three-dimensional parameter of the first three-dimensional point, the second three-dimensional parameter being used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified; the processing unit is used to generate a first rotation angle and / or a first translation amount using a convolutional neural network, and to calculate the second three-dimensional parameter according to the first rotation angle and / or the first translation amount to obtain the first three-dimensional parameter.
[0040] In another possible implementation, the acquisition unit is used to acquire a second near-infrared parameter of the first three-dimensional point, the second near-infrared parameter being used to indicate information about the near-infrared light reflected by the target to be identified at the position of the first three-dimensional point; the processing unit is used to generate a first rotation angle and / or a first translation amount using a convolutional neural network, and calculate the second near-infrared parameter according to the first rotation angle and / or the first translation amount to obtain the first near-infrared parameter.
[0041] In another possible implementation, the acquisition unit is used to acquire a third three-dimensional parameter on the first three-dimensional point, which is acquired by the camera; the processing unit is used to calculate a first normal vector of the third three-dimensional parameter on the first three-dimensional point based on the normal vector of the adjacent surface of the first three-dimensional point, and calculate the normal vectors of the components of the first normal vector on the X-axis, Y-axis and Z-axis respectively, and the second three-dimensional parameter includes the normal vectors on the X-axis, Y-axis and Z-axis.
[0042] Fifthly, an apparatus for training a neural network model is provided, comprising: an acquisition unit for acquiring first training data and first template data, the first training data including second three-dimensional parameters of each three-dimensional point on a target to be identified at different angles, the first template data being the second three-dimensional parameters of each three-dimensional point on the target to be identified at a first angle, the first angle being the angle at which the target to be identified can be identified; and a processing unit for inputting the first training data into a first model to obtain a second rotation angle and / or a second translation amount, rotating and / or translating the first training data according to the second rotation angle and / or the second translation amount to obtain second training data, calculating a first loss function value based on the difference between the second training data and the first template data, and adjusting the parameters of the first model based on the first loss function value.
[0043] In one possible implementation, the processing unit is used to output the second rotation angle and / or the second translation amount when the first loss function value is less than or equal to the first threshold.
[0044] In another possible implementation, the processing unit is used to input the first training data and the first template data into the first model to obtain the second rotation angle and / or the second translation amount.
[0045] A sixth aspect provides an apparatus for training a neural network model, comprising: an acquisition unit for acquiring a first training image and a first template image, the first training image including multiple incomplete images of a target to be identified, and the first template image being a complete image of the target to be identified; and a processing unit for inputting the first training image into a second model to obtain a first complete image, calculating a second loss function value based on the difference between the first complete image and the first template image, and adjusting the parameters of the second model based on the second loss function value.
[0046] In one possible implementation, the processing unit outputs the first complete image when the value of the second loss function is less than or equal to the second threshold.
[0047] A seventh aspect provides an electronic device comprising: an identity recognition device as described in any possible embodiment of the fourth aspect.
[0048] Eighthly, a computer-readable storage medium is provided for storing program instructions that, when executed by a computer, perform the methods in any of the possible embodiments of the first to third aspects described above.
[0049] A ninth aspect provides a computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method in any of the possible implementations of the first to third aspects described above. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of a system architecture provided in an embodiment of this application.
[0051] Figure 2 This is a schematic flowchart illustrating an identity recognition method provided in an embodiment of this application.
[0052] Figure 3 This is a schematic flowchart illustrating another identity recognition method provided in the embodiments of this application.
[0053] Figure 4 This is a schematic flowchart illustrating another identity recognition method provided in the embodiments of this application.
[0054] Figure 5 This is a schematic flowchart illustrating a method for training a neural network model provided in an embodiment of this application.
[0055] Figure 6 This is a schematic flowchart illustrating another method for training a neural network model provided in an embodiment of this application.
[0056] Figure 7 This is a schematic diagram of an identity recognition device provided in an embodiment of this application.
[0057] Figure 8 This is a schematic diagram of the hardware structure of an identity recognition device provided in an embodiment of this application. Detailed Implementation
[0058] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0059] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.
[0060] In this application, the terms "first," "second," and "third" are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "third," nor are they limited in quantity or execution order.
[0061] This application will present various aspects, embodiments, or features relating to systems that may include multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0062] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.
[0063] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0064] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0065] In this application, "at least one" means one or more, and "more than one" means two or more.
[0066] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0067] Since the embodiments of this application involve the application of neural networks, for ease of understanding, the relevant terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.
[0068] (1) Convolutional Neural Network
[0069] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as using a trainable filter to convolve with an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a CNN that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to some of its neighboring neurons. A convolutional layer typically contains several feature maps, each composed of rectangularly arranged neural units. Neural units on the same feature map share weights, which are the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The underlying principle is that the statistical information of one part of the image is the same as that of other parts. This means that image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations in the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0070] Convolutional kernels can be initialized as matrices of random size, and during the training of a convolutional neural network, they can learn appropriate weights. Furthermore, sharing weights directly reduces the number of connections between layers in the convolutional neural network, while also lowering the risk of overfitting.
[0071] (2) Generative adversarial networks (GAN)
[0072] GAN is a machine learning architecture proposed by Ian Goodfellow of the University of Montreal in 2014, which can be as creative and imaginative as artificial intelligence. Generally speaking, GAN mainly consists of two types of networks: a generator G and a discriminator D.
[0073] G is responsible for generating images. That is, after inputting a random code z, it outputs a fake image G(z) automatically generated by the neural network. D is responsible for receiving the image output by G as input and then judging whether the image is real or fake; it outputs 1 for real and 0 for fake. For example, GAN can be simply viewed as a game between two networks. D trains a binary classification neural network using data from real and fake images. G can fabricate a "fake image" based on a string of random numbers. G then uses this fabricated fake image to deceive D, which is responsible for identifying whether it is real or fake and giving a score. For example, if G generates an image and receives a high score from D, it means that G's generation ability is very successful; if D gives a low score, then G's performance is not good enough and its parameters need to be adjusted.
[0074] To overcome some limitations of traditional GANs, several generative networks have been developed, such as pix2pix, CycleGAN, and pix2pixHD.
[0075] The embodiments of this application are applicable to identity recognition systems, such as facial recognition systems, including but not limited to products based on optical facial imaging. This identity recognition system can be applied to various electronic devices with image acquisition devices (such as cameras), including personal computers, computer workstations, smartphones, tablets, smart cameras, media consumption devices, wearable devices, set-top boxes, game consoles, augmented reality (AR) / virtual reality (VR) devices, in-vehicle terminals, etc. The embodiments disclosed in this application do not limit this application.
[0076] It should be understood that the specific examples in this document are only intended to help those skilled in the art better understand the embodiments of this application, and are not intended to limit the scope of the embodiments of this application.
[0077] It should also be understood that the various implementation methods described in this specification can be implemented individually or in combination, and the embodiments of this application are not limited in this respect.
[0078] Unless otherwise stated, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.
[0079] To better understand the solutions of the embodiments of this application, the following will first combine... Figure 1 A brief introduction to the possible application scenarios of the embodiments of this application is provided.
[0080] like Figure 1 As shown, this application embodiment provides a system architecture 100. In Figure 1 In this embodiment, the data acquisition device 160 is used to collect training data. For the face recognition method according to this application, the training data may include training images or training videos.
[0081] After collecting the training data, the data acquisition device 160 stores the training data in the database 130, and the training device 120 trains the target model / rule 101 based on the training data maintained in the database 130.
[0082] The aforementioned target model / rule 101 can be used to implement the face recognition method of this application embodiment. Specifically, the target model / rule 101 in this application embodiment can be a neural network. It should be noted that in practical applications, the training data maintained in the database 130 may not all come from the data acquisition device 160; it may also be received from other devices. Furthermore, it should be noted that the training device 120 may not necessarily train the target model / rule 101 entirely based on the training data maintained in the database 130; it may also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0083] The target model / rule 101 trained using training device 120 can be applied to different systems or devices, such as... Figure 1 The execution device 110 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, etc., or it can be a server or cloud service. Figure 1 In this embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data may include the video or image to be processed input by the client device 140.
[0084] In some implementations, the client device 140 may be the same device as the execution device 110. For example, both the client device 140 and the execution device 110 may be terminal devices.
[0085] In other embodiments, the client device 140 may be a different device from the execution device 110. For example, the client device 140 may be a terminal device, while the execution device 110 may be a cloud device, a server, or other such device. The client device 140 may interact with the execution device 310 through a communication network of any communication mechanism / standard. The communication network may be a wide area network, a local area network, a point-to-point connection, or any combination thereof.
[0086] The calculation module 111 of the execution device 110 is used to process the input data (such as the image to be processed) received by the I / O interface 112. During the calculation and other related processing performed by the calculation module 111 of the execution device 110, the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0087] Finally, I / O interface 112 returns the processing result, such as the face recognition result obtained above, to client device 140, thereby providing it to the user.
[0088] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different objectives or tasks. The corresponding target models / rules 101 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.
[0089] exist Figure 1 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.
[0090] It is worth noting that, Figure 1This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 1 In this context, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed within the execution device 110.
[0091] like Figure 1 As shown, the target model / rule 101 is trained by the training device 120. In this embodiment of the application, the target model / rule 101 can be a neural network. Specifically, the neural network in this embodiment of the application can be a CNN, a region convolutional neural network (RCNN), a faster region convolutional neural network (faster RCNN), or other types of neural networks, etc. This application does not make any specific limitations on this.
[0092] The following is combined with Figure 2 The method 200 for identity verification provided in this application will be described in detail. Figure 2 This is a schematic flowchart illustrating an identity recognition method provided in an embodiment of this application.
[0093] 210. Obtain the first near-infrared parameter at the first three-dimensional point of the target to be identified. The first near-infrared parameter indicates the information of near-infrared light reflected by the target at the position of the first three-dimensional point.
[0094] In one embodiment, the first near-infrared parameter can be obtained directly from the camera.
[0095] When a camera captures an image of a target, it acquires information about the target point by point, such as pixels or 3D points. The user device can then convert this information into a data format that image recognition software can recognize, thereby enabling the identification of the target in the image.
[0096] For example, the target to be identified can be divided into multiple three-dimensional points. One of the three-dimensional points, such as the first three-dimensional point, may include a first near-infrared parameter. In the embodiments of this application, the first near-infrared parameter is used to indicate information about the near-infrared light reflected by the target at the position of the first three-dimensional point.
[0097] Specifically, the camera can emit near-infrared light towards the target to be identified, and then receive parameters such as the intensity or angle of the reflected near-infrared light to obtain information about the near-infrared light reflected by the target. If the target to be identified is divided into multiple three-dimensional points, it can be assumed that the target will reflect near-infrared light at each three-dimensional point, and the user equipment will record the corresponding information about the reflected near-infrared light at each three-dimensional point. Taking the first three-dimensional point among multiple three-dimensional points as an example, the user equipment can process the information about the near-infrared light reflected by the target at the first three-dimensional point to obtain the first near-infrared parameters.
[0098] It should be noted that the information reflected back to the camera by near-infrared light is not limited to intensity or angle; it can also be other information that can characterize the near-infrared light reflected by the target to be identified. The first near-infrared parameter can be obtained by the user equipment after processing the above information, or the relevant parameters acquired by the camera can be directly used as the first near-infrared parameter.
[0099] In another embodiment, the first near-infrared parameter may also be calculated from other parameters.
[0100] For example, in practical identity recognition applications, the image directly acquired by the camera will not be exactly the same as the pre-recorded image. That is, for the same target to be identified, the acquired image may differ significantly in angle from the pre-recorded image, resulting in a high false positive rate for the user device during the recognition process. Therefore, the parameters directly acquired by the camera can be the second near-infrared parameters of the first three-dimensional point. The user device can rotate and / or translate the first three-dimensional point so that it is in a similar position in terms of angle and location to the corresponding point in the pre-recorded image. During the rotation and / or translation of the first three-dimensional point, it will not lose the second near-infrared parameters it carries. Therefore, after the first three-dimensional point is rotated and / or translated, the second near-infrared parameters can be used to form the first near-infrared parameters.
[0101] Specifically, a first rotation angle and / or a first translation amount can be generated using a convolutional neural network. The first three-dimensional point is rotated and / or translated according to the first rotation angle and / or the first translation amount. The second near-infrared parameter is also calculated accordingly based on the first rotation angle and / or the first translation amount.
[0102] Alternatively, the first three-dimensional point can be projected onto a plane to form a first pixel. Then, the first near-infrared parameter can be used to form a first parameter, which is used to indicate the information of near-infrared light reflected by the target to be identified at the position of the first pixel.
[0103] Specifically, during the process of acquiring information about the target to be identified, the near-infrared light reflected back by the camera can carry information such as intensity or angle, as well as the time from the emission of near-infrared light to the reception of the reflected near-infrared light at each pixel, in order to calculate the relative position between each pixel. In this way, the user device can obtain not only the near-infrared information of each pixel, but also the distribution information of the pixels in three dimensions. During identity recognition, to simplify the process, the near-infrared information distributed in three dimensions can be projected onto a plane to facilitate rapid identification by the user device.
[0104] For example, at the location of a first 3D point, the camera can directly acquire information about the near-infrared light reflected from that point, forming a first near-infrared parameter. It can also acquire the point's position information in three dimensions. The first near-infrared parameter can be a parameter directly acquired by the camera, or a parameter obtained by processing the parameters acquired by the camera using the user equipment. After acquiring the information about the near-infrared light reflected from the first 3D point, the user equipment can project the first 3D point onto a plane to form a first 3D point, and project the first near-infrared parameter of the first 3D point onto the plane to form a first parameter.
[0105] It should be noted that the first pixel is obtained by projecting the first three-dimensional point onto the plane. Therefore, on the same two-dimensional plane, the two-dimensional coordinate position of the first pixel is the same as the two-dimensional coordinate position of the first three-dimensional point. In other words, the first parameter and the first near-infrared parameter indicate the information of the reflected near-infrared light at the same point.
[0106] 220. Obtain the first three-dimensional parameters of the first three-dimensional point. The first three-dimensional parameters are used to indicate the position information of the first three-dimensional point on the target to be identified.
[0107] Similarly, the first three-dimensional parameters can be obtained directly from the camera or calculated from other parameters.
[0108] When a camera acquires information about a target to be identified on a point-by-point basis, it can send laser pulses to the target. Based on the position and / or angle of the emitted laser pulses, and the position and / or angle of the received reflected laser pulses, the relative position information between different points on the target can be determined.
[0109] It should be understood that when the camera acquires the position information of three-dimensional points on the target to be identified, it is not limited to emitting laser pulses to acquire position information; other methods can also be used. For example, in conjunction with the near-infrared light emitted in step 210, the position and / or angle of the reflected near-infrared light can also be used to determine the relative positions between different points on the target to be identified. This application embodiment only uses the emission of laser pulses as an example for illustration.
[0110] For example, if the camera receives a laser pulse reflected from the target at a first three-dimensional point, it can obtain the position information of the first three-dimensional point on the target, i.e., the first three-dimensional parameters, based on the position and / or angle of the emitted and reflected laser pulses.
[0111] The first three-dimensional parameter can be obtained by processing the information acquired by the camera using the user equipment, or it can be a parameter directly acquired by the camera. The first three-dimensional parameter includes the coordinate information of a first three-dimensional point in a three-dimensional angle. This application embodiment does not limit the information used to determine the second parameter to the position and / or angle of the laser pulse reflected back to the camera; other information can also be used to determine the first three-dimensional parameter.
[0112] Alternatively, the first three-dimensional parameter can be calculated from other parameters.
[0113] For example, if the position and angle of the target relative to the camera differ significantly from its position and angle when the image was captured, the user device will struggle to match the relevant features, potentially leading to recognition failure. Therefore, the parameters directly acquired by the camera can be the second three-dimensional parameters at the first three-dimensional point. Before identity recognition, the user device can rotate and / or translate the first three-dimensional point to bring it closer to the corresponding point in the pre-captured image in terms of angle and position.
[0114] It is understandable that rotating and / or translating the first 3D point actually involves calculating the rotation and / or translation of the second 3D parameters on that first 3D point. Specifically, a convolutional neural network can be used to generate a first rotation angle and / or a first translation amount, and the first 3D point can be rotated and / or translated according to the first rotation angle and / or the first translation amount. In other words, the second 3D parameters can be calculated according to the first rotation angle and / or the first translation amount, and the calculated result is the first 3D parameter.
[0115] Specifically, a first rotation angle and / or a first translation amount can be generated using a convolutional neural network, and the first three-dimensional point can be rotated and / or translated according to the first rotation angle and / or the first translation amount. The convolutional neural network is used to train the model; for example, in this embodiment, the model used to output the rotation angle and / or translation amount can be called an image straightening model. Inputting the second three-dimensional parameters of the first three-dimensional point into the image straightening model yields the first rotation angle and / or the first translation amount as output.
[0116] Understandably, an image straightening model can output only rotation angles, which can include pitch, yaw, and roll. Pitch refers to the angle of rotation around the X-axis, yaw around the Y-axis, and roll around the Z-axis. For example, based on the input of the second 3D parameter, the image straightening model can output the first rotation angle. The second 3D parameter is then calculated using the first rotation angle to obtain the first 3D parameter, which is equivalent to rotating the first 3D point by a certain rotation angle. When each 3D point on the target to be identified is rotated, it can be seen that the entire image has been rotated in three dimensions. The purpose of rotating the image is to rotate the image of the target to be identified to an angle similar to the angle of the recorded image when the angle of the target to be identified relative to the camera differs significantly from the angle of the recorded image, so as to facilitate the comparison of features at similar locations.
[0117] Image straightening models can also output only the translation amount. For example, based on the input of the second 3D parameter, the image straightening model can output the first translation amount. The second 3D parameter is then calculated using the first translation amount to obtain the first 3D parameter, which is equivalent to translating the first 3D point by a certain amount. When each 3D point on the target to be identified is translated, it can be seen as the entire image being translated in a certain direction. When the image of the target to be identified obtained by the camera differs significantly from the pre-recorded image only on the plane facing the camera, translation can bring them to similar positions within a certain range. This makes feature comparison more direct and improves the recognition efficiency of the user's device.
[0118] Image straightening models can also simultaneously output rotation angles and translation amounts. For example, based on the input of the second 3D parameters, the image straightening model can output a first rotation angle and a first translation amount, that is, it calculates the rotation and translation simultaneously on the second 3D parameters to obtain the first 3D parameters. When each 3D point on the target to be identified has been rotated and translated, it can be regarded as rotating and translating the entire image.
[0119] It should be understood that if only the image is rotated, the rotated image may still differ from the recorded image in a two-dimensional plane; similarly, if only the image is translated, the acquired image of the target to be identified may still differ from the recorded image in terms of angle. Therefore, combining the two methods results in a higher similarity between the acquired image of the target to be identified and the image pre-recorded by the camera, which is more conducive to image recognition by user devices, improving the accuracy and efficiency of identity recognition.
[0120] Optionally, when the camera acquires the second three-dimensional parameters of the first three-dimensional point, it can use the value directly acquired by the camera as the second three-dimensional parameter, or it can obtain the second three-dimensional parameter through calculation.
[0121] For example, the third 3D parameter can be directly acquired from the first 3D point using a camera. The first normal vector of the third 3D parameter is then calculated based on the normal vectors of the surfaces adjacent to the first 3D point. The components of the first normal vector on the X, Y, and Z axes are then calculated separately, and the normal vector is recalculated for each component along each axis to obtain the second 3D parameter. In other words, the second 3D parameter includes parameters on the X, Y, and Z axes. Since vectors have direction, this directionality can be used to describe the curvature of the surface of the target object to be identified.
[0122] In addition, if the first three-dimensional point is projected onto a plane to form a first pixel, the first three-dimensional parameters can be used to form a second parameter, wherein the second parameter is used to indicate the position information of the first pixel on the target to be identified.
[0123] Specifically, the information directly acquired by the camera can be the position parameters of three-dimensional points on the target to be identified in three-dimensional angles. These position parameters include not only the coordinates on the plane facing the camera, but also depth information in the direction perpendicular to that plane. In actual identity recognition, image features on the plane facing the camera are typically extracted and compared with pre-recorded images, with depth information from the three-dimensional image rarely used. Therefore, to avoid increased computational load due to redundant data, the position information acquired in three-dimensional angles can be projected onto a plane, meaning only position information in two dimensions is used for identity recognition.
[0124] For example, the camera can directly acquire first three-dimensional parameters at a first three-dimensional point that can be used to indicate position information. The first three-dimensional parameters include coordinate information on the plane facing the camera and depth information in the direction perpendicular to the plane facing the camera. The first three-dimensional point is projected onto the plane to form a first pixel, and the first three-dimensional parameters at the first three-dimensional point are projected onto the plane to form second parameters. The second parameters still include the parameter values of the first three-dimensional parameters in three dimensions, but only the parameters on the two-dimensional plane are provided during the identity recognition process to simplify the computational workload of the identity recognition process.
[0125] It should be noted that the first pixel is obtained by projecting the first three-dimensional point onto the plane. Therefore, on the two-dimensional plane, the two-dimensional coordinate position of the first pixel is the same as the two-dimensional coordinate position of the first three-dimensional point. In other words, the first three-dimensional parameter and the second parameter indicate the position information of the same point.
[0126] It should be understood that in the above-mentioned alternative solutions, whether the location information is acquired directly by a camera or calculated through other parameters, the parameters related to the location information include parameters on the X-axis, Y-axis, and Z-axis. Even when a 3D point is projected onto a plane, each point still includes parameters on the X-axis, Y-axis, and Z-axis; however, only parameters from two of these channels are used during the identity recognition process.
[0127] 230. Identify the first image of the target to be identified. The first image includes multiple first three-dimensional points.
[0128] In steps 210 and 220 above, the user equipment can acquire the first near-infrared parameter and the first three-dimensional parameter of the target to be identified at the position of the first three-dimensional point through a camera, or acquire other parameters first through a camera, and then calculate the first near-infrared parameter and the first three-dimensional parameter from the other parameters. The first near-infrared parameter indicates the information of near-infrared light reflected by the target at the position of the first three-dimensional point, including parameters on one channel; the first three-dimensional parameter indicates the position information of the first three-dimensional point on the target to be identified, including parameters on three channels. That is, the first three-dimensional point includes parameters on four channels, meaning that each pixel among multiple first three-dimensional points can include parameters on all four channels. Multiple first three-dimensional points on the target to be identified can form a first image, which the user equipment then identifies.
[0129] The first near-infrared parameter and the first three-dimensional parameter can be acquired simultaneously by the same camera, sequentially by the same camera, or acquired from the same point by different cameras, and then superimposed. Alternatively, other parameters used to calculate the first near-infrared parameter and the first three-dimensional parameter can be acquired simultaneously by the same camera, sequentially by the same camera, or related parameters from the same point acquired by different cameras. In other words, the first near-infrared parameter and the first three-dimensional parameter used for identity recognition are parameters from the same point.
[0130] Specifically, in an identity recognition system, a user needs to pre-enter an image of the target to be identified into a database. The user device can then extract feature values from this pre-entered image and store them in the database as well. When identity recognition is required, feature values can be extracted from a newly acquired image of the target to be identified and compared with the feature values from the pre-entered image. If the user device finds an image in the database that has the highest similarity to the newly acquired image of the target to be identified and exceeds a certain threshold, then the target to be identified and the image found in the database can be identified as the same target. The threshold can be set manually as needed.
[0131] For example, in step 230, feature values can be extracted from the first image, and the feature values on the first image can be compared with the feature values of images pre-recorded in the user device to find the image with the highest similarity that exceeds a certain threshold.
[0132] If, in the aforementioned steps, the user equipment projects the first three-dimensional point onto the plane to form the first pixel, then multiple first pixels can form a third image, and the user equipment performs recognition by recognizing the third image.
[0133] The identity recognition method provided in this application combines heterogeneous images—that is, different types of image forms of the same target—for identification. This allows for the verification of the target's authenticity from multiple dimensions, resulting in higher security and reducing the likelihood of misidentifying fake substitutes as the real target. Especially when the target is a live object such as a face, the method effectively avoids attacks from non-live objects like paper or masks, making identity recognition more secure and efficient. Furthermore, when the target's front is not directly facing the camera but at an angle, rotation and translation can be used to allow the user device to recognize the rotated image, improving the accuracy of identity recognition.
[0134] Alternatively, the identity of the target to be identified can also be identified using method 300, such as... Figure 3 As shown. Figure 3 This is a schematic flowchart illustrating another identity recognition method provided in an embodiment of this application, wherein... Figure 3 Steps 310 and 320 shown in the figure are similar to those in the figure. Figure 2 Steps 210 and 220 shown are the same and will not be repeated here.
[0135] 330. Obtain the second near-infrared parameters at the second three-dimensional point of the target to be identified. The second three-dimensional point is generated by a generative adversarial network, and the second near-infrared parameters indicate the information about the near-infrared light reflected by the target at the position of the second three-dimensional point.
[0136] As can be seen from the descriptions of steps 210 and 220, the first near-infrared parameter and the first three-dimensional parameter of the first three-dimensional point are both obtained directly or indirectly through a camera. In step 330, the second three-dimensional point is generated by a generative adversarial network (GAN). It should be understood that generating the second three-dimensional point using a GAN is actually generating the parameters of the second three-dimensional point.
[0137] For example, when a camera acquires 3D point information on a target to be identified, occlusion of the camera may result in multiple 3D points being acquired that are not complete representations of the target. Alternatively, the camera may only acquire 3D points facing the camera and not those facing away from it. After rotating these points, those points that were originally facing the camera may now be facing away, while points that were originally facing away and not acquired may now be facing the camera. This results in incomplete images. Consequently, when a user device performs identity recognition, it may be unable to extract prominent feature points from the image, leading to slow and inaccurate image identification. Therefore, generative adversarial networks (GANs) can be used to fill in these missing points.
[0138] Specifically, generative adversarial networks (GANs) can be used to train the model. For example, the model used to fill in the missing parameters in this embodiment can be called an image completion model. The image completion model takes an incomplete image as input, such as a first image containing multiple first three-dimensional points, and outputs a complete image, such as a second image containing multiple first three-dimensional points and multiple second three-dimensional points. The output of the second image by the image completion model can also be understood as the model generating second near-infrared parameters for the second three-dimensional points.
[0139] When using the image completion model to generate parameters for the second 3D point, since the image completion model is a pre-trained model, the output second image after inputting the first image can be considered as a complete image that has been completed within a certain error range.
[0140] The image completion model can generate the second near-infrared parameter of the second three-dimensional point based on the information of the first near-infrared parameter of the first three-dimensional point. For example, the second near-infrared parameter of the second three-dimensional point in a certain direction can be inferred based on the variation pattern of the first near-infrared parameter of multiple first three-dimensional points along a certain direction; or, for some symmetrical targets to be identified, the symmetry of the target itself can be used to symmetrically process the first near-infrared parameter of the first three-dimensional point before inputting it into the image completion model to obtain the second near-infrared parameter of the second three-dimensional point. It should be understood that the above methods for generating the second three-dimensional point from the first three-dimensional point are only examples and are not limited.
[0141] Optionally, method 300 may also include step 340.
[0142] 340. Obtain the second three-dimensional parameters of the second three-dimensional point on the target to be identified. The second three-dimensional parameters indicate the position information of the second three-dimensional point on the target to be identified.
[0143] The method for generating the second three-dimensional parameters at the second three-dimensional point in step 340 is the same as the method for generating the second near-infrared parameters at the second three-dimensional point in step 330. It also utilizes a generative adversarial network to train a model, such as an image completion model, and then inputs the first image into this model. The output is the second image, which includes multiple first three-dimensional points and multiple second three-dimensional points. The output of the second image by the image completion model can also be understood as the model generating the second three-dimensional parameters at the second three-dimensional point.
[0144] It should be noted that the parameter information generated on the second three-dimensional point in this embodiment may only include the second near-infrared parameter in step 330, that is, the parameter used to indicate the information of near-infrared light reflected by the target to be identified at the position of the second three-dimensional point. It is understood that in some cases, such as when the user equipment is recognizing an image projected on a plane, it is only necessary to supplement the missing parameters related to the information of reflected near-infrared light, without needing to supplement the specific position information of each missing point.
[0145] The parameter information on the second three-dimensional point may also include only the second three-dimensional parameters from step 340, i.e., the position information of the second three-dimensional point on the target to be identified. It is understood that in some cases, such as when the user equipment prefers to identify the surface contour of the target, only the missing parameters related to the position information can be supplemented, and the relevant information on the reflection of near-infrared light at each position may not be used as an important basis for identification.
[0146] The parameter information on the second three-dimensional point can also include both the second near-infrared parameter in step 330 and the second three-dimensional parameter in step 340. By combining the second near-infrared parameter and the second three-dimensional parameter, the information represented by the second three-dimensional point on the second image is more complete. The user equipment can also comprehensively obtain the parameter information of the missing part in the first image through the relevant parameters of reflected near-infrared light and the position information of the pixels on the target to be identified, thereby improving the accuracy of subsequent identity recognition.
[0147] In addition, image completion models can be used to complete the pixels projected onto the plane.
[0148] Specifically, the third image can be input into the image completion model to output a complete fourth image. The third image is the image formed by multiple first pixels after the first 3D point is projected onto the plane to form the first pixel; the fourth image is the image formed by multiple first pixels and multiple second pixels after the relevant parameters on the second pixel are generated by the image completion model.
[0149] For example, similar to how an image completion model generates relevant parameters on a second 3D point, a third and / or fourth parameter can be generated on the second pixel point. The third parameter is used to indicate the information of near-infrared light reflected by the target at the position of the second pixel point, and the fourth parameter is used to indicate the position information of the second pixel point on the target.
[0150] 350. Identify the second image of the target to be identified. The second image includes multiple first three-dimensional points and multiple second three-dimensional points.
[0151] Specifically, through steps 330 and 340 above, relevant parameters on the second three-dimensional point can be obtained. Multiple first three-dimensional points and multiple second three-dimensional points can form a second image. When the user equipment performs identity recognition, it recognizes the second image.
[0152] Optionally, if the first and second pixels obtained through the above steps have been projected onto the plane, then multiple first pixels and multiple second pixels can form a fourth image, and the user equipment recognizes the fourth image when performing identity recognition.
[0153] Optionally, after obtaining the second image, the user equipment may project both the first and second three-dimensional points in the second image onto a plane. The second three-dimensional points, after being projected onto the plane, can form third pixels. Multiple third pixels, together with multiple first pixels, form a fifth image, where the first pixels are formed by projecting the first three-dimensional points onto the plane. Then, the user equipment performs identification based on the fifth image.
[0154] The identity recognition method provided in this application can combine different image forms of the same target to identify images, and can verify the authenticity of the target from multiple dimensions, making identity recognition more secure and less likely to mistakenly identify fake substitutes as real targets. At the same time, by using a model to fill in missing parts of the original image, the user device can extract more feature points to compare with images in the database during the identity recognition process, thus improving the accuracy and efficiency of identity recognition.
[0155] It should be understood that the identity recognition method provided in this application embodiment can be applied to the recognition process of various targets. For example, the target to be recognized can be a symmetrical living object, such as a human face or an animal face; it can also be an asymmetrical living object, such as a human hand or an ear; or it can be a non-living object, such as daily necessities. It should also be understood that this application embodiment performs identity recognition on the target based on its image, not just on its attributes. That is, after comparing the identity recognition method provided in this application embodiment with images in a pre-recorded image database, it also needs to associate the content in the image with the identity to which it belongs. For example, a user device can identify the content in an image as a face, a hand, an animal face, or a daily necessities, and it also needs to associate the face or hand with a specific person, a person with an identity distinct from others. Similarly, if the identified content in the image is an animal face or a daily necessities, it also needs to associate the content in the image with a certain identity.
[0156] The following describes the method provided in this application embodiment using a face as an example. It should be understood that the method provided in this application embodiment can be well applied in the field of face recognition, but this application embodiment does not limit the scope of application of the method.
[0157] Figure 4 This is a schematic flowchart illustrating another identity recognition method provided in the embodiments of this application. Figure 4 The method 400 shown is a preferred example of this application, and the scope of protection of this application shall be determined by the scope of the claims.
[0158] 410. Obtain the first face image of the person to be identified.
[0159] It should be noted that, in practical applications, the method provided in this application embodiment can directly manipulate the parameters on three-dimensional points or pixels without first forming an image. The first face image and the second face image below are only for descriptive convenience and are used only to distinguish the various parameter information obtained in different steps. They do not limit the method provided in this application embodiment.
[0160] An image formed by multiple first three-dimensional points is called a first face image, where the first three-dimensional points include second three-dimensional parameters and second near-infrared parameters. The second three-dimensional parameters and second near-infrared parameters can be parameters directly acquired by the camera, or the second three-dimensional parameters can be obtained by vector calculation of the positional information related to the three-dimensional points directly acquired by the camera.
[0161] Specifically, in order to acquire parameters representing the first face image, the camera can be a 3D camera, such as a Time-of-Flight (TOF) camera. A TOF camera can obtain depth parameters for different points on the target by emitting laser pulses towards the target, receiving the reflected laser pulses, and calculating the time difference or phase difference between the emission and reflection of the laser pulses back to the camera. Based on the set of all points on the target, the first face image can be represented in three dimensions.
[0162] The laser pulses emitted by the 3D camera can be either infrared or near-infrared light. When the emitted laser pulses are near-infrared light, information about the near-infrared light reflected from the face to be identified can also be obtained using the near-infrared light, i.e., the second near-infrared parameter. The image quality obtained using near-infrared light is not affected by ambient light, and high-quality images can be obtained, especially in low-light environments.
[0163] The second near-infrared parameter mainly identifies the target by using a grayscale image of the face to be identified. The grayscale image reflects the contrast between light and dark areas of the target. Compared to using color images for identification, near-infrared images can avoid the problem of being unable to identify the target under ideal lighting conditions, such as strong light, weak light, backlight, or no light.
[0164] The second three-dimensional parameters on each first three-dimensional point acquired by the camera can indicate the relative position information between each three-dimensional point on the face to be identified, such as the three-dimensional coordinates of each three-dimensional point in the same coordinate system, or the coordinate vector of each three-dimensional point in the same coordinate system.
[0165] It should be understood that the second three-dimensional parameters and the second near-infrared parameters on the first face image can be obtained simultaneously using the same three-dimensional camera, or the second three-dimensional parameters and the second near-infrared parameters on the first face image can be obtained sequentially and simultaneously using the same three-dimensional camera, or two separate cameras can be used to obtain the two types of parameters on the same three-dimensional point. This application does not limit the specific methods used. The second three-dimensional parameters and the second near-infrared parameters are parameters on the same point. The second three-dimensional parameters include parameters from three channels, and the second near-infrared parameters include parameters from one channel; that is, each three-dimensional point includes parameters from four channels.
[0166] 420. Rotate and translate the first face image according to the rotation angle and / or translation amount output by the image normalization model to obtain the second face image. The image normalization model is trained by a convolutional neural network.
[0167] By inputting the parameter information of multiple first 3D points on the first face image into the image normalization model, the output can be a first rotation angle and a first translation amount. Since all 3D points on the first face image are in the same coordinate system, the output values of the rotation angle and translation amount at each 3D point are the same.
[0168] Based on a first rotation angle and a first translation amount, the parameters of the 3D points in the first face image are rotated and translated. The image formed by the rotated and translated 3D points is called the second face image. It should be understood that the 3D points in the second face image are still the same points as those in the first face image; only the data at those points has changed. That is, for the same 3D point, the coordinates of that 3D point in the second face image are calculated from the coordinates of that 3D point in the first face image. The coordinates of each 3D point are calculated to obtain new 3D point coordinates, and all the 3D points are combined to form the second face image.
[0169] In other words, the second three-dimensional parameters obtained in step 410 can be rotated and translated to obtain the first three-dimensional parameters, and the second near-infrared parameters can be rotated and translated to obtain the first near-infrared parameters.
[0170] The first rotation angle includes pitch, yaw, and roll, which adjust the position and orientation of three-dimensional points on the first face image at rotation angles around the X, Y, and Z axes, respectively. After these three angle adjustments, the face orientation in the second face image can be adjusted to be approximately the same as the face orientation in the pre-recorded face image sample. "Approximately the same" means that the similarity between the angles of some feature points and the angles of some feature points at the same location in the sample image is greater than or equal to a certain threshold.
[0171] The first translation is used to adjust the deviation between the rotated first face image and the sample image on the plane. For example, after rotation, the orientation of the first face image may be basically the same as the sample image, but they may not overlap, and there may be a certain deviation in the two-dimensional plane. In this case, translation can be used to make the rotated first face image have more overlap points with the sample image, so that face recognition can be performed more accurately in the subsequent feature value extraction and comparison process.
[0172] In the process of rotating and translating the first face image into the second face image, typical features on the face can be used for comparison to ensure that the angle difference between the second face image after being straightened and the face image samples in the database is small during subsequent recognition and comparison.
[0173] For example, feature points around the facial features can be selected such that the three-dimensional and near-infrared parameters of these feature points after rotation and translation are similar to the three-dimensional and near-infrared parameters of the feature points around the facial features in the sample images in the database, which is greater than or equal to a certain threshold.
[0174] Alternatively, if the position-related parameters of key points in the second face image, such as points around the facial features, have an error less than or equal to a certain threshold compared to the position-related parameters of key points in the template face image, then the second face image can be considered the output image after the first face image has been corrected. That is, the user device can pre-set template face images in a database so that during the process of correcting the first face image to the second face image, it can be corrected to a position easily recognizable by the user device.
[0175] 430. Obtain a third-party face image using an image completion model. The image completion model is trained using a generative adversarial network.
[0176] If, when a camera acquires information from a face image, the face's frontal direction is not directly facing the camera's direction, but rather at an angle, the user device will only obtain the parameters of the 3D points on the face facing the camera, and will not be able to acquire the 3D points on the side facing away from the camera. After the first face image is rotated and translated into a second face image, the 3D points that were originally facing the camera may rotate to the side facing away from the camera, while the data points that were originally facing away from the camera and were not acquired will rotate to the side facing the camera. These missing data points will result in a certain amount of data loss in the second face image. Therefore, it is necessary to fill in the parameters of the missing face points to form a complete face image, which is convenient for subsequent feature value extraction and comparison.
[0177] The parameters on the second face image can be supplemented in several ways. One method is to generate a second 3D point based on the second face image using an image completion model. The second 3D point may include a second 3D parameter and / or a second near-infrared parameter. The second 3D parameter indicates the position information of the second 3D point on the second face image, and the second near-infrared parameter indicates the information of near-infrared light that the face to be identified may reflect at the position of the second 3D point.
[0178] The relevant parameters on the second 3D point can be randomly generated within the range of the second face image and judged entirely by the discriminator in the generative adversarial network. When the judgment is true, the third face image is output. Alternatively, some possible parameter values on the second 3D point can be inferred based on the variation pattern of the relevant parameters on the first 3D point, and then these parameter values can be input into the model to obtain the third face image. Or, considering that the face has a certain degree of symmetry, the known relevant parameters on the first 3D point can be symmetrically processed before being input into the model to obtain the third face image.
[0179] Another way to complete the parameters on the second face image is to first project the first three-dimensional point onto a plane to form the first pixel, and then complete the image based on the image formed by multiple first pixels. The first near-infrared parameter then becomes the first parameter, used to indicate the information of near-infrared light reflected at the position of the target to be identified at the first pixel; the first three-dimensional parameter becomes the second parameter, used to indicate the position information of the first pixel on the face to be identified.
[0180] The image completion model completes the second pixel point, which includes a third parameter and a fourth parameter. The third parameter is used to indicate the information of the near-infrared light that the target to be identified may reflect at the position of the second pixel point, and the fourth parameter is used to indicate the position information of the second pixel point on the face to be identified.
[0181] The generation of the second pixel is similar to the generation of the second three-dimensional point. It can be generated randomly within the image range after projection onto the plane; or it can be generated by first inferring some possible parameter values of the second pixel based on the variation law of relevant parameters of the first pixel, and then corrected by using the image completion model; or it can be generated by symmetrically processing the known relevant parameters of the first pixel and then inputting them into the image completion model to obtain the relevant parameters of the output second pixel.
[0182] The image completion model processes the input second face image and outputs a third face image, which can also be understood as multiple first pixels and multiple second pixels forming a third face image.
[0183] It should be understood that the methods for obtaining third-party facial images described above are merely examples and are not intended to limit the scope of the methods.
[0184] 440. Perform face recognition based on a third-party face image.
[0185] After steps 410 to 430, a complete frontal image of the face to be identified, i.e., the third face image, is obtained. Feature values are then extracted from this third face image and compared with pre-recorded face image samples in the database to complete face recognition.
[0186] In one implementation, the third face image may include a first three-dimensional point and a second three-dimensional point, which are equivalent to the second image described in method 200 or method 300, and the user equipment can recognize the second image.
[0187] In another implementation, the third face image may include a first pixel and a second pixel. This corresponds to the fourth image described in method 200 or method 300, which the user equipment can recognize.
[0188] In another embodiment, a third face image including a first three-dimensional point and a second three-dimensional point can be projected onto a plane to form a fifth image as described in method 200 or method 300, which the user equipment can recognize.
[0189] In the method provided in this application embodiment, the three-dimensional coordinate parameters obtained during the face recognition process can be combined with near-infrared parameters to verify the authenticity of the face to be recognized from multiple dimensions, making face recognition more secure and less likely to mistakenly identify non-living objects such as masks or paper as real faces. Simultaneously, when the face image acquired by the camera is a large-angle profile, rotation, translation, and padding can be used to transform the large-angle profile into a frontal, complete face image that is easily recognized by the user device, improving the accuracy and efficiency of face recognition. This application embodiment also utilizes a neural network to train the model, ensuring that the error between the output parameters and the actual parameters is within a certain range, which also improves the accuracy and efficiency of face recognition.
[0190] In the identity recognition method provided in this application embodiment, a convolutional neural network can be used to train an image straightening model, and a generative adversarial network can be used to train an image completion model. The training methods of these two models are described below, wherein the image straightening model is the first model and the image completion model is the second model.
[0191] Figure 5 This is a schematic flowchart illustrating a method for training a neural network model provided in an embodiment of this application.
[0192] 510. Obtain the first training data and the first template data. The first training data includes the second three-dimensional parameters of each three-dimensional point on the target to be identified from different angles, and the first template data includes the second three-dimensional parameters of each three-dimensional point on the target to be identified from the first angle, where the first angle is the angle at which the target to be identified can be identified.
[0193] Specifically, before training the first model, the first template data needs to be entered. The first template data is used to verify whether the output data of the first model is close to the real data. It includes at least the second three-dimensional parameters of each three-dimensional point on the target to be identified. The first template data is usually the data that the camera can collect when the target to be identified is facing the camera directly, or it can be at a first angle to the camera. This first angle should also be the angle at which the target to be identified can be successfully identified.
[0194] Simultaneously, the initial training data needs to be acquired during model training. The initial training data includes data on the same target to be identified at different angles to the direction the camera is facing, such as the second three-dimensional parameters of each three-dimensional point on the target to be identified at different angles.
[0195] Multiple sets of initial training data from different perspectives are beneficial for adjusting parameters during model training, resulting in a more accurate trained model.
[0196] 520. Input the first training data into the first model to obtain the second rotation angle and / or the second translation amount.
[0197] In a convolutional neural network, certain parameters are pre-set for the first model. When the first training data is received as input, the second rotation angle and / or the second translation amount can be output.
[0198] Alternatively, the first training data and the first template data can be simultaneously input into the first model. The first model can then directly calculate the rotation angle and / or translation based on the parameters at points with similar positions between the two, and process the large number of calculated rotation angles and / or translations to obtain the second rotation angle and / or the second translation.
[0199] 530. Rotate and / or translate the first training data according to the second rotation angle and / or the second translation amount to obtain the second training data.
[0200] It should be understood that the first model outputs a rotation angle and / or translation amount, while the parameters used for identity recognition are three-dimensional parameters and / or near-infrared parameters. Furthermore, the first template data used to verify whether the output data of the first model is close to the real data also includes three-dimensional parameters on three-dimensional points. Therefore, it is necessary to rotate and / or translate the first training data according to the second rotation angle and / or second translation amount output by the first model to obtain the second training data, and then compare the second training data with the first template data.
[0201] 540. Calculate the first loss function value based on the difference between the second training data and the first template data.
[0202] The first loss function is typically used to characterize the difference between the data output by the model and the sample data. The value of the first loss function can be used to determine whether the error between the second training data and the real data is acceptable.
[0203] 550. Determine whether the value of the first loss function is less than or equal to the first threshold.
[0204] If so, proceed to step 551 and output the second rotation angle and / or the second translation amount;
[0205] If not, proceed to step 552 to adjust the parameters of the first model.
[0206] It is understandable that if the value of the first loss function is less than or equal to the first threshold, it means that the difference between the second training data obtained based on the output rotation angle and / or translation amount and the first template data is small. In the actual identity recognition process, identity recognition based on the second training data can obtain the expected recognition result with a high probability.
[0207] Conversely, if the value of the first loss function is greater than the first threshold, it indicates that there is a large difference between the second training data obtained from the output rotation angle and / or translation amount and the first template data. In this case, it is necessary to further adjust the parameters in the first model so that data with smaller error can be output.
[0208] During model training, the judgment in step 550 can be skipped, and the parameters of the first model can be adjusted directly based on the value of the first loss function within a certain number of loops.
[0209] Figure 6 This is a schematic flowchart illustrating another method for training a neural network model provided in an embodiment of this application.
[0210] 610. Obtain the first training image and the first template image. The first training image includes multiple incomplete images of the target to be identified, and the first template image is a complete image of the target to be identified.
[0211] 620. Input the first training image into the second model to obtain the first complete image.
[0212] 630. Calculate the second loss function value based on the difference between the first complete image and the first template image.
[0213] 640. Determine whether the value of the second loss function is less than or equal to the second threshold.
[0214] If so, proceed to step 641 and output the first complete image;
[0215] If not, proceed to step 642 to adjust the parameters of the second model.
[0216] Specifically, when training the second model, only the first training image needs to be input into the second model, and then the output image is continuously compared with the first template image until the discriminator in the generative adversarial network determines that the output image is a real image, at which point the image can be output. Step 630 above is the process of comparing the output image with the first template image, and step 640 is the process of the discriminator in the generative adversarial network making the judgment.
[0217] The above combines Figures 1 to 6 The technical solutions provided in the embodiments of this application have been described in detail below. Figures 7 to 8 This application describes a communication device provided in an embodiment.
[0218] Figure 7 This is a schematic diagram of an identity recognition device provided in an embodiment of this application. As shown in the figure, the device 700 includes an acquisition unit 701 and a processing unit 702.
[0219] In one implementation, device 700 can serve as an identity recognition device to achieve the combination described above. Figures 2 to 4 The method described.
[0220] In one embodiment, the acquisition unit 701 is used to acquire a first near-infrared parameter on a first three-dimensional point of the target to be identified. The first near-infrared parameter is used to indicate information about the near-infrared light reflected at the position of the target to be identified at the first three-dimensional point. The first three-dimensional parameter is used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified. The processing unit 702 is used to identify a first image of the target to be identified. The first image includes a plurality of first three-dimensional points.
[0221] In another embodiment, the acquisition unit 701 is used to acquire a second near-infrared parameter on a second three-dimensional point of the target to be identified. The second three-dimensional point is generated by a generative adversarial network, and the second near-infrared parameter is used to indicate information about the near-infrared light reflected by the target at the position of the second three-dimensional point. The processing unit 702 is used to identify a second image of the target to be identified. The second image includes a plurality of first three-dimensional points and a plurality of second three-dimensional points.
[0222] In another embodiment, the acquisition unit 701 is used to acquire second three-dimensional parameters on a second three-dimensional point of the target to be identified. The second three-dimensional point is generated by a generative adversarial network, and the second three-dimensional parameters are used to indicate the position information of the second three-dimensional point on the target to be identified. The processing unit 702 is used to identify a second image of the target to be identified. The second image includes a plurality of first three-dimensional points and a plurality of second three-dimensional points.
[0223] In another embodiment, the second three-dimensional point is generated by a generative adversarial network based on the first three-dimensional point.
[0224] In another embodiment, the second three-dimensional point is generated by a generative adversarial network after performing symmetric processing on the first three-dimensional point.
[0225] In another embodiment, the processing unit 702 is used to project a first three-dimensional point onto a plane to form a first pixel. The first pixel includes a first parameter and a second parameter. The first parameter is formed by a first near-infrared parameter, and the second parameter is formed by a first three-dimensional parameter. The processing unit 702 is used to identify a third image of the target to be identified. The third image includes a plurality of first pixels.
[0226] In another embodiment, the acquisition unit 701 is used to acquire a third parameter on a second pixel of the target to be identified, the second pixel being generated by a generative adversarial network, and the third parameter being used to indicate information about the near-infrared light reflected by the target at the position of the second pixel; the processing unit 702 is used to identify a fourth image of the target to be identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
[0227] In another embodiment, the acquisition unit 701 is used to acquire a fourth parameter on the second pixel of the target to be identified, the fourth parameter being used to indicate the position information of the second pixel on the target to be identified; the processing unit 702 is used to identify a fourth image of the target to be identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
[0228] In another embodiment, the processing unit 702 is used to project the second three-dimensional point onto a plane to form a third pixel point. The third pixel point includes a fifth parameter and a sixth parameter. The fifth parameter is formed by the second near-infrared parameter, and the sixth parameter is formed by the second three-dimensional parameter. The fifth image of the target to be identified is used to identify the target. The fourth image includes a plurality of first pixel points and a plurality of third pixel points.
[0229] In another embodiment, the acquisition unit 701 is used to acquire the second three-dimensional parameters of the first three-dimensional point, the second three-dimensional parameters being used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified; the processing unit 702 is used to generate a first rotation angle and / or a first translation amount using a convolutional neural network, and calculate the second three-dimensional parameters according to the first rotation angle and / or the first translation amount to obtain the first three-dimensional parameters.
[0230] In another embodiment, the acquisition unit 701 is used to acquire a second near-infrared parameter of a first three-dimensional point, the second near-infrared parameter being used to indicate information about the near-infrared light reflected at the position of the target to be identified at the first three-dimensional point; the processing unit 702 is used to generate a first rotation angle and / or a first translation amount using a convolutional neural network, and calculate the second near-infrared parameter according to the first rotation angle and / or the first translation amount to obtain the first near-infrared parameter.
[0231] In another embodiment, the acquisition unit 701 is used to acquire a third three-dimensional parameter on the first three-dimensional point, the third three-dimensional parameter being acquired by a camera; the processing unit 702 is used to calculate a first normal vector of the third three-dimensional parameter on the first three-dimensional point based on the normal vector of the adjacent surface of the first three-dimensional point, and to calculate the normal vectors of the components of the first normal vector on the X-axis, Y-axis and Z-axis respectively, the second three-dimensional parameter including the normal vectors on the X-axis, Y-axis and Z-axis.
[0232] In another implementation, device 700 can serve as a device for training a neural network model to achieve the combination described above. Figure 5 The method described.
[0233] In one embodiment, the acquisition unit 701 is used to acquire first training data and first template data. The first training data includes second three-dimensional parameters of each three-dimensional point on the target to be identified at different angles. The first template data is the second three-dimensional parameters of each three-dimensional point on the target to be identified at a first angle, where the first angle is the angle at which the target to be identified can be identified. The processing unit 702 is used to input the first training data into the first model to obtain a second rotation angle and / or a second translation amount, rotate and / or translate the first training data according to the second rotation angle and / or the second translation amount to obtain second training data, calculate a first loss function value based on the difference between the second training data and the first template data, and adjust the parameters of the first model based on the first loss function value.
[0234] In another embodiment, the processing unit 702 is configured to output a second rotation angle and / or a second translation amount when the value of the first loss function is less than or equal to a first threshold.
[0235] In another embodiment, the processing unit 702 is used to input the first training data and the first template data into the first model to obtain the second rotation angle and / or the second translation amount.
[0236] In another implementation, device 700 can serve as a device for training a neural network model to achieve the combination described above. Figure 6 The method described.
[0237] In one embodiment, the acquisition unit 701 is used to acquire a first training image and a first template image. The first training image includes multiple incomplete images of the target to be identified, and the first template image is a complete image of the target to be identified. The processing unit 702 is used to input the first training image into the second model to obtain a first complete image, calculate a second loss function value based on the difference between the first complete image and the first template image, and adjust the parameters of the second model based on the second loss function value.
[0238] In another embodiment, when the value of the second loss function is less than or equal to the second threshold, the processing unit 702 is used to output the first complete image.
[0239] Figure 8 This is a schematic diagram of the hardware structure of an identity recognition device provided in an embodiment of this application. Figure 8 The device 800 shown (which may specifically be a computer device) includes a memory 801, a processor 802, a communication interface 803, and a bus 804. The memory 801, processor 802, and communication interface 803 are interconnected via the bus 804.
[0240] In one implementation, device 800 can serve as an identity verification device.
[0241] The memory 801 can be a ROM, static storage device, or RAM. The memory 801 can store programs, and when the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 803 are used to execute the various steps of the animation generation method of this embodiment. Specifically, the processor 802 can execute the steps described above. Figure 2 Steps 210 to 230 in the method shown, or performing the above... Figure 3 Steps 310 to 350 in the method shown, or performing the above... Figure 4 Steps 410 to 440 in the method shown.
[0242] The processor 802 may be a general-purpose processor, CPU, microprocessor, ASIC, GPU, or one or more integrated circuits, used to execute relevant programs to achieve the functions required by the units in the identity recognition device of this application embodiment.
[0243] The processor 802 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the identity recognition method in this application embodiment can be accomplished through integrated logic circuits in the processor 802 or through software instructions.
[0244] The processor 802 described above can also be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 801. The processor 802 reads the information in memory 801 and, in conjunction with its hardware, completes the functions required by the units included in the identity recognition apparatus of the embodiments of this application, or executes the identity recognition method of the method embodiments of this application.
[0245] In another implementation, device 800 can be used as a training device for a neural network.
[0246] The memory 801 can be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 801 can store a program. When the program stored in the memory 801 is executed by the processor 802, the processor 802 executes the various steps of the neural network training method of this embodiment. Specifically, the processor 802 can execute the steps described above... Figure 5 Steps 520 to 550, as well as steps 551 and 552, in the method shown Figure 6 Steps 620 to 640, as well as steps 641 and 642, in the method shown.
[0247] The processor 802 may be a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute related programs to implement the neural network training method of the method embodiment of this application.
[0248] The processor 802 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the neural network training method of this application can be completed through the integrated logic circuitry in the processor 802 or through software instructions.
[0249] The processor 802 described above can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 801. The processor 802 reads the information in memory 801 and, in conjunction with its hardware, completes the functions required by the units included in the training device in the embodiments of this application, or executes the methods of the embodiments of this application. Figure 5 and Figure 6 The training method for the neural network is shown.
[0250] The communication interface 803 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between the device 800 and other devices or communication networks. For example, training data can be acquired through the communication interface 803.
[0251] Bus 804 may include a pathway for transmitting information between various components of device 800 (e.g., memory 801, processor 802, communication interface 803).
[0252] This application also provides a platform system that includes the aforementioned identity recognition device.
[0253] This application also provides a computer-readable medium having a computer program stored thereon, which, when executed by a computer, implements the method of any of the above method embodiments.
[0254] This application also provides a computer program product that, when executed by a computer, implements the method of any of the above method embodiments.
[0255] This application also provides an electronic device that may include the identity recognition device described in the above-described application embodiments.
[0256] For example, electronic devices include smart locks, mobile phones, computers, access control systems, and other devices that require identity verification. The identity verification device includes both software and hardware components used for identity verification within the electronic device.
[0257] Optionally, the electronic device may also include a near-infrared image acquisition device and / or a three-dimensional point cloud acquisition device.
[0258] It is understood that the near-infrared image acquisition device and / or the three-dimensional point cloud acquisition device can be any acquisition device in the related technology, and the embodiments of this application do not specifically limit it.
[0259] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0260] As used in this specification, the terms "unit," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0261] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0262] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0263] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0264] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0265] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0266] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identity recognition, characterized in that, include: Obtain a first near-infrared parameter at a first three-dimensional point of the target to be identified. The first near-infrared parameter is used to indicate information about the near-infrared light reflected by the target at the position of the first three-dimensional point. Obtain first three-dimensional parameters at the first three-dimensional point of the target to be identified. The first three-dimensional parameters are used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified. The first three-dimensional parameters include coordinate information on the plane facing the camera and depth information in the direction perpendicular to the plane facing the camera. Obtaining the first three-dimensional parameters at the first three-dimensional point includes: Obtain the second three-dimensional parameters of the first three-dimensional point. The second three-dimensional parameters are used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified. The second three-dimensional parameters include parameters on the three channels of the X-axis, Y-axis and Z-axis. The first rotation angle and / or the first translation amount are generated using the image straightening model; The second three-dimensional parameters are calculated based on the first rotation angle and / or the first translation amount to obtain the first three-dimensional parameters; The second near-infrared parameter is obtained at the second three-dimensional point of the target to be identified. The second three-dimensional point is generated by the image completion model. The second near-infrared parameter is used to indicate the information of near-infrared light reflected by the target at the position of the second three-dimensional point. Identifying the first image of the target to be identified, wherein the first image includes a plurality of first three-dimensional points, the identification of the first image of the target to be identified includes: The second image of the target to be identified is identified, and the second image includes a plurality of first three-dimensional points and a plurality of second three-dimensional points.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the second three-dimensional parameters on the second three-dimensional point of the target to be identified, wherein the second three-dimensional parameters are used to indicate the position information of the second three-dimensional point on the target to be identified; The process of identifying the first image of the target to be identified includes: The second image of the target to be identified is then identified.
3. The method according to claim 1 or 2, characterized in that, The second three-dimensional point is generated by the image completion model based on the first three-dimensional point.
4. The method according to claim 1 or 2, characterized in that, The second three-dimensional point is generated by the image completion model after performing symmetrical processing on the first three-dimensional point.
5. The method according to claim 1 or 2, characterized in that, The method further includes: The first three-dimensional point is projected onto a plane to form a first pixel. The first pixel includes a first parameter and a second parameter. The first parameter is formed by the first near-infrared parameter, and the second parameter is formed by the first three-dimensional parameter. The process of identifying the first image of the target to be identified includes: The third image of the target to be identified is identified, and the third image includes a plurality of the first pixels.
6. The method according to claim 5, characterized in that, The method further includes: Obtain a third parameter on the second pixel of the target to be identified, wherein the second pixel is generated by the image completion model, and the third parameter is used to indicate information about the near-infrared light reflected by the target at the position of the second pixel; The identification of the third image of the target to be identified includes: The fourth image of the target to be identified is identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
7. The method according to claim 6, characterized in that, The method further includes: Obtain a fourth parameter on the second pixel of the target to be identified, the fourth parameter being used to indicate the position information of the second pixel on the target to be identified; The identification of the third image of the target to be identified includes: The fourth image of the target to be identified is identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
8. The method according to claim 5, characterized in that, The method further includes: The second three-dimensional point is projected onto a plane to form a third pixel. The third pixel includes a fifth parameter and a sixth parameter. The fifth parameter is formed by the second near-infrared parameter, and the sixth parameter is formed by the second three-dimensional parameter. The identification of the second image of the target to be identified includes: The fifth image of the target to be identified is identified, and the fifth image includes a plurality of first pixels and a plurality of third pixels.
9. The method according to claim 7, characterized in that, Obtain the first near-infrared parameters at the first three-dimensional point, including: Obtain the second near-infrared parameter of the first three-dimensional point, the second near-infrared parameter being used to indicate information about the near-infrared light reflected by the target to be identified at the position of the first three-dimensional point; The first rotation angle and / or the first translation amount are generated using the image straightening model; The second near-infrared parameter is calculated based on the first rotation angle and / or the first translation amount to obtain the first near-infrared parameter.
10. The method according to claim 6, characterized in that, Obtaining the second three-dimensional parameters of the first three-dimensional point includes: Obtain the third three-dimensional parameter on the first three-dimensional point, wherein the third three-dimensional parameter is acquired by the camera; Calculate the first normal vector of the third three-dimensional parameter on the first three-dimensional point based on the normal vector of the plane adjacent to the first three-dimensional point; The normal vectors of the components of the first normal vector on the X-axis, Y-axis and Z-axis are calculated respectively. The second three-dimensional parameter includes the normal vectors on the X-axis, Y-axis and Z-axis.
11. The method according to claim 1, characterized in that, The image straightening model is trained according to the following steps: Acquire first training data and first template data. The first training data includes the second three-dimensional parameters of each three-dimensional point on the target to be identified from different angles. The first template data is the second three-dimensional parameters of each three-dimensional point on the target to be identified from the first angle. The first angle is the angle at which the target to be identified can be identified. The first training data is input into the image straightening model to obtain the second rotation angle and / or the second translation amount; The first training data is rotated and / or translated according to the second rotation angle and / or the second translation amount to obtain the second training data; The first loss function value is calculated based on the difference between the second training data and the first sample data; The parameters of the image straightening model are adjusted based on the first loss function value.
12. The method according to claim 11, characterized in that, The method further includes: When the value of the first loss function is less than or equal to the first threshold, the second rotation angle and / or the second translation amount are output.
13. The method according to claim 11 or 12, characterized in that, The step of inputting the first training data into the image straightening model to obtain the second rotation angle and / or the second translation amount includes: The first training data and the first template data are input into the image straightening model to obtain the second rotation angle and / or the second translation amount.
14. The method according to claim 1, characterized in that, The image completion model is trained according to the following steps: Acquire a first training image and a first template image, wherein the first training image includes multiple incomplete images of the target to be identified, and the first template image is a complete image of the target to be identified; The first training image is input into the image completion model to obtain the first complete image; The second loss function value is calculated based on the difference between the first complete image and the first sample image; The parameters of the image completion model are adjusted based on the value of the second loss function.
15. The method according to claim 14, characterized in that, The method further includes: When the value of the second loss function is less than or equal to the second threshold, the first complete image is output.
16. An identity recognition device, characterized in that, include: The acquisition unit is used to acquire a first near-infrared parameter on a first three-dimensional point of the target to be identified, wherein the first near-infrared parameter is used to indicate information about the near-infrared light reflected by the target at the position of the first three-dimensional point; The acquisition unit is used to acquire the first three-dimensional parameters on the first three-dimensional point of the target to be identified. The first three-dimensional parameters are used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified. The first three-dimensional parameters include coordinate information on the plane facing the camera and depth information in the direction perpendicular to the plane facing the camera. The acquisition unit is used to acquire the second three-dimensional parameters of the first three-dimensional point. The second three-dimensional parameters are used to indicate the three-dimensional position information of the first three-dimensional point on the target to be identified. The second three-dimensional parameters include parameters on three channels: X-axis, Y-axis and Z-axis. The acquisition unit is used to acquire the second near-infrared parameter on the second three-dimensional point of the target to be identified. The second three-dimensional point is generated by the image completion model. The second near-infrared parameter is used to indicate the information of the target to be identified reflecting near-infrared light at the position of the second three-dimensional point. A processing unit is used to identify a first image of the target to be identified, wherein the first image includes a plurality of the first three-dimensional points; The processing unit is used to identify the second image of the target to be identified, the second image including a plurality of first three-dimensional points and a plurality of second three-dimensional points; The processing unit is used to generate a first rotation angle and / or a first translation amount using an image straightening model; The processing unit is used to calculate the second three-dimensional parameters according to the first rotation angle and / or the first translation amount to obtain the first three-dimensional parameters.
17. The apparatus according to claim 16, characterized in that, The acquisition unit is used to acquire the second three-dimensional parameters on the second three-dimensional point of the target to be identified, and the second three-dimensional parameters are used to indicate the position information of the second three-dimensional point on the target to be identified; The processing unit is used to identify the second image of the target to be identified.
18. The apparatus according to claim 16 or 17, characterized in that, The second three-dimensional point is generated by the image completion model based on the first three-dimensional point.
19. The apparatus according to claim 16 or 17, characterized in that, The second three-dimensional point is generated by the image completion model after performing symmetrical processing on the first three-dimensional point.
20. The apparatus according to claim 16 or 17, characterized in that, The processing unit is used to project the first three-dimensional point onto a plane to form a first pixel. The first pixel includes a first parameter and a second parameter. The first parameter is formed by the first near-infrared parameter, and the second parameter is formed by the first three-dimensional parameter. The processing unit is used to identify a third image of the target to be identified, the third image including a plurality of the first pixels.
21. The apparatus according to claim 20, characterized in that, The acquisition unit is used to acquire a third parameter on the second pixel of the target to be identified. The second pixel is generated by the image completion model. The third parameter is used to indicate the information of near-infrared light reflected by the target to be identified at the position of the second pixel. The processing unit is used to identify the fourth image of the target to be identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
22. The apparatus according to claim 21, characterized in that, The acquisition unit is used to acquire a fourth parameter on the second pixel of the target to be identified, and the fourth parameter is used to indicate the position information of the second pixel on the target to be identified; The processing unit is used to identify the fourth image of the target to be identified, the fourth image including a plurality of first pixels and a plurality of second pixels.
23. The apparatus according to claim 20, characterized in that, The processing unit is used to project the second three-dimensional point onto a plane to form a third pixel point. The third pixel point includes a fifth parameter and a sixth parameter. The fifth parameter is formed by the second near-infrared parameter, and the sixth parameter is formed by the second three-dimensional parameter. The processing unit is used to identify the fifth image of the target to be identified, the fifth image including a plurality of first pixels and a plurality of third pixels.
24. The apparatus according to claim 22, characterized in that, The acquisition unit is used to acquire the second near-infrared parameter of the first three-dimensional point, and the second near-infrared parameter is used to indicate the information of the near-infrared light reflected by the target to be identified at the position of the first three-dimensional point; The processing unit is used to generate a first rotation angle and / or a first translation amount using an image straightening model; The processing unit is used to calculate the second near-infrared parameter according to the first rotation angle and / or the first translation amount to obtain the first near-infrared parameter.
25. The apparatus according to claim 21, characterized in that, The acquisition unit is used to acquire a third three-dimensional parameter on the first three-dimensional point, and the third three-dimensional parameter is acquired by the camera. The processing unit is used to calculate the first normal vector of the third three-dimensional parameter on the first three-dimensional point based on the normal vector of the adjacent surface of the first three-dimensional point; The processing unit is used to calculate the normal vectors of the components of the first normal vector on the X-axis, Y-axis and Z-axis respectively, and the second three-dimensional parameter includes the normal vectors on the X-axis, Y-axis and Z-axis.
26. The apparatus according to claim 16, characterized in that, The acquisition unit is used to acquire first training data and first template data. The first training data includes the second three-dimensional parameters of each three-dimensional point on the target to be identified from different angles. The first template data is the second three-dimensional parameters of each three-dimensional point on the target to be identified from a first angle. The first angle is the angle at which the target to be identified can be identified. The processing unit is used to input the first training data into the image straightening model to obtain a second rotation angle and / or a second translation amount; The processing unit is used to rotate and / or translate the first training data according to the second rotation angle and / or the second translation amount to obtain the second training data; The processing unit is used to calculate a first loss function value based on the difference between the second training data and the first sample data; The processing unit is used to adjust the parameters of the image straightening model based on the first loss function value.
27. The apparatus according to claim 26, characterized in that, The processing unit is used to output the second rotation angle and / or the second translation amount when the value of the first loss function is less than or equal to the first threshold.
28. The apparatus according to claim 26 or 27, characterized in that, The processing unit is used to input the first training data and the first template data into the image straightening model to obtain the second rotation angle and / or the second translation amount.
29. The apparatus according to claim 16, characterized in that, The acquisition unit is used to acquire a first training image and a first template image. The first training image includes multiple incomplete images of the target to be identified, and the first template image is a complete image of the target to be identified. The processing unit is used to input the first training image into the image completion model to obtain the first complete image; The processing unit is used to calculate a second loss function value based on the difference between the first complete image and the first template image; The processing unit is used to adjust the parameters of the image completion model based on the second loss function value.
30. The apparatus according to claim 29, characterized in that, When the value of the second loss function is less than or equal to the second threshold, the processing unit outputs the first complete image.
31. An electronic device, characterized in that, include: The device for identity recognition as described in any one of claims 16 to 30.
32. A computer-readable storage medium, characterized in that, Used to store program instructions, which, when executed by a computer, perform the method as described in any one of claims 1 to 15.
33. A computer program product containing instructions, characterized in that, When the instructions are executed by a computer, the computer performs the method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Three-dimensional face recognition method and device, terminal equipment and computer readable medium
CN110852310A
Image processing method, device and equipment and storage medium
CN111340943A