Identity recognition method, device and computer-readable storage medium

By performing live body recognition and face recognition on the back-end device, the problems of complexity and difficulty in development in the existing technology are solved, and a simplified identity recognition process and convenient identification effect are achieved.

CN114154093BActive Publication Date: 2025-06-24GUANGZHOU YUNCONG DINGWANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111413651.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-06-24
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

In the prior art, the identity recognition method requires living body recognition and face recognition on the front-end device and the back-end server respectively, resulting in complex processes and the need to develop different software for different types of front-end devices, which increases the development difficulty.

Method used

The video acquisition device of the front-end device is called through the HTML page to collect video information, and the video information is sent to the back-end device for live recognition and face recognition. The back-end device determines the identity identification result based on the identification result and sends it to the front-end device.

Benefits of technology

The process of identity identification methods is simplified, and the need to develop different identification software for different types of front-end devices is avoided, thus achieving convenient and reliable identity identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154093B_ABST
    Figure CN114154093B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video processing, and specifically provides an identity recognition method, device, and computer-readable storage medium, aiming to solve the problem of how to perform identity recognition conveniently and reliably. For this purpose, the method of the present invention includes: receiving video information sent by a front-end device, where the video information is the video information collected by the front-end device through its own video acquisition device by calling an HTML page; respectively performing liveness recognition and face recognition on the video information; determining an identity recognition result according to the results of liveness recognition and face recognition and sending the identity recognition result to the front-end device. Even if the types of front-end devices are different, the video acquisition device can be called through an HTML page to collect video information. At the same time, the front-end device only needs to receive the identity recognition result sent by the back-end device after completing liveness recognition and identity recognition. Therefore, there is no need to develop different recognition software for different types of front-end devices, and thus identity recognition can be completed conveniently and reliably.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and specifically provides an identity recognition method, device, and computer-readable storage medium. Background Art

[0002] In application scenarios such as financial services and security management, it is usually necessary to identify the identity of personnel. When performing identity recognition, two recognition methods, namely face recognition and liveness recognition, can be used simultaneously for identity recognition. That is, only when both face recognition and liveness recognition are passed can it be determined that the identity recognition is successful. Currently, the identity recognition method that simultaneously uses the above two recognition methods mainly installs liveness recognition software in the front-end device. Through this liveness recognition software, the camera device of the front-end device is called to collect the image information of the personnel, and liveness recognition is performed based on the collected image information. After the liveness recognition is passed, this image information is sent to the back-end server. The back-end server will perform face recognition on this image information and send the identity recognition success information to the front-end device for output after the face recognition is passed. Since it is necessary to perform liveness recognition and face recognition separately in the front-end device and the back-end server to complete the identity recognition, the complexity of the method process is increased. At the same time, it is also necessary to develop different liveness recognition software for different types of front-end devices, which also increases the development and design difficulty of the identity recognition method. Summary of the Invention

[0003] In order to overcome the above-mentioned defects, the present invention is proposed to provide an identity recognition method, device, and computer-readable storage medium that solve or at least partially solve the technical problem of how to conveniently and reliably complete the identity recognition of personnel through liveness recognition and face recognition.

[0004] In a first aspect, the present invention provides an identity recognition method, which is applied to a back-end device, and the method includes:

[0005] Receiving video information sent by a front-end device, where the video information is the video information collected by the front-end device through its own video acquisition device by calling an HTML page;

[0006] Performing liveness recognition and face recognition on the video information respectively;

[0007] Determining an identity recognition result according to the results of liveness recognition and face recognition and sending the identity recognition result to the front-end device.

[0008] In a technical solution of the above identity recognition method, the step of "performing liveness recognition on the video information" specifically includes:

[0009] Obtain video features of different types in the video information, classify actions based on all video features to determine the type of action actually completed by the person in the video information;

[0010] If the type of action actually completed is consistent with the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body;

[0011] And / or,

[0012] The step of "performing face recognition on the video information" specifically includes:

[0013] After determining that the person in the video information is a live body through the live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0014] In a technical solution of the above identity recognition method, the video information is the action key frames extracted from the video information after the front-end device collects the video information by calling its own video acquisition device through an HTML page;

[0015] And / or, the HTML page includes at least an HTML5 page.

[0016] In a second aspect, an identity recognition method is provided, which is applied to a front-end device, and the method includes:

[0017] Call its own video acquisition device through an HTML page to collect video information and send the video information to the back-end device;

[0018] Receive the identity recognition result fed back by the back-end device according to the video information, where the identity recognition result is the identity recognition result determined by the back-end device after respectively performing live body recognition and face recognition on the video information.

[0019] In a technical solution of the above identity recognition method, the step of "sending the video information to the back-end device" specifically includes:

[0020] Extract the action key frames in the video information and send the action key frames to the back-end device;

[0021] And / or, the HTML page includes at least an HTML5 page.

[0022] In a technical solution of the above identity recognition method, the back-end device is configured to perform live body recognition on the video information by executing the following steps:

[0023] Obtain video features of different types in the video information, classify actions based on all the video features to determine the type of action actually completed by the person in the video information;

[0024] If the type of action actually completed is the same as the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body;

[0025] and / or,

[0026] The back-end device is further configured to perform face recognition on the video information by executing the following steps:

[0027] After determining that the person in the video information is a live body through the live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0028] In a third aspect, there is provided an identity recognition device applied to a back-end device, and the device includes:

[0029] A first information receiving module, configured to receive video information sent by a front-end device, where the video information is the video information collected by the front-end device through its own video collection device by calling an HTML page;

[0030] An information recognition module, configured to perform live body recognition and face recognition on the video information respectively;

[0031] A result processing module, configured to determine an identity recognition result according to the results of the live body recognition and the face recognition and send the identity recognition result to the front-end device.

[0032] In a technical solution of the above identity recognition device, the video information is an action key frame extracted from the video information by the front-end device after collecting the video information through its own video collection device by calling an HTML page;

[0033] and / or, the HTML page includes at least an HTML 5 page;

[0034] and / or, the information recognition module includes a live body recognition sub-module and / or a face recognition sub-module;

[0035] The live body recognition sub-module is configured to perform the following operations:

[0036] Obtain video features of different types in the video information, classify actions based on all the video features to determine the type of action actually completed by the person in the video information;

[0037] If the type of the actually completed action is the same as the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body.

[0038] The face recognition sub-module is configured to perform the following operations:

[0039] After determining that the person in the video information is a live body through the live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0040] In a fourth aspect, there is provided an identity recognition device, which is applied to a front-end device. The device includes:

[0041] An information collection module, which is configured to collect video information through its own video collection device by calling an HTML page and send the video information to a back-end device;

[0042] A second information receiving module, which is configured to receive the identity recognition result fed back by the back-end device according to the video information. The identity recognition result is the identity recognition result determined by the back-end device based on the results of live body recognition and face recognition of the video information respectively.

[0043] In a technical solution of the above identity recognition device, the information collection module is further configured to extract the action key frames in the video information and send the action key frames to the back-end device;

[0044] And / or, the HTML page includes at least an HTML 5 page;

[0045] And / or, the back-end device is configured to perform live body recognition on the video information by executing the following steps:

[0046] Obtain different types of video features in the video information, classify actions according to all video features to determine the type of the action actually completed by the person in the video information;

[0047] If the type of the actually completed action is the same as the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body;

[0048] And / or, the back-end device is further configured to perform face recognition on the video information by executing the following steps:

[0049] After determining that the person in the video information is a live body through the live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0050] In a fifth aspect, a control device is provided. The control device includes a processor and a storage device. The storage device is adapted to store multiple program codes, and the program codes are adapted to be loaded and run by the processor to execute the identity recognition method according to any one of the technical solutions of the above identity recognition method.

[0051] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores multiple program codes, and the program codes are adapted to be loaded and run by a processor to execute the identity recognition method according to any one of the technical solutions of the above identity recognition method.

[0052] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:

[0053] In an embodiment of the present invention, the identity recognition method can be applied to a backend device. The identity recognition method may include the following steps: receiving video information sent by a frontend device, where the video information may be video information collected by the frontend device through its own video capture device by calling an HTML page; performing liveness recognition and face recognition on the video information respectively; determining an identity recognition result according to the results of liveness recognition and face recognition and sending the identity recognition result to the frontend device. The HTML page refers to a web page that can be loaded by the frontend device created using Hyper Text Markup Language (HTML) in the field of computer technology. By directly calling the frontend device's own video capture device through an HTML page to collect video information of a person, it can be applied to all frontend devices that can load HTML pages. Even if the types of frontend devices are different, the video capture device can be called through the HTML page to collect video information. At the same time, after the frontend device collects the video information, it will send the video information to the backend device. The backend device completely completes liveness recognition and face recognition, and the frontend device only needs to receive the identity recognition result determined by the backend device after completing liveness recognition and face recognition, which greatly simplifies the method process of the identity recognition method. At the same time, it also overcomes the defect in the prior art that different recognition software needs to be developed for different types of frontend devices, so that the identity recognition of a person can be conveniently and reliably completed through liveness recognition and face recognition. Further, the frontend device can extract key action frames from the video information and send the key action frames to the backend device for liveness recognition and face recognition. By extracting key action frames, the interference of non-key action frames on liveness recognition can be removed, and the accuracy of liveness recognition can be improved.

[0054] Further, in another technical solution of implementing the present invention, the identity recognition method can be applied to a front-end device, and the identity recognition method may include the following steps: calling the video acquisition device of itself through an HTML page to acquire video information and sending the video information to a back-end device; receiving the identity recognition result fed back by the back-end device according to the video information, where the identity recognition result is the identity recognition result determined by the back-end device based on the results of live body recognition and face recognition of the video information respectively. The beneficial effects of this technical solution are similar to those of the foregoing technical solution and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] With reference to the accompanying drawings, the disclosure of the present invention will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. Among them:

[0056] Figure 1 is a schematic diagram of the main step flow of an identity recognition method according to an embodiment of the present invention;

[0057] Figure 2 is a schematic diagram of the main step flow of a live body recognition method according to an embodiment of the present invention;

[0058] Figure 3 is a schematic diagram of the main step flow of an identity recognition method according to another embodiment of the present invention;

[0059] Figure 4 is a schematic diagram of the main step flow of an identity recognition method according to still another embodiment of the present invention;

[0060] Figure 5 is a schematic diagram of the main step flow of an identity recognition method according to yet another embodiment of the present invention;

[0061] Figure 6 is a schematic diagram of an application scenario according to the present invention;

[0062] Figure 7 is Figure 6 a schematic diagram of a functional interface of live body detection in the application scenario shown;

[0063] Figure 8 is Figure 6 another schematic diagram of a functional interface of live body detection in the application scenario shown;

[0064] Figure 9 is a schematic diagram of the main structural block diagram of an identity recognition device according to an embodiment of the present invention;

[0065] Figure 10Schematic diagram of the main structure of an identity recognition device according to another embodiment of the present invention. Detailed implementation manners

[0066] Some embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.

[0067] In the description of the present invention, a "module" and a "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, memories, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A non-transitory computer-readable storage medium includes any suitable medium for storing program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "this" may also include the plural form.

[0068] Refer to the attached Figure 1 , Figure 1 Schematic diagram of the main step flow of an identity recognition method according to an embodiment of the present invention. In the embodiment of the present invention, the identity recognition method can be applied to a backend device, which refers to a device that does not directly interact with a person when performing identity recognition on the person. The device may be a computer, a server, etc. As Figure 1 shown, the identity recognition method in the embodiment of the present invention mainly includes the following steps S101-step S103.

[0069] Step S101: Receive video information sent by a front-end device.

[0070] The front-end device refers to a device that directly interacts with a person when performing identity recognition on the person. The device may be a mobile device such as a person's mobile phone, a mobile computer, etc., or a non-mobile device such as a device fixedly installed at a certain position. Through the video acquisition device of the front-end device, video information of the person during identity recognition can be acquired.

[0071] Video information refers to the video information collected by the front-end device through its own video acquisition device via an HTML page. The HTML page refers to a web page that can be loaded by the front-end device and is created using Hyper Text Markup Language (HTML) in the field of computer technology. In this embodiment, a conventional HTML page in the field of computer technology can be used to create the HTML page, which will not be elaborated here. The HTML page can at least include an HTML5 page, that is, a web page created using the fifth-generation hypertext markup language and can be loaded by the front-end device.

[0072] When performing identity recognition on a person, the person can execute the specified type of action within the acquisition range of the video acquisition device of the front-end device according to the specified type of action prompted by the front-end device. The front-end device can call the video acquisition device through the HTML page to collect the video information of the person when performing this specified type of action.

[0073] Step S102: Perform liveness recognition and face recognition on the video information respectively.

[0074] Liveness recognition refers to recognizing whether the person in the video information is a live person, and face recognition refers to recognizing which person in the preset person library the person in the video information is.

[0075] Step S103: Determine the identity recognition result according to the results of liveness recognition and face recognition and send the identity recognition result to the front-end device.

[0076] When the result of liveness recognition is that the person in the video information is a live person and the result of face recognition is that the person in the video information is a person in the preset person library, it can be determined that this person has successfully passed the identity recognition. The identity recognition result can include information on whether this person has passed the identity recognition, the face image of this person, and in addition, it can also include the person information of this person in the preset person library, and the person information can uniquely indicate which person this person is.

[0077] Based on the above steps S101 - S103, the video information of a person can be collected by directly calling the video capture device of the front - end device through an HTML page. This can be applied to all front - end devices that can load an HTML page. Even if the types of front - end devices are different, the video capture device can be called through the HTML page to collect video information. At the same time, after the front - end device collects the video information, it will send the video information to the back - end device. The back - end device completely completes the live body recognition and face recognition, and the front - end device only needs to receive the identity recognition result determined by the back - end device after completing the live body recognition and face recognition. This greatly simplifies the method flow of the identity recognition method and also overcomes the defect in the prior art that different recognition software needs to be developed for different types of front - end devices, so that the identity of a person can be conveniently and reliably recognized.

[0078] The above steps S101 and S102 will be further described below.

[0079] In one embodiment of the above step S101, the front - end device can extract the action key frames from the video information after collecting the video information by calling its own video capture device through an HTML page, and then send the extracted action key frames to the back - end device. The action key frame refers to the video frame in which the video information contains the picture of a person performing an action. By extracting the action key frames, the interference of non - action key frames on the live body recognition can be removed, and the accuracy of the live body recognition can be improved.

[0080] In one embodiment, the inter - frame distance between any two video frames in the video information can be calculated by the method shown in the following formula (1), and then the two video frames with the largest inter - frame distance are selected as the action key frames.

[0081]

[0082] The meanings of the parameters in formula (1) are as follows:

[0083] x (m) represents the m - th video frame, x (n) represents the n - th video frame, dis(x (m) , x (n) ) represents the inter - frame distance between the video frame x (m) and the video frame x (n) , N represents the total number of part nodes of a person in the video information, and the part nodes include but are not limited to: face nodes such as eyes and mouth, and non - face nodes such as hands and arms; represents the abscissa of the feature point corresponding to the i - th node in the m - th video frame, represents the abscissa of the feature point corresponding to the i - th node in the n - th video frame, represents the ordinate of the feature point corresponding to the i - th node in the m - th video frame, Denotes the vertical coordinate of the feature point corresponding to the \(i\)-th node in the \(n\)-th video frame, ω i Denotes the contribution degree of the \(i\)-th node, ω i The calculation formula is shown in the following formula (2):

[0084]

[0085] The meanings of the parameters in formula (2) are as follows:

[0086] Denotes the coordinate variance of the \(i\)-th node Denotes the sum of the coordinate variances of all nodes, and the variance of the \(i\)-th node The calculation formula is shown in the following formula (3):

[0087]

[0088] The meanings of the parameters in formula (3) are as follows:

[0089] \(k\) denotes the total number of video frames Denotes the horizontal coordinate of the feature point corresponding to the \(i\)-th node in the \(f\)-th video frame Denotes the vertical coordinate of the feature point corresponding to the \(i\)-th node in the \(f\)-th video frame Denotes the average value of the horizontal coordinates of the feature points corresponding to the \(i\)-th node in \(k\) video frames Denotes the average value of the vertical coordinates of the feature points corresponding to the \(i\)-th node in \(k\) video frames

[0090] In an embodiment of the above step S102, the video information can be subjected to liveness recognition through the following steps 11 to 12

[0091] Step 11: Obtain different types of video features in the video information, classify the actions according to the video features, so as to determine the type of action actually completed by the person in the video information

[0092] The types of video features include but are not limited to: video features extracted based on RGB image information, video features extracted based on optical flow information, and video features extracted based on node coordinate information. RGB image information refers to the red, green, and blue primary color information extracted from the video frames in the video information. Optical flow information refers to the optical flow information extracted from the video frames in the video information by using the optical flow method in the field of image processing technology. Node coordinate information refers to the horizontal and vertical coordinates of the feature points corresponding to the body part nodes of the person in the video frames in the embodiment described in the foregoing step S101

[0093] It should be noted that in this embodiment, a classification model capable of performing action classification based on the above new video features can be trained first using a conventional image classification method in the field of machine technology, and then this classification model can be used to perform action classification on the input video features. The model structure and training method of the classification model will not be elaborated here.

[0094] Step 12: If the type of the actually completed action is the same as the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body.

[0095] Refer to the appendix Figure 2 In the live body recognition method according to an embodiment of the present invention, when the key frame sequence sent by the front-end device is received, the RGB image information, optical flow information, and node coordinate information of each key frame in the key frame sequence can be extracted in sequence, and then the RGB image information, optical flow information, and node coordinate information are respectively input into the spatial information flow network, temporal information flow network, and spatio-temporal graph convolutional network for video feature extraction, so as to obtain the video features extracted based on the RGB image information, the video features extracted based on the optical flow information, and the video features extracted based on the node coordinate information. Further, after fusing the above video feature information, action classification can be performed according to the newly formed video features after fusion to obtain a classification result, that is, the type of the action actually completed by the person in the video information.

[0096] It should be noted that the above spatial information flow network, temporal information flow network, and spatio-temporal graph convolutional network can all adopt conventional network models in the field of image processing technology, as long as these network models can extract video features from the RGB image information, optical flow information, and node coordinate information respectively, and will not be elaborated here.

[0097] In an implementation manner of the above step S102, after determining that the person in the video information is a live body through live body recognition, a face image of the person can be extracted from the video information and face recognition can be performed on the face image. In this implementation manner, a face image that meets the preset image quality requirements can be extracted from the video information for face recognition, so as to accurately detect the face information in the face image, and then other face recognition operations such as face comparison can be performed according to the detected face information. In this implementation manner, those skilled in the art can flexibly set the specific conditions of the image quality requirements according to actual needs, such as the brightness being greater than a certain value, the clarity being greater than a certain value, etc., as long as the face information can be accurately detected from this face image.

[0098] Refer to the appendix Figure 3 , Figure 3It is a schematic flowchart of the main steps of an identity recognition method according to another embodiment of the present invention. In the embodiment of the present invention, the identity recognition method can be applied to a front-end device. As Figure 3 shown, the identity recognition method in the embodiment of the present invention mainly includes the following steps S201 - step S202.

[0099] Step S201: Call the video capture device of itself through an HTML page to capture video information and send the video information to the back-end device. The HTML page can at least include an HTML5 page.

[0100] Step S202: Receive the identity recognition result feedback by the back-end device according to the video information. The identity recognition result is the identity recognition result determined by the back-end device based on the results of live body recognition and face recognition after respectively performing live body recognition and face recognition on the video information.

[0101] The front-end device, back-end device, HTML page, video capture device, video information, live body recognition, face recognition, and identity recognition result in this embodiment are respectively the same as the relevant contents in the foregoing identity recognition method example, and will not be elaborated herein.

[0102] Based on the above steps S201 - step S202, directly calling the video capture device of the front-end device itself through an HTML page to capture the video information of personnel can be applied to all front-end devices that can load an HTML page. Even if the types of front-end devices are different, the video capture device can be called through the HTML page to capture video information. At the same time, after the front-end device captures the video information, it will send the video information to the back-end device. The back-end device completely completes the live body recognition and face recognition. The front-end device only needs to receive the identity recognition result determined by the back-end device after completing the live body recognition and face recognition, which greatly simplifies the method process of the identity recognition method and also overcomes the defect in the prior art that different recognition software needs to be developed for different types of front-end devices, so that the identity of personnel can be conveniently and reliably recognized.

[0103] The following further elaborates on the above steps S201 and step S202 respectively.

[0104] In an implementation manner of the above step S201, the front-end device can extract the action key frames in the video information and send the action key frames to the back-end device. In this implementation manner, the method for extracting the action key frames is the same as the relevant method in the foregoing identity recognition method embodiment, and will not be elaborated herein.

[0105] In an embodiment of the above step S202, the backend device can perform liveness recognition on the video information through the following steps: obtain different types of video features in the video information, classify actions according to all the video features to determine the type of action actually completed by the person in the video information; if the type of action actually completed is the same as the specified type sent by the front-end device, it is determined that the person in the video information is a live body; otherwise, it is determined that the person is not a live body. The above steps are the same as steps 11 to 12 in the foregoing method embodiment, and will not be elaborated here.

[0106] In an embodiment of the above step S202, after determining that the person in the video information is a live body through liveness recognition, the backend device can also extract the face image of the person from the video information and perform face recognition on the face image. The above steps are the same as the related methods in the foregoing method embodiment, and will not be elaborated here.

[0107] Refer to the appendix Figure 4 , Figure 4 is a schematic diagram of the main step flow of the identity recognition method according to still another embodiment of the present invention. As Figure 4 shown, the identity recognition method in the embodiment of the present invention mainly includes the following steps S301 - step S304.

[0108] Step S301: The front-end device obtains a video of a specified action.

[0109] Step S302: The front-end device obtains the key frames of the video.

[0110] Step S303: The backend device performs liveness recognition on the key frames.

[0111] Step S304: The backend device performs face recognition on the key frames.

[0112] After performing the above steps S301 to S304, it is also possible to control the backend device to determine the identity recognition result according to the results of liveness recognition and face recognition, and send the identity recognition result to the front-end device.

[0113] It should be noted that the methods described in the above steps S301 to S304 are respectively the same as the related methods in the foregoing Figures 1-3 described method embodiments, and will not be elaborated here.

[0114] Refer to the appendix Figure 5 , Figure 5 is a schematic diagram of the main step flow of the identity recognition method according to yet another embodiment of the present invention. As Figure 5 shown, the identity recognition method in the embodiment of the present invention mainly includes the following steps S401 - step S404.

[0115] Step S401: The front-end device acquires the video of the specified action.

[0116] Step S402: The front-end device obtains the key frames of the video based on the node weights.

[0117] Step S403: Perform action classification based on multiple types of video features of the video frames.

[0118] Step S404: Obtain the action classification result.

[0119] After performing the above steps S401 to S404, the back-end device can also be controlled to perform face recognition on the video information selectively according to the action classification result to determine the face recognition result.

[0120] It should be noted that the methods described in the above steps S401 to S404 are respectively the same as the relevant methods in the foregoing Figures 1-3 method embodiments, and will not be elaborated herein.

[0121] Refer to the appendix Figures 6 to 8 , Figures 6 to 8 which shows an application scenario of the present invention. As Figure 6 shown, in this application scenario, the identity recognition application can be called through the HTML page of the front-end device. The functions of this identity recognition application mainly include live detection, face comparison, face recognition, document OCR, attribute analysis, real-time analysis, display of logs, and parameter setting. Among them, the live detection is the same as the live recognition in the foregoing method embodiments; face recognition refers to detecting the face information of the face image in the video information, and face comparison refers to comparing the face information detected according to the face recognition with the personnel in the preset personnel library to determine which personnel in the preset personnel library the current personnel is. That is to say, the face comparison and face recognition in this embodiment jointly implement the face recognition in the foregoing method embodiments. Document OCR recognition refers to performing text recognition on the document image of the personnel by using the optical character recognition technology (OCR). Attribute analysis refers to performing attribute analysis on the personnel according to the results of live detection, face comparison, face recognition, and document OCR, such as analyzing the age and gender of the personnel. Real-time analysis refers to performing real-time analysis on the results of live detection, face comparison, face recognition, and document OCR in response to the received analysis requirements. Display of logs refers to displaying the application logs of the identity recognition application, and parameter setting refers to setting the application parameters of the identity recognition application, such as adjusting the size of the display interface, etc.

[0122] When the personnel selects live detection, it will enter Figure 7The interface shown, in which the matters needing attention during the live detection will be prompted to the personnel. When the personnel confirm that they can start the live detection, they can click "Start Detection", and then enter Figure 8 the interface shown. In Figure 8 the interface shown, the personnel will be prompted to complete actions of a specific type, such as Figure 8 turning the head to the left as shown. After clicking "Detect", the identity recognition application can call the video capture device of the front-end device through the HTML page to capture the video information of the personnel turning the head to the left, and send the video information to the back-end device for live detection.

[0123] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present invention, different steps do not necessarily have to be executed in such an order, and they can be executed simultaneously (in parallel) or in other orders, and these changes are all within the protection scope of the present invention.

[0124] Furthermore, the present invention also provides an identity recognition device.

[0125] Referring to the appendix Figure 9 , Figure 9 is the main structural block diagram of the identity recognition device according to an embodiment of the present invention. As Figure 9 shown, the identity recognition device in the embodiment of the present invention mainly includes a first information receiving module, an information recognition module, and a result processing module. The first information receiving module can be configured to receive the video information sent by the front-end device, and the video information is the video information captured by the front-end device through its own video capture device by calling the HTML page; the information recognition module can be configured to perform live body recognition and face recognition on the video information respectively; the result processing module can be configured to determine the identity recognition result according to the results of the live body recognition and face recognition and send the identity recognition result to the front-end device.

[0126] In one embodiment, the video information can be the key action frames extracted from the video information after the front-end device captures the video information through its own video capture device by calling the HTML page.

[0127] In one embodiment, the HTML page at least includes an HTML 5 page;

[0128] In one embodiment, the information recognition module may include a live body recognition sub-module and / or a face recognition sub-module. The live body recognition sub-module may be configured to perform the following operations: obtain different types of video features in the video information, classify actions based on all the video features to determine the type of action actually completed by the person in the video information; if the type of action actually completed is consistent with the specified type sent by the front-end device, determine that the person in the video information is a live body; otherwise, determine that the person is not a live body. The face recognition sub-module may be configured to perform the following operations: after determining that the person in the video information is a live body through live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0129] Furthermore, the present invention also provides another identity recognition device.

[0130] Refer to the appendix Figure 10 , Figure 10 is the main structural block diagram of the identity recognition device according to an embodiment of the present invention. As Figure 10 shown, the identity recognition device in the embodiment of the present invention mainly includes an information acquisition module and a second information receiving module. The information acquisition module may be configured to collect video information through its own video acquisition device by calling an HTML page and send the video information to the back-end device; the second information receiving module may be configured to receive the identity recognition result fed back by the back-end device according to the video information, and the identity recognition result is the identity recognition result determined by the back-end device after performing live body recognition and face recognition on the video information respectively according to the results of live body recognition and face recognition.

[0131] In one embodiment, the information acquisition module may be further configured to extract the action key frames in the video information and send the action key frames to the back-end device.

[0132] In one embodiment, the HTML page includes at least an HTML 5 page.

[0133] In one embodiment, the back-end device may be configured to perform live body recognition on the video information by executing the following steps: obtain different types of video features in the video information, classify actions based on all the video features to determine the type of action actually completed by the person in the video information; if the type of action actually completed is consistent with the specified type sent by the front-end device, determine that the person in the video information is a live body; otherwise, determine that the person is not a live body.

[0134] In one embodiment, the back-end device may also be configured to perform face recognition on the video information by executing the following steps: after determining that the person in the video information is a live body through live body recognition, extract the face image of the person from the video information and perform face recognition on the face image.

[0135] The aboveFigures 9 to 10 The identity recognition device shown is used to execute Figures 1 to 3 the embodiment of the identity recognition method shown. The technical principles, technical problems solved and technical effects produced by both are similar. Those skilled in the art of this technology can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the identity recognition device can refer to the content described in the embodiment of the identity recognition method, and will not be elaborated here.

[0136] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0137] Furthermore, the present invention also provides a control device. In an embodiment of the control device according to the present invention, the control device includes a processor and a storage device. The storage device can be configured to store a program for executing the identity recognition method of the above-mentioned method embodiment, and the processor can be configured to execute the program in the storage device. The program includes, but is not limited to, the program for executing the identity recognition method of the above-mentioned method embodiment. For the convenience of description, only the part related to the embodiment of the present invention is shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The control device can be a control device formed by various electronic devices.

[0138] Furthermore, the present invention also provides a computer-readable storage medium. In an embodiment of the computer-readable storage medium according to the present invention, the computer-readable storage medium may be configured to store a program for executing the identity recognition method in the above method embodiment. This program can be loaded and run by a processor to implement the above identity recognition method. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The computer-readable storage medium may be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present invention is a non-transitory computer-readable storage medium.

[0139] Furthermore, it should be understood that since the setting of each module is only for illustrating the functional units of the device of the present invention, the corresponding physical devices of these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only illustrative.

[0140] Those skilled in the art can understand that the various modules in the device can be adaptively split or combined. Such splitting or combination of specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combination will all fall within the protection scope of the present invention.

[0141] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. An identity recognition method, characterized in that, Applied to a backend device, the method includes: Receiving an action key frame sent by a frontend device, where the action key frame is extracted from video information by the frontend device after collecting video information through its own video capture device via an HTML page; Performing liveness recognition and face recognition on the action key frame respectively; Determining an identity recognition result based on the results of liveness recognition and face recognition and sending the identity recognition result to the frontend device; Wherein, The action key frame is obtained by the following method: Calculate the inter-frame distance dis(x (m) and x (n) ) between any two video frames x (m) ,x (n) ) in the video information respectively, and select the two video frames with the largest inter-frame distance as the key action frames; and respectively represent the abscissa of the feature point corresponding to the i-th part node in video frames x (m) and x (n) , and the ordinate of the feature point corresponding to the i-th part node in video frames x and respectively represent the ordinate of the feature point corresponding to the i-th part node in video frames x (m) and x (n) . ω i represents the contribution degree of the i-th part node, N represents the total number of part nodes, and the part nodes include face nodes and non-face nodes; represents the coordinate variance of the i-th part node, represents the sum of the coordinate variances of all part nodes, k represents the total number of video frames, and respectively represent the abscissa and ordinate of the feature point corresponding to the i-th part node in the f-th video frame, and respectively represent the average abscissa and average ordinate of the feature points corresponding to the i-th part node in k video frames.

2. The identity recognition method according to claim 1, wherein The step of "performing liveness recognition on the action key frame" specifically includes: Obtaining different types of video features in the action key frame, classifying actions based on all video features to determine the type of action actually completed by the person in the action key frame; If the type of action actually completed is consistent with the specified type sent by the frontend device, it is determined that the person in the action key frame is a live person; otherwise, it is determined that the person is not a live person; And / or, The step of "performing face recognition on the action key frame" specifically includes: After determining that the person in the action key frame is a live person through the liveness recognition, extracting the face image of the person from the action key frame and performing face recognition on the face image.

3. The identity recognition method according to claim 1, wherein The HTML page includes at least an HTML5 page.

4. An identity recognition method, characterized in that, Applied to a frontend device, the method includes: Collecting video information by calling its own video capture device through an HTML page, extracting an action key frame from the video information and sending the action key frame to the backend device; Receiving the identity recognition result fed back by the backend device according to the action key frame, where the identity recognition result is the identity recognition result determined by the backend device based on the results of liveness recognition and face recognition performed on the action key frame respectively; Wherein, The extracting of the action key frame from the video information includes: Calculate the inter-frame distance dis(x (m) and x (n) ) between any two video frames x (m) , x (n) ) in the video information respectively, and select the two video frames with the largest inter-frame distance as the action key frames; and respectively represent the abscissa of the feature point corresponding to the i-th part node in video frames x (m) and x (n) . and respectively represent the ordinate of the feature point corresponding to the i-th part node in video frames x (m) and x (n) . ω i represents the contribution degree of the i-th part node, N represents the total number of part nodes, and the part nodes include face nodes and non-face nodes; represents the coordinate variance of the i-th part node, represents the sum of the coordinate variances of all part nodes, k represents the total number of video frames, and respectively represent the abscissa and ordinate of the feature point corresponding to the i-th part node in the f-th video frame, and respectively represent the average abscissa and average ordinate of the feature points corresponding to the i-th part node in k video frames.

5. The identity recognition method according to claim 4, wherein The HTML page includes at least an HTML5 page.

6. The identity recognition method according to claim 4, wherein The backend device is configured to perform liveness recognition on the action key frame by executing the following steps: Obtaining different types of video features in the action key frame, classifying actions based on all video features to determine the type of action actually completed by the person in the action key frame; If the type of action actually completed is consistent with the specified type sent by the frontend device, it is determined that the person in the action key frame is a live person; otherwise, it is determined that the person is not a live person; And / or, The backend device is further configured to perform face recognition on the action key frame by executing the following steps: After determining that the person in the action key frame is a live person through the liveness recognition, extracting the face image of the person from the action key frame and performing face recognition on the face image.

7. An identity recognition device, characterized in that, Applied to a backend device, the device includes: A first information receiving module, configured to receive an action key frame sent by a front-end device, where the action key frame is extracted from video information by the front-end device after the front-end device calls its own video acquisition device to acquire video information through an HTML page; An information recognition module, configured to perform liveness recognition and face recognition on the action key frame respectively; A result processing module, configured to determine an identity recognition result according to the results of liveness recognition and face recognition and send the identity recognition result to the front-end device; Wherein, The action key frame is obtained by the following method: Calculate the inter-frame distance dis(x (m) and x (n) ) between any two video frames x (m) ,x (n) in the video information respectively, and select the two video frames with the largest inter-frame distance as the action key frames; and respectively represent the abscissa of the feature point corresponding to the i-th part node in video frames x (m) and x (n) . The ordinate of the feature point corresponding to the i-th part node in video frames x and respectively represent the ordinate of the feature point corresponding to the i-th part node in video frames x (m) and x (n) . ω i represents the contribution degree of the i-th part node, N represents the total number of part nodes, and the part nodes include face nodes and non-face nodes; represents the coordinate variance of the $i$-th part node, represents the sum of the coordinate variances of all part nodes, $k$ represents the total number of video frames, and respectively represent the abscissa and ordinate of the feature point corresponding to the $i$-th part node in the $f$-th video frame, and respectively represent the average abscissa and average ordinate of the feature points corresponding to the $i$-th part node in $k$ video frames.

8. The identity recognition device according to claim 7, wherein The HTML page includes at least an HTML 5 page; And / or, the information recognition module includes a liveness recognition sub-module and / or a face recognition sub-module; The liveness recognition sub-module is configured to perform the following operations: Obtain different types of video features in the action key frame, classify actions according to all video features to determine the action type actually completed by the person in the action key frame; If the actually completed action type is consistent with the specified type sent by the front-end device, it is determined that the person in the action key frame is a live body; otherwise, it is determined that the person is not a live body; The face recognition sub-module is configured to perform the following operations: After determining that the person in the action key frame is a live body through the liveness recognition, extract the face image of the person from the action key frame and perform face recognition on the face image.

9. An identity recognition device, characterized in that, Applied to a front-end device, the device includes: An information acquisition module, configured to call its own video acquisition device to acquire video information through an HTML page, extract an action key frame from the video information and send the action key frame to a back-end device; A second information receiving module, configured to receive the identity recognition result fed back by the back-end device according to the action key frame, where the identity recognition result is the identity recognition result determined by the back-end device according to the results of liveness recognition and face recognition after respectively performing liveness recognition and face recognition on the action key frame; Wherein, The extracting the action key frame from the video information includes: Calculate the inter-frame distance dis(x (m) and x (n) ) between any two video frames x (m) , x (n) ) in the video information respectively, and select the two video frames with the largest inter-frame distance as the action key frames; and respectively represent the abscissa of the feature point corresponding to the $i$-th part node in video frames $x$ (n) and $x$ (n) . The ordinate of the feature point corresponding to the $i$-th part node in video frames $x$ and respectively represent the ordinate of the feature point corresponding to the $i$-th part node in video frames $x$ (m) and $x$ (n) . $\omega$ i represents the contribution degree of the $i$-th part node, $N$ represents the total number of part nodes, and the part nodes include face nodes and non-face nodes; represents the coordinate variance of the i-th part node, represents the sum of the coordinate variances of all part nodes, k represents the total number of video frames, and respectively represent the abscissa and ordinate of the feature point corresponding to the i-th part node in the f-th video frame, and respectively represent the average abscissa and average ordinate of the feature points corresponding to the i-th part node in k video frames.

10. The identity recognition device according to claim 9, characterized in that, The HTML page includes at least an HTML 5 page; And / or, the back-end device is configured to perform liveness recognition on the action key frame by executing the following steps: Obtain different types of video features in the action key frame, classify actions according to all video features to determine the action type actually completed by the person in the action key frame; If the actually completed action type is consistent with the specified type sent by the front-end device, it is determined that the person in the action key frame is a live body; otherwise, it is determined that the person is not a live body; And / or, the back-end device is further configured to perform face recognition on the action key frame by executing the following steps: After determining that the person in the action key frame is a live body through the liveness recognition, extract the face image of the person from the action key frame and perform face recognition on the face image.

11. A control device, comprising a processor and a storage device, the storage device being adapted to store a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by the processor to execute the identity recognition method according to any one of claims 1 to 6.

12. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is adapted to be loaded and run by a processor to perform the identity recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Human face recognition method, server and computer readable storage medium

    CN108269333A

  • Face recognition system based on video image heart rate detection and living body detection

    CN112396011A