Identity authentication method and device, storage medium and electronic equipment

By automatically capturing customers' facial images during video calls for identity verification, the problem of low efficiency and high error rate caused by manual operation in existing technologies is solved, and an efficient and accurate identity verification process is achieved.

CN116405227BActive Publication Date: 2026-03-20INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, identity verification by manually capturing customer facial images suffers from low efficiency and a high error rate.

Method used

By establishing a video connection between the customer and the customer service representative, the system automatically collects video streams within a preset time range, uses facial recognition technology to obtain the target image to be authenticated from the video stream, and performs identity verification, reducing manual operations.

Benefits of technology

It improves the efficiency and accuracy of identity authentication, reduces manpower and time costs, and enhances the customer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405227B_ABST
    Figure CN116405227B_ABST
Patent Text Reader

Abstract

The application discloses an identity authentication method and device, a storage medium and electronic equipment, and relates to the field of financial technology. The method comprises the following steps: in response to a video call request, a video connection between a first object and a second object is established, wherein the video call request is a request for the first object to initiate a video call with the second object through a first device; according to the video connection, a target video stream is obtained, and according to the target video stream, a target to-be-authenticated image is determined, wherein the target video stream is a video stream within a preset time range; according to the target to-be-authenticated image, identity authentication is performed on the first object, and an authentication result is obtained, wherein the authentication result is used to provide data reference for the second object. The application solves the technical problem that the existing technology manually intercepts the face image of a customer, resulting in low auditing efficiency when the customer is authenticated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of financial technology, in particular to an identity authentication method and device, a storage medium and an electronic device. BACKGROUND

[0002] With the promotion of non-contact financial service mode, financial institutions (for example, banks) transfer some businesses (for example, identity authentication, etc.) originally required to be handled by customers offline through manual operation to online through audio and video review. At present, when customers handle identity authentication review business through remote video call, the related technology mainly guides customers to take pictures through the oral instruction of the seat customer service to obtain the identification pictures required for identity authentication. When the taken pictures do not meet the identification standard, the seat customer service needs to manually re-capture the pictures. In addition, when performing identity authentication, the seat customer service needs to manually submit the comparison. The seat customer service needs to perform multiple manual operations in the business handling process, the review efficiency is low, and the error rate is high, and the interactive experience of the customers is not good.

[0003] At present, no effective solution has been proposed for the above problems. SUMMARY

[0004] The embodiments of the present application provide an identity authentication method, device, storage medium and electronic device to at least solve the technical problem that the existing technology causes low efficiency in identity authentication of customers by manually intercepting the face images of the customers.

[0005] According to an aspect of the embodiments of the present application, an identity authentication method is provided, including: establishing a video connection between a first object and a second object in response to a video call request, wherein the video call request is a request of the first object to initiate a video call with the second object through a first device; obtaining a target video stream according to the video connection, and determining a target to-be-authenticated image according to the target video stream, wherein the target video stream is a video stream within a preset time range; performing identity authentication on the first object according to the target to-be-authenticated image, and obtaining an authentication result, wherein the authentication result is used to provide data reference for the second object.

[0006] Further, obtaining the target video stream according to the video connection includes: collecting a first image of the first object according to the video connection, and displaying the first image to the second object; and obtaining the target video stream in the case that the first image is displayed abnormally.

[0007] Further, the target to-be-authenticated image is determined according to the target video stream, including: collecting an image from the target video stream to obtain a second image of the first object; performing image preprocessing on the second image to obtain a processed image; determining a face position of the first object according to the processed image to generate face coordinate information of the first object, wherein the face coordinate information is information for representing a position of a feature point of a face of the first object; performing face alignment processing on the processed image according to the face coordinate information to obtain an image conforming to a preset shape; performing face encoding processing on the image conforming to the preset shape to obtain a first feature vector of the first object; obtaining a reference feature vector of the first object, and determining the target to-be-authenticated image according to the first feature vector and the reference feature vector, wherein the reference feature vector is an archived feature vector of the first object.

[0008] Further, the target to-be-authenticated image is determined according to the first feature vector and the reference feature vector, including: performing similarity calculation on the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; comparing the first similarity with a first threshold to obtain a first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, taking the second image as the target to-be-authenticated image; if the first comparison result is that the first similarity is less than or equal to the first threshold, re-executing the step of collecting an image from the target video stream until the first similarity is greater than the first threshold, and taking a current image as the target to-be-authenticated image.

[0009] Further, the target to-be-authenticated image is determined according to the first feature vector and the reference feature vector, including: performing similarity calculation on the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; comparing the first similarity with a first threshold to obtain a first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, taking the second image as the target to-be-authenticated image; if the first comparison result is that the first similarity is less than or equal to the first threshold, re-executing the step of collecting an image from the target video stream until the first similarity is greater than the first threshold, and taking a current image as the target to-be-authenticated image.

[0010] Further, the target to-be-authenticated image is determined according to the first feature vector and the reference feature vector, including: performing similarity calculation on the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; comparing the first similarity with a first threshold to obtain a first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, taking the second image as the target to-be-authenticated image; if the first comparison result is that the first similarity is less than or equal to the first threshold, re-executing the step of collecting an image from the target video stream until the first similarity is greater than the first threshold, and taking a current image as the target to-be-authenticated image.

[0011] Furthermore, before responding to a video call request and establishing a video connection between the first object and the second object, the method further includes: receiving a video call request; determining whether the network environment of the first device is in the target network environment; if the network environment of the first device is in the target network environment, then entering the waiting response page corresponding to the video call request; if the network environment of the first device is not in the target network environment, then generating target prompt information, wherein the target prompt information is used to prompt the first object to switch network environments.

[0012] According to another aspect of the present invention, an identity authentication device is also provided, comprising: a first response module, configured to respond to a video call request and establish a video connection between a first object and a second object, wherein the video call request is a request initiated by the first object through a first device to conduct a video call with the second object; a first determination module, configured to acquire a target video stream based on the video connection and determine a target image to be authenticated based on the target video stream, wherein the target video stream is a video stream within a preset duration range; and a first processing module, configured to perform identity authentication on the first object based on the target image to be authenticated and obtain an authentication result, wherein the authentication result is used to provide data reference for the second object.

[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described authentication method at runtime.

[0014] According to another aspect of the present invention, an electronic device is also provided, the electronic device including one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are configured to run the programs, wherein the programs are configured to execute the above-described authentication method during runtime.

[0015] In this embodiment of the invention, a method is adopted to automatically capture the customer's facial image from a video stream. First, a video call request is responded to, establishing a video connection between a first object and a second object. Then, based on the video connection, a target video stream is acquired, and based on the target video stream, a target image to be authenticated is determined. Finally, based on the target image to be authenticated, the first object is authenticated to obtain an authentication result. The video call request is initiated by the first object through a first device to conduct a video call with the second object. The target video stream is a video stream within a preset duration range, and the authentication result is used to provide data reference for the second object.

[0016] In the above process, by establishing a video connection between the customer and the customer service representative, a data foundation is provided for obtaining the target video stream. Based on the video connection, the target video stream can be obtained, and the customer's facial image can be automatically captured from the video stream to obtain the target image to be authenticated. Then, based on the target image to be authenticated, the customer's identity can be verified and the authentication result can be obtained. This achieves the acquisition of facial images without the customer's awareness, greatly reducing labor and time costs, improving the efficiency and accuracy of audio and video review and authentication, and enhancing the customer experience of the audio and video review process.

[0017] Therefore, the technical solution of this invention achieves the goal of reducing redundant operations of customer service representatives during business processing by automatically collecting customer facial images, reducing the error rate of manual operations, and improving customer experience. This achieves the technical effect of improving the efficiency of review and authentication, and solves the technical problem of low review efficiency when manually capturing customer facial images in the prior art. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0019] Figure 1 This is a flowchart of an optional identity authentication method according to an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of an optional identity authentication and verification process according to an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of an optional automatic image acquisition process according to an embodiment of the present invention;

[0022] Figure 4 This is a schematic diagram of an optional identity authentication device according to an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of an optional electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] It should be noted that all relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this invention are information and data authorized by the user or fully authorized by all parties. For example, this system has an interface with the relevant user or organization. Before obtaining relevant information, it needs to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving consent from the aforementioned user or organization.

[0027] Example 1

[0028] According to an embodiment of the present invention, an embodiment of an identity authentication method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0029] Figure 1 This is a flowchart of an optional identity authentication method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0030] Step S101: Respond to the video call request and establish a video connection between the first object and the second object, wherein the video call request is a request initiated by the first object through the first device to conduct a video call with the second object.

[0031] In the above steps, video call requests can be responded to by application systems, processors, electronic devices, etc. Optionally, video call requests can be responded to by an identity authentication system. The first object can be a customer who conducts business remotely through mobile banking, and the second object can be a video customer service representative. The first device can be the terminal device used by the customer, such as a mobile phone, tablet computer, laptop computer, PC, etc.

[0032] Specifically, when a customer enters the mobile banking video verification prompt page and initiates a video call request, the video agent confirms the connection to the video service through a terminal device (e.g., a PC), responds to the video call request through the identity authentication system associated with the terminal device used by the video agent, displays the video agent's service page on the customer's terminal device, and establishes a video connection between the customer and the video agent.

[0033] Step S102: Based on the video connection, obtain the target video stream, and based on the target video stream, determine the target image to be authenticated, wherein the target video stream is a video stream within a preset duration range.

[0034] Specifically, the target video stream can be a video stream within a preset duration range, such as a video stream within one minute after the video connection is established. The target image to be authenticated can be a facial image of a customer used for identity verification and comparison. Based on the target video stream, the target image to be authenticated can be determined; for example, by capturing images from the target video stream, the target image to be authenticated can be determined from the captured images.

[0035] Step S103: Based on the target image to be authenticated, the first object is authenticated to obtain the authentication result, wherein the authentication result is used to provide data reference for the second object.

[0036] Specifically, the system can perform identity recognition comparison between the target image to be authenticated and the customer's archived reference image (i.e., a face image of the customer in an upright posture) to achieve customer identity authentication. Optionally, the identity authentication system uses a pre-trained face recognition model to perform identity recognition comparison between the target image to be authenticated and the customer's archived reference image to obtain the comparison result, thereby obtaining the authentication result based on the comparison result.

[0037] Specifically, after the first object is authenticated based on the target image to be authenticated and the authentication result is obtained, the authentication result is displayed in the service interface of the second device (i.e. the terminal device used by the video customer service representative). The video customer service representative then determines the final review result based on the authentication result.

[0038] Based on the scheme defined in steps S101 to S103 above, it can be understood that in this embodiment of the invention, the method of automatically acquiring the customer's facial image from the video stream is adopted. First, a video connection is established between the first object and the second object in response to the video call request. Then, based on the video connection, the target video stream is obtained, and based on the target video stream, the target image to be authenticated is determined. Finally, based on the target image to be authenticated, the first object is authenticated to obtain the authentication result. Here, the video call request is a request initiated by the first object through the first device to conduct a video call with the second object, the target video stream is a video stream within a preset duration range, and the authentication result is used to provide data reference for the second object.

[0039] It is noteworthy that, in the above process, establishing a video connection between the customer and the customer service representative provides the data foundation for obtaining the target video stream. Based on the video connection, the target video stream can be obtained, and the customer's facial image can be automatically captured from the video stream to obtain the target image to be authenticated. Then, based on the target image to be authenticated, the customer's identity can be verified, and the authentication result can be obtained. This achieves the acquisition of facial images without the customer's awareness, greatly reducing labor and time costs, improving the efficiency and accuracy of audio and video review and authentication, and enhancing the customer experience of the audio and video review process.

[0040] Therefore, the technical solution of this invention achieves the goal of reducing redundant operations of customer service representatives during business processing by automatically collecting customer facial images, reducing the error rate of manual operations, and improving customer experience. This achieves the technical effect of improving the efficiency of review and authentication, and solves the technical problem of low review efficiency when manually capturing customer facial images in the prior art.

[0041] In one optional embodiment, acquiring a target video stream based on a video connection includes: acquiring a first image of a first object based on the video connection and displaying the first image to a second object; and acquiring the target video stream if the first image is displayed abnormally.

[0042] Specifically, in the process of obtaining the target video stream based on the video connection, the first image of the customer is captured based on the video connection. That is, after the video call is connected, the identity authentication system automatically captures a facial image of the customer (i.e., the first image mentioned above) and displays the image to the video customer service representative through the service interface of the terminal device used by the video customer service representative.

[0043] Specifically, if the first image captured can be displayed normally in the service interface, the video customer service representative clicks the "Submit for Comparison" option, and the identity authentication system uses a pre-trained face recognition model to compare the first image with the customer's archived baseline image to obtain the comparison result; if the first image captured cannot be displayed normally in the service interface, that is, the first image is displayed abnormally, the identity authentication system obtains the target video stream, for example, obtains the video stream within 1 minute after the video connection is established.

[0044] It should be noted that, based on the video connection, the system automatically captures the customer's facial image, enabling the acquisition of portrait images without the customer's awareness. This significantly reduces labor and time costs, improves video review and authentication efficiency, and enhances the customer experience.

[0045] In one optional embodiment, determining the target image to be authenticated based on the target video stream includes: acquiring an image from the target video stream to obtain a second image of a first object; performing image preprocessing on the second image to obtain a processed image; determining the face position of the first object based on the processed image and generating face coordinate information of the first object, wherein the face coordinate information is information used to characterize the position of face feature points of the first object; performing face alignment processing on the processed image based on the face coordinate information to obtain an image conforming to a preset shape; performing face encoding processing on the image conforming to the preset shape to obtain a first feature vector of the first object; obtaining a reference feature vector of the first object, and determining the target image to be authenticated based on the first feature vector and the reference feature vector, wherein the reference feature vector is an archived feature vector of the first object.

[0046] Specifically, in the process of determining the target image to be authenticated based on the target video stream, the image is first captured from the target video stream to obtain the second image of the first object. For example, the image is captured from the video stream within 1 minute after the video connection is established to obtain the customer's face image (i.e., the second image).

[0047] Further, the second image is preprocessed to obtain a processed image. Optionally, the identity authentication system uses a pre-trained face recognition model to preprocess the second image to obtain a processed image. Specifically, the original image (i.e., the second image) is input into the pre-trained face recognition model, and the pre-trained face recognition model converts the second image to grayscale to obtain the processed image.

[0048] Furthermore, based on the processed image, the position of the first object's face can be determined, generating the face coordinate information of the first object. Specifically, the face coordinate information is information used to characterize the position of the customer's facial feature points, such as the position information of the facial features. The pre-trained face recognition model can calculate the gradient of pixels in the processed image to obtain a gradient histogram, and use the gradient histogram to detect the position of the face in the image, cut out the face and mark the coordinate position to generate the customer's face coordinate information.

[0049] Furthermore, based on the facial coordinate information, face alignment processing is performed on the processed image to obtain an image that conforms to a preset shape. Specifically, the pre-trained face recognition model can locate the feature points on the face based on the facial coordinate information, and align each feature point to a standard shape through geometric transformations (e.g., rotation, scaling, etc.), that is, align the customer's facial feature points to a preset standard position (e.g., the customer's face position in an upright posture), to obtain an image that conforms to the preset shape.

[0050] Furthermore, by performing face encoding processing on an image that conforms to a preset shape, the first feature vector of the first object can be obtained. Specifically, a pre-trained face recognition model can re-mark the coordinate positions of an image that conforms to a preset shape to obtain face coordinate information. Then, through feature extraction processing, the pixel values ​​of the face image are converted into feature vectors to obtain the customer's first feature vector.

[0051] Furthermore, the baseline feature vector of the first object is obtained, and the target image to be authenticated is determined based on the first feature vector and the baseline feature vector. Specifically, the baseline feature vector is the customer's archived feature vector, that is, the feature vector corresponding to the customer's archived baseline image. The pre-trained face recognition model can determine the target image to be authenticated based on the first feature vector and the baseline feature vector.

[0052] It should be noted that, based on facial recognition technology, the system automatically collects customers' facial images. During the video review process, it acquires images of people without the customer's knowledge, providing a better review experience, reducing the manual workload of video customer service staff, improving the efficiency of video review and authentication, and increasing the work efficiency of video customer service staff.

[0053] In one optional embodiment, determining the target image to be authenticated based on the first feature vector and the reference feature vector includes: calculating the similarity between the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; comparing the first similarity with a first threshold to obtain a first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, then the second image is used as the target image to be authenticated; if the first comparison result is that the first similarity is less than or equal to the first threshold, then the step of acquiring the image from the target video stream is repeated until the first similarity is greater than the first threshold, and the current image is used as the target image to be authenticated.

[0054] Optionally, in the process of determining the target image to be authenticated based on the first feature vector and the reference feature vector, a similarity calculation is performed on the first feature vector and the reference feature vector to obtain the first similarity between the first feature vector and the reference feature vector. Specifically, the first similarity between the first feature vector and the reference feature vector is obtained by performing a similarity calculation on the first feature vector and the reference feature vector through a pre-trained face recognition model.

[0055] Further, the first similarity is compared with the first threshold to obtain the first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, then the second image is used as the target image to be authenticated; if the first comparison result is that the first similarity is less than or equal to the first threshold, then the step of acquiring images from the target video stream is repeated until the first similarity is greater than the first threshold, and the current image is used as the target image to be authenticated.

[0056] Specifically, based on business needs, a first threshold is preset, for example, 85%. The first similarity is compared with the first threshold. If the calculated first similarity is greater than 85%, the second image is considered to have a similarity greater than 85% with the customer's archived reference image, and the second image is selected for identity recognition comparison, that is, the second image is used as the target image to be authenticated. If the calculated first similarity is less than or equal to 85%, the second image is considered to have a similarity less than or equal to 85% with the customer's archived reference image, and the image is re-acquired from the target video stream, that is, an image with a similarity greater than 85% with the customer's archived reference image is acquired, and the current image is used as the target image to be authenticated.

[0057] It should be noted that by calculating the similarity between the first feature vector and the baseline feature vector, the target image to be authenticated can be determined, providing a data foundation for subsequent identity authentication and recognition. This reduces redundant operations for video customer service representatives during business processing, lowers the error rate of manual operations, and improves the accuracy of video review and authentication.

[0058] In one optional embodiment, the first object is authenticated based on the target image to be authenticated to obtain an authentication result, including: comparing the target image to be authenticated with the reference image corresponding to the reference feature vector to obtain a second comparison result; if the second comparison result is a successful comparison, the authentication result is determined to be successful; if the second comparison result is a failed comparison, the number of failures is obtained, and the authentication result is determined based on the number of failures.

[0059] Specifically, during the process of authenticating the first object based on the target image to be authenticated and obtaining the authentication result, the target image to be authenticated is compared with the reference image corresponding to the reference feature vector to obtain a second comparison result. Optionally, after determining the target image to be authenticated, the target image to be authenticated is displayed on the service interface of the terminal device used by the video customer service representative. The video customer service representative clicks the "Submit Comparison" option, and the identity authentication system uses a pre-trained face recognition model to compare the target image to be authenticated with the reference image to achieve customer identity authentication.

[0060] Specifically, the second comparison result can characterize the similarity between the face in the target image to be authenticated and the face in the reference image. According to business needs, a recognition threshold for identity authentication is preset, for example, the recognition threshold is 95%. If the similarity between the target image to be authenticated and the reference image is greater than 95%, the second comparison result is a successful comparison. Further, the authentication result is determined to be authentication passed. Optionally, the video customer service representative informs the customer that the identity authentication review has been passed through a video call. If the similarity between the target image to be authenticated and the reference image is less than or equal to 95%, the second comparison result is a failed comparison. Further, the number of failed comparisons is obtained, and the authentication result is determined based on the number of failed comparisons.

[0061] It should be noted that by comparing the target image to be authenticated with the benchmark image corresponding to the benchmark feature vector, customer identity authentication is achieved, reducing the error rate of manual operation and improving the accuracy of video review and authentication.

[0062] In one optional embodiment, determining the authentication result based on the number of failures includes: if the number of failures is less than or equal to a second threshold, then re-execute the step of acquiring images from the target video stream, and determine the authentication result based on the re-acquired image; if the number of failures is greater than the second threshold, then based on the video connection, the second object sends a shooting instruction to the first object, and a third image of the first object is captured based on the video connection, and the authentication result is determined based on the third image, wherein the shooting instruction is used to instruct the first object to adjust its face position.

[0063] Specifically, based on business needs, a second threshold is preset, for example, two attempts. If the number of failures is less than or equal to two, i.e., less than or equal to two unsuccessful comparisons, automatic image acquisition is performed again from the target video stream, and the authentication result is determined based on the re-acquired image. If the number of failures is greater than two, i.e., three unsuccessful comparisons, the process enters the manual screenshot stage by the video customer service representative. The video customer service representative verbally guides the customer to adjust their position via video call (i.e., the second party issues a shooting instruction to the first party). The video customer service representative clicks the "screenshot" option in the service interface, captures a third image of the customer based on the video connection, and determines the authentication result based on the third image.

[0064] It should be noted that in this embodiment, automatic image acquisition is used first to verify the customer's identity, supplemented by manual image acquisition, to ensure the smooth progress of the identity verification process and provide customers with a better experience in handling verification procedures.

[0065] In one optional embodiment, before responding to a video call request and establishing a video connection between the first object and the second object, a video call request is received; it is determined whether the network environment of the first device is in the target network environment; if the network environment of the first device is in the target network environment, the user enters the waiting response page corresponding to the video call request; if the network environment of the first device is not in the target network environment, a target prompt message is generated, wherein the target prompt message is used to prompt the first object to switch network environments.

[0066] Specifically, the target network environment can be a Wi-Fi environment. When a customer enters the mobile banking video verification prompt page and initiates a video call request, the system receives the video call request and determines whether the customer's terminal device is in a Wi-Fi environment. If it is in a Wi-Fi environment, the customer is taken to the page for calling a remote video customer service representative, i.e., the waiting page for the video call request. If the customer is not in a Wi-Fi environment, a target prompt message is generated, such as a pop-up prompting the customer to switch network environments. If the customer clicks "continue," they are taken to the page for calling a remote video customer service representative; if they click "cancel," they are returned to the mobile banking video verification prompt page.

[0067] It should be noted that by determining whether the network environment of the first device is in the target network environment, the customer can be prompted to switch network environments, providing a better experience for customers in handling audit business and improving customer satisfaction.

[0068] Figure 2 This is a schematic diagram of an optional identity authentication and verification process according to an embodiment of the present invention, such as... Figure 2 As shown, the main steps include the following:

[0069] Step 1: The customer enters the mobile banking video verification prompt page and initiates an audio / video call.

[0070] Step 2: Determine if the customer is operating in a Wi-Fi environment. If they are, proceed to the remote audio / video customer service page. If they are not in a Wi-Fi environment, a pop-up message will appear. If the customer clicks "Continue," they will be taken to the remote audio / video customer service page. If they click "Cancel," they will be returned to the mobile banking video verification prompt page.

[0071] Step 3: The audio / video customer service representative confirms the connection to the audio / video service. The customer then sees the audio / video customer service page and proceeds to Step 4.

[0072] Step 4: After the video connection is established, first capture an image and display it on the customer service side. If the facial recognition photo displayed on the customer service side can be displayed normally, proceed to Step 6 to enter the submission and comparison stage.

[0073] Step 5: If the customer's facial recognition photo displayed on the customer service side is abnormal, the system will collect the first image with a similarity greater than 85% to the customer's facial recognition baseline photo from the video stream within 1 minute after the video connection is established and fill it in. If the collection is successful, a pop-up message will appear saying "Automatic collection completed, please submit for comparison", and the system will proceed to Step 6; if the collection fails, a pop-up message will appear saying "Automatic collection failed, please take a screenshot manually", and the system will proceed to Step 7.

[0074] Step 6: Customer service clicks "Submit for Comparison." The system compares the collected image with the customer's archived image. If the comparison is successful, the customer service will be notified that the comparison result is "Passed," and the process will proceed to Step 8 for the final review. If the first and second system comparisons fail, the customer service will be notified that the comparison result is "Failed," and the process will proceed to Step 7 for the customer service to manually take a screenshot. If the third system comparison fails, the customer will be informed that the process cannot be completed, and the process will proceed to Step 8 for the final review.

[0075] Step 7: The customer service representative clicks the "Prepare Screenshot" button on the page. The customer then raises the face outline frame. The customer service representative verbally guides the customer to adjust the position and clicks the "Screenshot" button. The system automatically uploads the screenshot to the agent's page and proceeds to Step 6 to enter the submission and comparison stage.

[0076] Step 8: The customer service representative selects the final review result based on the comparison results and clicks submit. The customer or the customer service representative hangs up the call, and the review result page is displayed on the customer's side, ending the process.

[0077] Figure 3 This is a schematic diagram of an optional automatic image acquisition process according to an embodiment of the present invention, such as... Figure 3As shown, when the customer's facial recognition photo is displayed abnormally on the customer service side, image acquisition is performed. That is, after the customer establishes a video connection with the audio and video customer service, the customer's image is captured through the video stream. Then, the image is preprocessed using a pre-trained facial recognition model. The original image is input, converted to grayscale, and the gradient of the pixels in the image is calculated. The gradient histogram is used to detect the face position, the face is cut and the coordinate position is marked, and then face alignment is performed to locate the feature points on the face. Geometric transformations (such as rotation, scaling, etc.) are used to align each feature point to a standard shape. Then, face encoding and feature extraction are performed. The pixel values ​​of the converted standard-shape face image are marked with coordinates, and the pixel values ​​of the face image are converted into a discernible feature vector. Then, face matching is performed. The feature vector of the customer's archived self-image is compared with the feature vector extracted from the original image to obtain a similarity score.

[0078] It should be noted that, in this embodiment, based on facial recognition technology, facial images are acquired first during the video verification process without the customer's awareness, improving the efficiency of remote bank audio and video verification and enhancing the customer experience of the audio and video verification process. Through the facial recognition-based audio and video verification method, customer service representatives can automatically acquire the required facial images during the initial greeting phase of the video connection, ultimately completing the audio and video verification operation smoothly. This provides a good customer experience, reduces the manual workload of audio and video customer service personnel, and improves their work efficiency. By acquiring the customer's biometric information through facial recognition and automatically comparing it with online verification photos to fill in a matching facial recognition image, the manual operation content of customer service personnel is reduced, improving the efficiency of audio and video verification and reducing the error rate of manual operations.

[0079] Therefore, the technical solution of this invention achieves the goal of reducing redundant operations of customer service representatives during business processing by automatically collecting customer facial images, reducing the error rate of manual operations, and improving customer experience. This achieves the technical effect of improving the efficiency of review and authentication, and solves the technical problem of low review efficiency when manually capturing customer facial images in the prior art.

[0080] Example 2

[0081] According to an embodiment of the present invention, an embodiment of an identity authentication device is provided, wherein, Figure 4 This is a schematic diagram of an optional identity authentication device according to an embodiment of the present invention, such as... Figure 4As shown, the device includes: a first response module 401, used to respond to a video call request and establish a video connection between a first object and a second object, wherein the video call request is a request initiated by the first object through a first device to conduct a video call with the second object; a first determination module 402, used to obtain a target video stream based on the video connection and determine a target image to be authenticated based on the target video stream, wherein the target video stream is a video stream within a preset duration range; and a first processing module 403, used to perform identity authentication on the first object based on the target image to be authenticated and obtain an authentication result, wherein the authentication result is used to provide data reference for the second object.

[0082] It should be noted that the first response module 401, the first determination module 402 and the first processing module 403 mentioned above correspond to steps S101 to S103 in the above embodiments. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1.

[0083] Optionally, the first determining module includes: a first acquisition module, configured to acquire a first image of a first object based on a video connection, and display the first image to a second object; and a first acquisition module, configured to acquire a target video stream if the first image is displayed abnormally.

[0084] Optionally, the first determining module further includes: a second acquisition module, used to acquire images from the target video stream to obtain a second image of the first object; a second processing module, used to preprocess the second image to obtain a processed image; a second determining module, used to determine the face position of the first object based on the processed image and generate face coordinate information of the first object, wherein the face coordinate information is information used to characterize the position of face feature points of the first object; a third processing module, used to perform face alignment processing on the processed image based on the face coordinate information to obtain an image conforming to a preset shape; a fourth processing module, used to perform face encoding processing on the image conforming to the preset shape to obtain a first feature vector of the first object; and a third determining module, used to obtain a reference feature vector of the first object and determine the target image to be authenticated based on the first feature vector and the reference feature vector, wherein the reference feature vector is the archived feature vector of the first object.

[0085] Optionally, the third determining module includes: a first calculation module, used to calculate the similarity between the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; a first comparison module, used to compare the first similarity with a first threshold to obtain a first comparison result; a fourth determining module, used to take the second image as the target image to be authenticated if the first comparison result is that the first similarity is greater than the first threshold; and a fifth determining module, used to re-execute the step of acquiring images from the target video stream if the first comparison result is that the first similarity is less than or equal to the first threshold, until the first similarity is greater than the first threshold, and take the current image as the target image to be authenticated.

[0086] Optionally, the first processing module includes: a second comparison module, used to compare the target image to be authenticated with the reference image corresponding to the reference feature vector to obtain a second comparison result; a sixth determination module, used to determine the authentication result as successful if the second comparison result is successful; and a seventh determination module, used to obtain the number of comparison failures if the second comparison result is a failure, and determine the authentication result based on the number of failures.

[0087] Optionally, the seventh determining module includes: an eighth determining module, used to re-execute the step of image acquisition from the target video stream if the number of failures is less than or equal to the second threshold, and determine the authentication result based on the re-acquired image; and a ninth determining module, used to send a shooting instruction from the second object to the first object based on the video connection if the number of failures is greater than the second threshold, and capture a third image of the first object based on the video connection, and determine the authentication result based on the third image, wherein the shooting instruction is used to instruct the first object to adjust the face position.

[0088] Optionally, the identity authentication device further includes: a receiving module, used to receive a video call request before establishing a video connection between the first object and the second object in response to the video call request; a judging module, used to judge whether the network environment of the first device is in the target network environment; a switching module, used to enter the waiting response page corresponding to the video call request if the network environment of the first device is in the target network environment; and a generating module, used to generate target prompt information if the network environment of the first device is not in the target network environment, wherein the target prompt information is used to prompt the first object to switch network environments.

[0089] Example 3

[0090] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described authentication method at runtime.

[0091] Example 4

[0092] According to another aspect of the present invention, an electronic device is also provided, wherein, Figure 5 This is a schematic diagram of an optional electronic device according to an embodiment of the present invention, such as... Figure 5 As shown, the electronic device includes one or more processors; a memory for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to run the programs, wherein the programs are configured to execute the aforementioned authentication method during runtime. When the processor executes the program, it performs the following steps: responding to a video call request and establishing a video connection between a first object and a second object, wherein the video call request is a request initiated by the first object through a first device to conduct a video call with the second object; based on the video connection, acquiring a target video stream, and based on the target video stream, determining a target image to be authenticated, wherein the target video stream is a video stream within a preset duration range; and performing authentication on the first object based on the target image to be authenticated, obtaining an authentication result, wherein the authentication result is used to provide data reference for the second object.

[0093] Optionally, the processor may further perform the following steps when executing the program: acquiring a target video stream based on a video connection, including: acquiring a first image of a first object based on a video connection and displaying the first image to a second object; and acquiring the target video stream if the first image is displayed abnormally.

[0094] Optionally, the processor, when executing the program, further implements the following steps: determining the target image to be authenticated based on the target video stream, including: acquiring an image from the target video stream to obtain a second image of the first object; performing image preprocessing on the second image to obtain a processed image; determining the face position of the first object based on the processed image and generating face coordinate information of the first object, wherein the face coordinate information is information used to characterize the position of the face feature points of the first object; performing face alignment processing on the processed image based on the face coordinate information to obtain an image conforming to a preset shape; performing face encoding processing on the image conforming to the preset shape to obtain a first feature vector of the first object; obtaining a reference feature vector of the first object, and determining the target image to be authenticated based on the first feature vector and the reference feature vector, wherein the reference feature vector is an archived feature vector of the first object.

[0095] Optionally, the processor further implements the following steps when executing the program: determining the target image to be authenticated based on the first feature vector and the reference feature vector, including: calculating the similarity between the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; comparing the first similarity with a first threshold to obtain a first comparison result; if the first comparison result is that the first similarity is greater than the first threshold, then the second image is used as the target image to be authenticated; if the first comparison result is that the first similarity is less than or equal to the first threshold, then the step of acquiring the image from the target video stream is re-executed until the first similarity is greater than the first threshold, and the current image is used as the target image to be authenticated.

[0096] Optionally, when the processor executes the program, it further implements the following steps: based on the target image to be authenticated, the first object is authenticated to obtain an authentication result, including: comparing the target image to be authenticated with the reference image corresponding to the reference feature vector to obtain a second comparison result; if the second comparison result is a successful comparison, the authentication result is determined to be successful; if the second comparison result is a failed comparison, the number of failures is obtained, and the authentication result is determined based on the number of failures.

[0097] Optionally, the processor, when executing the program, further implements the following steps: determining the authentication result based on the number of failures, including: if the number of failures is less than or equal to a second threshold, then re-execute the step of acquiring images from the target video stream, and determine the authentication result based on the re-acquired image; if the number of failures is greater than the second threshold, then, based on the video connection, the second object sends a shooting instruction to the first object, and captures a third image of the first object based on the video connection, and determines the authentication result based on the third image, wherein the shooting instruction is used to instruct the first object to adjust its face position.

[0098] Optionally, the processor, when executing the program, also implements the following steps: before responding to the video call request and establishing a video connection between the first object and the second object, receiving the video call request; determining whether the network environment of the first device is in the target network environment; if the network environment of the first device is in the target network environment, entering the waiting response page corresponding to the video call request; if the network environment of the first device is not in the target network environment, generating target prompt information, wherein the target prompt information is used to prompt the first object to switch network environments.

[0099] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0100] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0101] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0104] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0106] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An identity authentication method, characterized in that, include: In response to a video call request, a video connection is established between a first object and a second object, wherein the video call request is a request initiated by the first object through a first device to conduct a video call with the second object; Based on the video connection, a target video stream is obtained, and based on the target video stream, a target image to be authenticated is determined, wherein the target video stream is a video stream within a preset duration range, and the target image to be authenticated is a facial image of a customer used for identity recognition and comparison. Based on the target image to be authenticated, the first object is authenticated to obtain an authentication result, wherein the authentication result is used to provide data reference for the second object; Based on the target video stream, the target image to be authenticated is determined, including: A second image of the first object is obtained by acquiring an image from the target video stream; The second image is preprocessed to obtain the processed image; Based on the processed image, the face position of the first object is determined, and the face coordinate information of the first object is generated, wherein the face coordinate information is information used to characterize the position of the face feature points of the first object; Based on the face coordinate information, the processed image is subjected to face alignment processing to obtain an image that conforms to a preset shape; The image conforming to the preset shape is subjected to face encoding processing to obtain the first feature vector of the first object; Obtain the baseline feature vector of the first object, and determine the target image to be authenticated based on the first feature vector and the baseline feature vector, wherein the baseline feature vector is the archived feature vector of the first object; Determining the target image to be authenticated based on the first feature vector and the reference feature vector includes: A similarity calculation is performed on the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; The first similarity is compared with the first threshold to obtain the first comparison result; If the first comparison result is that the first similarity is greater than the first threshold, then the second image is used as the target image to be authenticated; If the first comparison result is that the first similarity is less than or equal to the first threshold, then the step of acquiring images from the target video stream is repeated until the first similarity is greater than the first threshold, and the current image is used as the target image to be authenticated.

2. The method according to claim 1, characterized in that, Based on the video connection, the target video stream is obtained, including: Based on the video connection, a first image of the first object is captured, and the first image is displayed to the second object; If the first image displays abnormally, the target video stream is acquired.

3. The method according to claim 1, characterized in that, Based on the target image to be authenticated, the first object is authenticated to obtain an authentication result, including: The target image to be authenticated is compared with the reference image corresponding to the reference feature vector to obtain a second comparison result; If the second comparison result is a successful comparison, then the authentication result is determined to be successful. If the second comparison result is a comparison failure, then the number of comparison failures is obtained, and the authentication result is determined based on the number of failures.

4. The method according to claim 3, characterized in that, The authentication result is determined based on the number of failures, including: If the number of failures is less than or equal to the second threshold, the step of acquiring images from the target video stream is re-executed, and the authentication result is determined based on the re-acquired images. If the number of failures exceeds the second threshold, the second object sends a shooting instruction to the first object based on the video connection, and a third image of the first object is captured based on the video connection. The authentication result is determined based on the third image, wherein the shooting instruction is used to instruct the first object to adjust its face position.

5. The method according to claim 1, characterized in that, Before establishing a video connection between the first and second objects in response to a video call request, the method further includes: Receive the video call request; Determine whether the network environment of the first device is within the target network environment; If the network environment of the first device is in the target network environment, then the user will enter the waiting response page corresponding to the video call request. If the network environment of the first device is not in the target network environment, a target prompt message is generated, wherein the target prompt message is used to prompt the first object to switch network environments.

6. An identity authentication device, characterized in that, include: The first response module is used to respond to a video call request and establish a video connection between the first object and the second object, wherein the video call request is a request initiated by the first object through the first device to conduct a video call with the second object; The first determining module is used to obtain a target video stream based on the video connection, and to determine a target image to be authenticated based on the target video stream, wherein the target video stream is a video stream within a preset duration range, and the target image to be authenticated is a face image of a customer used for identity recognition and comparison. The first processing module is used to perform identity authentication on the first object based on the target image to be authenticated, and obtain an authentication result, wherein the authentication result is used to provide data reference for the second object; The first determining module further includes: a second acquisition module, used to acquire images from the target video stream to obtain a second image of the first object; a second processing module, used to preprocess the second image to obtain a processed image; a second determining module, used to determine the face position of the first object based on the processed image and generate face coordinate information of the first object, wherein the face coordinate information is information used to characterize the position of face feature points of the first object; a third processing module, used to perform face alignment processing on the processed image based on the face coordinate information to obtain an image conforming to a preset shape; a fourth processing module, used to perform face encoding processing on the image conforming to the preset shape to obtain a first feature vector of the first object; and a third determining module, used to obtain a reference feature vector of the first object and determine the target image to be authenticated based on the first feature vector and the reference feature vector, wherein the reference feature vector is the archived feature vector of the first object. The third determining module includes: a first calculation module, used to calculate the similarity between the first feature vector and the reference feature vector to obtain a first similarity between the first feature vector and the reference feature vector; a first comparison module, used to compare the first similarity with a first threshold to obtain a first comparison result; a fourth determining module, used to take the second image as the target image to be authenticated if the first comparison result is that the first similarity is greater than the first threshold; and a fifth determining module, used to re-execute the step of acquiring images from the target video stream if the first comparison result is that the first similarity is less than or equal to the first threshold, until the first similarity is greater than the first threshold, and take the current image as the target image to be authenticated.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the authentication method according to any one of claims 1 to 5 when it is run.

8. An electronic device, characterized in that, The electronic device includes one or more processors; A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to be configured to run the programs, wherein the programs are configured to execute the authentication method of any one of claims 1 to 5 at runtime.

Citation Information

Patent Citations

  • Face recognition method and device and training method and device of face recognition system

    CN111401344A

  • Identity authentication method and device, computing equipment and medium

    CN112367314A

  • Customer identity verification method and device, electronic equipment and storage medium

    CN112507314A