Face signing method, device, electronic device and storage medium

By extracting the user's mouth status and audio feature information from the video face-to-face signing data and sending relevant instructions for face-to-face signing authentication, the problems of low security and poor user experience in the prior art are solved, and higher security and user experience are achieved.

CN114973058BActive Publication Date: 2025-05-30CHENGDU NEW HOPE FINANCIAL INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210378288.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-05-30
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

There are problems with low security and poor user experience during the existing video interview process, especially for those with language dysfunction and users with accent or dialect, it is difficult to effectively conduct interview authentication.

Method used

By obtaining mouth status information and audio data from the user's face-to-face video data, sending face-to-face instructions to the user based on these feature information, instructing the user to make corresponding answers or complete related actions, and performing face-to-face authentication based on user feedback.

Benefits of technology

It improves the security and user experience of the interview process, prevents fraud in interviews, and adapts to the interview recognition of voice dysfunction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973058B_ABST
    Figure CN114973058B_ABST
Patent Text Reader

Abstract

The present application provides a face-to-face signing method, device, electronic device, and storage medium, which relate to the technical field of human-computer interaction. The method includes: obtaining feature information from the user's face-to-face signing video data, where the feature information includes the user's mouth state information and audio data; sending a face-to-face signing instruction to the user based on the feature information; and receiving and recognizing the user's feedback on the face-to-face signing instruction to obtain a recognition result, and performing face-to-face signing authentication on the user according to the recognition result. Using the method provided in the embodiments of the present application for face-to-face signing can collect more feature information from the user's face-to-face signing video data, and ask questions or give instructions to the user based on the feature information, and determine whether to pass the face-to-face signing authentication based on the user's feedback, thereby improving the security of the face-to-face signing process and improving the user experience by increasing the diversity and randomness of the questions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction. Specifically, it relates to a face-to-face signing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, in the industry, video face-to-face signing is unmanned. It relies on AI technology to perform voice-to-text voice broadcasts on users and uses speech recognition to identify users' answers to external questions, thereby completing intelligent face-to-face signing. The process of users' face-to-face signing often involves reading a paragraph or answering simple questions generated by the system.

[0003] However, in such a face-to-face signing scenario without customer service participation, if only relying on users to read a piece of text or answer questions, it is very easy to be attacked by attackers. Attackers can collect users' information in advance and perform face-to-face signing fraud by reading after or answering during face-to-face signing. For some special groups, such as users with language function disorders, it is impossible for such users to answer questions or read a paragraph; moreover, users may have accents and dialects when answering questions, which may lead to incorrect or inaccurate speech recognition. Therefore, there are currently problems of low security and poor user experience during face-to-face signing. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a face-to-face signing method, apparatus, electronic device, and storage medium, which are used to collect more feature information from the face-to-face signing video data of users, and based on the feature information, ask questions or give instructions to users, and judge whether to pass the face-to-face signing authentication based on the feedback of users. By increasing the diversity and randomness of questions, the security of the face-to-face signing process is improved.

[0005] In a first aspect, the embodiments of this application provide a face-to-face signing method, including:

[0006] Obtain feature information from the face-to-face signing video data of the user, where the feature information includes the mouth state information and audio data of the user;

[0007] Send a face-to-face signing instruction to the user based on the feature information;

[0008] Receive and identify the feedback of the user on the face-to-face signing instruction to obtain an identification result, and perform face-to-face signing authentication on the user according to the identification result.

[0009] In the above implementation process, feature information collected from the user's in-person signing video can be used, and based on the feature information, an in-person signing instruction related to the feature information can be sent to the user, instructing the user to give corresponding answers or complete relevant actions. By receiving the user's feedback on the in-person signing instruction, the user can be authenticated for in-person signing. Since in each in-person signing process in the embodiments of the present application, the user is authenticated based on the feature information in the in-person signing video, each in-person signing instruction can be a different one, so fraud in in-person signing can be effectively prevented. Since various feature information in the video can be collected in the embodiments of the present application instead of the traditional method of instructing the user to read after or answer questions, the security in the in-person signing process can be effectively improved and the user experience can be enhanced.

[0010] Optionally, the in-person signing instruction includes a first action instruction, and the body movements include at least one of head movements, torso movements, and limb movements;

[0011] Before sending the in-person signing instruction to the user based on the feature information, the method includes:

[0012] Determining whether the user is a person with language dysfunction based on the user's mouth state information and the audio data;

[0013] The sending of the in-person signing instruction to the user based on the feature information includes:

[0014] When the user is a person with language dysfunction, sending a first action instruction to the user, specifying that the user complete the body movement, receiving and identifying the user's feedback on the first action instruction to obtain the identification result, and authenticating the user for in-person signing according to the identification result.

[0015] In the above implementation process, it is possible to combine the user's mouth state and audio data to determine whether the user has a language disorder, and an instruction can be sent to instruct the user to complete head movements and / or hand movements. By identifying the user's actions, the in-person signing authentication can be completed, which can improve the flexibility of in-person signing and the user's in-person signing experience.

[0016] Optionally, the determining whether the user is a person with language dysfunction based on the user's mouth state information and the audio data includes:

[0017] Identifying the user's mouth state in the in-person signing video data based on a mouth state recognition model to obtain mouth state information;

[0018] When the number of words indicating that the user opens their mouth in the mouth state information reaches a threshold, obtain audio stream data from the face signing video data; obtain face signing audio from the audio stream data according to a preset feedback time, compare the face signing audio with the data segment of the user opening their mouth in the face signing video data, and obtain a comparison result; and determine whether the user is a person with a language disorder based on the comparison result.

[0019] In the above implementation process, it is possible to compare the user's mouth state and audio data to determine whether the user has a language disorder, so that when it is recognized that the user is a person with a language disorder, a face signing method for language disorder patients can be used to conduct face signing authentication for them, which can improve the face signing efficiency and the user's face signing experience.

[0020] Optionally, the head movement includes a facial movement and a head posture change movement. The facial movement is a movement that changes the posture of each facial organ, and the head posture change movement is a movement that changes the facial orientation.

[0021] The step of instructing the user to complete the limb movement to receive and recognize the user's feedback on the first action instruction, obtain the recognition result, and conduct face signing authentication for the user based on the recognition result includes:

[0022] Receive the image data of the user completing the facial movement and / or the head posture change movement.

[0023] Process the image data based on image classification recognition to obtain the first recognition result of the user completing the facial movement.

[0024] Based on head pose estimation, recognize the angle value of at least one of the user's head elevation angle, roll angle, and yaw angle, and obtain the second recognition result of the user completing the head posture change movement based on the angle value; and

[0025] Conduct face signing authentication for the user based on the first recognition result and the second recognition result.

[0026] In the above implementation process, it is possible to combine image classification and head pose estimation to recognize the user's facial movement feedback and head posture change movement feedback on the face signing instruction, thereby improving the applicability and flexibility of face signing.

[0027] Optionally, the step of instructing the user to complete the limb movement to receive and recognize the user's feedback on the first action instruction, obtain the recognition result, and conduct face signing authentication for the user based on the recognition result includes:

[0028] Receive the image data of the user completing the hand movement.

[0029] Identify multiple key points of the user's hand from the image data, and obtain a third recognition result of the user completing the hand movement based on the positional relationship between the multiple key points; and

[0030] Perform in-person signing authentication on the user based on the third recognition result.

[0031] In the above implementation process, on the basis of recognizing the user's head movement, a method of recognizing the user's hand movement is added, which can perform diverse and complete movement recognition on the user, enrich the recognition method, and thus improve the user's in-person signing experience.

[0032] Optionally, after comparing the in-person signing audio with the data segment of the user opening the mouth in the in-person signing video data to obtain a comparison result, the method further includes:[[]]

[0033] When the comparison result indicates that the user is not a person with speech dysfunction and the audio recognition result of the user does not match the preset answer to the question multiple times, confirm that the user has an accent and / or dialect; and send a first action instruction to the user to specify that the user complete the limb movement.

[0034] In the above implementation process, it is possible to judge whether the user conducts in-person signing authentication through dialect or whether an accent causes an error in the speech recognition result. By switching to sending an action instruction to the user to ensure the normal progress of in-person signing, the user experience can be improved.

[0035] Optionally, the feature information further includes at least one of the user's location information, image text information, scene information, and portrait information; wherein, the location information includes information representing the geographical location of the user, the image text information includes text information representing the natural image in the in-person signing video data, the scene information includes the type of scene where the user is located, and the portrait information includes information representing the dressing characteristics of the user;

[0036] The sending the in-person signing instruction to the user based on the feature information includes:[[]]

[0037] Send an in-person signing instruction to the user based on at least one of the location information, image text information, scene information, and portrait information.

[0038] In the above implementation process, multiple types of feature information can be extracted from the user's in-person signing video, and the user can be questioned according to the extracted feature information, which can increase the diversity and randomness of the questions in video in-person signing, reduce fraud in the video in-person signing process, and thus improve the security factor in unmanned and self-service video in-person signing.

[0039] Second aspect, an embodiment of the present application provides a face-to-face signing device, including:

[0040] An acquisition module, configured to acquire feature information from the face-to-face signing video data of a user, where the feature information includes the mouth state information and audio data of the user;

[0041] An instruction sending module, configured to send a face-to-face signing instruction to the user based on the feature information; and

[0042] An identification module, configured to receive and identify the feedback of the user on the face-to-face signing instruction, obtain an identification result, and perform face-to-face signing authentication on the user according to the identification result.

[0043] In the above implementation process, feature information collected from the face-to-face signing video of the user can be used, and a face-to-face signing instruction related to the feature information can be sent to the user based on the feature information, instructing the user to give corresponding answers or complete relevant actions. The user is authenticated for face-to-face signing by receiving the feedback of the user on the face-to-face signing instruction. Since each face-to-face signing process in the embodiment of the present application authenticates the user according to the feature information in the face-to-face signing video, each face-to-face signing instruction can be a different face-to-face signing instruction, so fraud in face-to-face signing can be effectively prevented. Since various feature information in the video can be collected in the embodiment of the present application instead of the traditional method of instructing the user to read after or answer questions, face-to-face signing recognition can also be performed on users with speech function disorders, thereby effectively improving the security and flexibility in the face-to-face signing process.

[0044] Optionally, the face-to-face signing instruction includes a first action instruction, and the body movement includes at least one of a head movement, a torso movement, and a limb movement;

[0045] The face-to-face signing device may further include:

[0046] A judgment module, configured to determine whether the user is a person with speech function disorder based on the mouth state information and the audio data of the user.

[0047] The instruction sending module may be specifically configured to:

[0048] When the user is a person with speech function disorder, send a first action instruction to the user, specifying the user to complete the body movement, so as to receive and identify the feedback of the user on the first action instruction, obtain the identification result, and perform face-to-face signing authentication on the user according to the identification result.

[0049] In the above implementation process, it is possible to determine whether the user has a language function disorder by combining the user's mouth state and audio data. When it is detected that the user has a language function disorder, an instruction can be sent to indicate the user to complete a head movement and / or a hand movement, and the face signing authentication can be completed by recognizing the user's movement, which can improve the flexibility and applicability of face signing.

[0050] Optionally, the determination module may be specifically configured to:

[0051] Recognize the user's mouth state of the face signing video data based on the mouth state recognition model to obtain mouth state information; when the number of words with the user's mouth open represented by the mouth state information reaches a threshold, obtain audio stream data from the face signing video data; obtain face signing audio from the audio stream data according to a preset feedback time, compare the face signing audio with the data segment of the user's mouth opening in the face signing video data, and obtain a comparison result; and determine whether the user is a person with a language function disorder based on the comparison result.

[0052] In the above implementation process, it is possible to compare the user's mouth state and audio data to determine whether the user has a language function disorder, so that when it is recognized that the user is a person with a language function disorder, a face signing method for language disorder patients can be used to conduct face signing authentication for them, which can improve the face signing efficiency and flexibility.

[0053] Optionally, the head movement includes a facial movement and a head posture change movement, the facial movement is a movement that changes the postures of various facial organs, and the head posture change movement is a movement that changes the facial orientation;

[0054] The recognition module may be specifically configured to:

[0055] Receive the image data of the user completing the facial movement and / or the head posture change movement; process the image data based on image classification recognition to obtain a first recognition result of the user completing the facial movement; estimate and recognize at least one of the elevation angle, roll angle, and yaw angle of the user's head based on head posture, and obtain a second recognition result of the user completing the head posture change movement based on the angle value; and conduct face signing authentication for the user based on the first recognition result and the second recognition result.

[0056] In the above implementation process, it is possible to combine image classification and head posture estimation to recognize the user's facial movement feedback and head posture change movement feedback for the face signing instruction, so as to improve the applicability and flexibility of face signing.

[0057] Optionally, the recognition module may also be specifically configured to:

[0058] Receive the image data of the user completing the hand movement; identify multiple key points of the user's hand from the image data, obtain a third recognition result of the user completing the hand movement based on the positional relationship between the multiple key points; and perform in-person signing authentication on the user based on the third recognition result.

[0059] In the above implementation process, on the basis of recognizing the user's head movement, a way of recognizing the user's hand movement is added, which can perform diverse and complete movement recognition on the user, enrich the recognition methods, and thus improve the user's in-person signing experience.

[0060] Optionally, the recognition module may specifically be used for:

[0061] When the comparison result indicates that the user is not a person with language dysfunction and the audio recognition result of the user does not match the preset answer to the question multiple times, confirm that the user has an accent and / or dialect; and send a first action instruction to the user to specify that the user complete the limb movement.

[0062] In the above implementation process, it is possible to determine whether the user conducts in-person signing authentication through dialect or whether an accent causes an error in the speech recognition result. By switching to sending an action instruction to the user, the normal progress of in-person signing can be ensured, and the user experience can be improved.

[0063] Optionally, the feature information further includes at least one of the user's location information, image text information, scene information, and portrait information; wherein, the location information includes information representing the geographical location of the user, the image text information includes the text information of the natural image in the in-person signing video data, the scene information includes the type of scene where the user is located, and the portrait information includes information representing the dressing characteristics of the user;

[0064] The instruction sending module may specifically be used for:

[0065] Send an in-person signing instruction to the user based on at least one of the location information, image text information, scene information, and portrait information.

[0066] In the above implementation process, various feature information can be extracted from the user's in-person signing video, and the user can be questioned according to the extracted feature information, which can increase the diversity and randomness of the questions in video in-person signing, reduce fraud in the video in-person signing process, and thus improve the security factor in unmanned and self-service video in-person signing.

[0067] In a third aspect, an embodiment of the present application provides an electronic device, which includes a memory and a processor. When the processor reads and runs the program instructions stored in the memory, it executes the steps in any of the above implementation manners.

[0068] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are read and run by a processor, the steps in any of the above implementation manners are executed. Description of the Drawings

[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 Schematic diagram of the steps of the face-to-face signing method provided by the embodiment of the present application;

[0071] Figure 2 Schematic diagram of the face-to-face signing steps applied to the scenario for language-disabled persons provided by the embodiment of the present application;

[0072] Figure 3 Schematic diagram of the steps for determining whether a user has a language disorder provided by the embodiment of the present application;

[0073] Figure 4 Schematic diagram of the steps for recognizing the head movements of a user provided by the embodiment of the present application;

[0074] Figure 5 Schematic diagram of the face-to-face signing steps for instructing a user to perform hand movements provided by the embodiment of the present application;

[0075] Figure 6 Schematic diagram of a face-to-face signing device provided by the embodiment of the present application. Detailed Embodiments

[0076] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, the functional modules in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0077] During the research process, the applicant found that the current problems in video face-to-face signing in the industry only include statement-type question recognition and question-and-answer type question recognition. Among them, the essence of a statement-type question is to specify that the user reads a passage aloud, and use speech recognition to recognize the speech audio read by the user. By determining whether the result text of the speech recognition is consistent with the text of the read question, the recognition of this question is completed; while the answer-type question is to ask the user a question, recognize the speech of the user's answer to obtain the recognition text result, and determine whether the user has answered the question based on whether the recognition text result contains the keywords in the answer or perform semantic-level recognition.

[0078] Currently, the types of questions in the face-to-face signing process in the industry are relatively few, and it is unable to well identify whether the person answering the questions is fraudulent, and the user experience is also relatively poor. Especially now, the mainstream video face-to-face signing is unmanned, relying on AI technology to perform voice broadcasts of text-to-speech for users, and using speech recognition to intelligently recognize the questions answered by users, so as to complete intelligent face-to-face signing.

[0079] In a first aspect, in such an intelligent face-to-face signing scenario without customer service participation, if relying solely on the user to read a text or answer questions, it is easily attacked by attackers. The attackers can completely collect the user's information in advance and perform shadowing or answering during the face-to-face signing to cause fraud. In a second aspect, for some special groups, such as users with speech function disorders during face-to-face signing, it is impossible for such users to answer questions or read a passage; moreover, the user may have an accent and dialect when answering questions, which may lead to incorrect or inaccurate speech recognition and affect the user experience. In a third aspect, when the user is unclear about the answer to a question, making a gesture is easier to understand than answering the question. Therefore, to solve the above problems, the embodiments of the present application provide a more secure face-to-face signing question design solution, which divides the face-to-face signing questions into four categories: statement questions, Q&A questions, instruction questions, and spontaneous questions. How to design and identify these questions is the problem to be solved by the present application.

[0080] First, the embodiments of the present application provide a face-to-face signing method, which collects and identifies the image and voice feature information in the user's face-to-face signing video, and sends a face-to-face signing instruction to the user according to the collected feature information to specify the user to make corresponding answers or complete relevant actions, so as to perform face-to-face signing authentication on the user, thereby enabling face-to-face signing recognition of users with speech function disorders and preventing fraud during the face-to-face signing process, improving the security of face-to-face signing and enhancing the user experience. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the steps of the face-to-face signing method provided by the embodiments of the present application. The implementation steps of face-to-face signing may include:

[0081] In step S11, feature information is obtained from the user's face-to-face signing video data, where the feature information includes the user's mouth state information and audio data.

[0082] The face-to-face signing video data may be a video recorded when confirming questions face-to-face with the user, a face-to-face signing video of an AI virtual human having a video call with the user, or video data obtained by the user taking a self-portrait for face-to-face signing independently. Face-to-face signing may be a procedure for the user to go to the loan bank to pay the fees required for the loan and conduct face-to-face interviews and sign. It can be applied to various types of video recording scenarios that require uploading standardized documents for business processes, such as entrusted payment and co-borrower signing, and can also be applied to risk warnings for large-amount and high-risk customer groups of ordinary consumer loans and business loans, as well as other scenarios that require the user to confirm information and save audio and video records.

[0083] Among them, the user can perform in-person signing authentication through a mobile terminal, or through a fixed terminal set up by a bank or other institutions. The mobile terminal can be an electronic device with an Internet connection function, and this electronic device can be a configurator of engineering equipment, a mobile phone, a tablet computer, a personal digital assistant, or a dedicated in-person signing terminal. The fixed terminal can be a computer, a server, etc. A camera device or a device capable of externally connecting a camera device can be set on the mobile terminal and the fixed terminal. The camera device can be a camera, and the in-person signing video data of the user is collected through the camera.

[0084] In step S12, an in-person signing instruction is sent to the user based on the feature information.

[0085] Among them, the feature information in this application can include features such as the user's mouth state, audio data extracted from the in-person signing video, portrait features, user location information, natural images in the in-person signing video, and the corresponding environment where the user is located.

[0086] Specifically, the audio data can be collected in real time, or it can be synthetic audio data obtained by segmenting and synthesizing the overall audio data of the user; the portrait features can be the user's clothing or the user's ornaments, such as the identification of the user wearing a hat, a scarf, and earrings. The identification method will be described in the corresponding part later and will not be elaborated here; the user location information can be obtained by allowing the user to authorize the location information to obtain the user's current location information. The natural image in the in-person signing video can be the background that appears in the user's in-person signing video. The environment where the user is located can include the type of scene where the user is located, and can include scene types such as indoors, outdoors, in a car, in an office, and in a mall.

[0087] The way to send the in-person signing instruction can be a screen pop-up window, can be a screen dynamic text prompt, can be played to the user based on a recorded voice, or can be a voice played to the user based on voice synthesis.

[0088] In step S13, the feedback of the user on the in-person signing instruction is received and recognized to obtain a recognition result, and the user is authenticated for in-person signing according to the recognition result.

[0089] Exemplarily, an in-person signing instruction can be sent to the user to instruct the user to complete corresponding hand movements, mouth movements, head movements, and to instruct the user to answer in-person signing questions. The user's feedback is the relevant actions taken or the corresponding questions answered after receiving the in-person signing instruction. The recognition result is the result indicating whether the user passes the in-person signing obtained after recognizing the user's feedback information. There are multiple ways for the user to pass the recognition. It can be to send multiple in-person signing instructions to the user and judge based on the proportion of the instructions completed by the user, or it can be to send one in-person signing instruction and judge based on the completion degree of the user's completion of this in-person signing instruction.

[0090] Identifying the hand movements of the user may include identifying through key point detection technology; the methods for identifying the mouth movements of the user may include identifying based on a mouth state recognition model and identifying the user's mouth shape based on DLIB, where DLIB is an open-source C++ toolkit containing machine learning algorithms; the method for identifying the head movements of the user may be identifying based on head pose estimation technology; identifying the user's answer to a question may be through speech recognition based on the speech data collected during the video face signing process, converting the speech into text, and matching the text with a preset question answer.

[0091] It can be seen that the embodiments of the present application are based on the feature information collected from the user's face signing video, and based on the feature information, send a face signing instruction related to the feature information to the user, instructing the user to make a corresponding answer or complete a related action, and authenticate the user by receiving the feedback of the user on the face signing instruction. Since each face signing process in the embodiments of the present application authenticates the user according to the feature information in the face signing video, each face signing instruction can be a different face signing instruction, so that fraud in face signing can be effectively prevented. Since various feature information in the video can be collected in the embodiments of the present application rather than the traditional method of instructing the user to read after or answer questions, face signing recognition can also be performed on speech function-impaired persons, thereby effectively improving the security during the face signing process and improving the user experience.

[0092] Next, on the basis of the above embodiments, the steps that may be included in the face signing method are continued to be introduced. Exemplarily, when the embodiments of the present application are applied to a scenario for persons with language function impairments, the face signing instruction may include a first action instruction, and the first action instruction is an instruction for specifying the user to complete a body movement, and the body movement includes at least one of head movement, torso movement, and limb movement. Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of the face signing steps provided by the embodiments of the present application for a scenario for persons with language function impairments. The face signing steps for persons with language function impairments may include:

[0093] In step S21, determine whether the user is a person with language function impairment based on the mouth state information of the user and the audio data.

[0094] In step S22, when the user is a person with language function impairment, send a first action instruction to the user, specifying the user to complete the body movement, so as to receive and identify the feedback of the user on the first action instruction, obtain the identification result, and authenticate the user according to the identification result.

[0095] Among them, obtaining feature information from the user's video data for in-person signing and sending an in-person signing instruction to the user based on the feature information before step S21 can refer to the above steps S11 to S12, which will not be elaborated here.

[0096] Here, the description directly starts from obtaining the user's mouth state and audio data. Exemplarily, the "language disordered person" in the embodiments of the present application refers to a user with speech disorders, whose language fluency and clarity are affected, including users with unclear pronunciation, stuttering or aphasia disorders.

[0097] Since users with language disorders cannot undergo in-person signing based on the conventional methods of following reading and answering questions, the applicant has designed a method for in-person signing authentication for users based on directive questions during the research process. By giving a certain instruction to the user to receive the user's feedback for in-person signing authentication. For example, if the user is a language disordered person, they cannot open their mouths to speak, and it is not realistic to ask the user to answer questions or follow a passage. Moreover, when the user answers questions, considering issues such as accents and dialects, the speech recognition may be incorrect or inaccurate, resulting in a poor user experience, and the user may try many times and still not pass. Additionally, when the user does not know the answer to the question, making a gesture is easier to understand than answering the question, which can improve the user's in-person signing experience.

[0098] Based on this, the embodiments of the present application collect the mouth state and audio data of the user during the in-person signing process, and pre-judge whether the user has a language disorder. When it is found that there are multiple overlapping pronunciations between the user's mouth shape and voice, the recognition rate between the mouth shape and voice is low, or the user only opens the mouth but has no sound, it can be judged that the user has a language disorder. Therefore, for this scenario, the in-person signing process is switched to the issue of asking the user to perform a certain action instruction. For example, the first action instruction sent in the embodiments of the present application instructs the user to complete a head movement or a hand movement.

[0099] Exemplarily, an instruction indicating an action can be randomly sent to the user, or an instruction with a preset order and quantity can be sent to the user. The head movements instructed for the user to complete can include opening the mouth, blinking, shaking the head, nodding, raising the head, covering the eyes, and covering the mouth, etc. The hand movements instructed for the user to complete can include making a fist, making an OK gesture, giving a thumbs up, making a 1 gesture, making a 2 gesture, making a 3 gesture, making a 4 gesture, and making a 5 gesture, etc. It can be understood that the above head movements and hand movements are only exemplary, and can be specifically set according to the situation during the actual in-person signing process.

[0100] In addition, the above-mentioned face-to-face signing steps applied to scenarios for people with language impairments can also be applied to the face-to-face signing process of normal users, or when the user's mouth is moving and there is sound, but the voice recognition results are wrong many times, it can be considered that the user does not understand the question, or the user's accent and dialect are relatively serious, resulting in more voice recognition errors. At this time, it can also be switched to the above-mentioned face-to-face signing steps.

[0101] It can be seen that the embodiments of the present application can combine the user's mouth state and audio data to determine whether the user has a language dysfunction. When it is detected that the user has a language dysfunction, instructions can be sent to instruct the user to complete head movements and / or hand movements. The face-to-face authentication is completed based on the identification of the user's movements, which can improve the flexibility of the face-to-face authentication and improve the user's face-to-face authentication experience.

[0102] In an optional embodiment, the present application embodiment also provides a method for determining whether a user has a language dysfunction. Figure 3 , Figure 3 A schematic diagram of the steps for determining whether a user has a language dysfunction is provided in an embodiment of the present application. The steps for determining whether a user has a language dysfunction may include:

[0103] In step S31, the user's mouth state is recognized on the face-to-face signing video data based on the mouth state recognition model to obtain the mouth state information.

[0104] Exemplarily, the mouth state recognition model can be a pre-trained convolutional neural network model, and the steps of training the mouth state recognition model can include detecting the face area based on the haar feature and AdaBoost algorithm or other face detection algorithms; extracting the feature points of the mouth on the obtained face area by combining random forest and linear regression; detecting the mouth area of ​​the face using the LBF feature based on the face feature points and combining the regularization method; constructing the core structure convolution layer of SR-Net; constructing the downsampling layer of SR-Net; using the rectified linear unit to construct the fully connected layer of SR-Net; setting the output values ​​of some neurons in the hidden layer to 0 with a preset probability (usually set to 0.5), and designing the Dropout of SR-Net; constructing a training sample set and selecting the corresponding network structure and number of iterations for training, so as to obtain the above-mentioned mouth state recognition model when the training is completed.

[0105] When recognizing the mouth state of a user, the face signature video data can be first segmented into frame-by-frame images. The user's mouth in each frame of the image is recognized, and the state is counted. For example, in which frames the user's mouth is in an open state and in which frames the user's mouth is in a closed state. Whether the user can open his mouth to speak is determined based on the number of frames in which the user's mouth is in an open state and the number of frames in which the user's mouth is in a closed state.

[0106] In step S32, when the number of words indicating that the user opens his mouth in the mouth state information reaches a threshold, audio stream data is obtained from the face signature video data; face signature audio is obtained from the audio stream data according to a preset feedback time, and the face signature audio is compared with the data segment of the user opening his mouth in the face signature video data to obtain a comparison result; and based on the comparison result, it is determined whether the user is a person with language dysfunction.

[0107] Exemplarily, the mouth state recognition result can be the number of frames in which the user's mouth is in an open state, or the duration in which the user's mouth is in an open state within a preset time period. When it is determined that the user can open his mouth to speak according to the mouth state recognition result, audio data is extracted from the face signature video data. Among them, the meaning of the preset feedback time is the time when it is expected that the user will receive an instruction to give feedback and answer questions. The preset feedback time can be set according to specific face signature questions. For example, when the face signature question is to ask the user's location, scene, identity and other information, the user's answering time is short. Therefore, the preset feedback time can be exemplarily set to 3 seconds or 5 seconds. When the face signature question is to require the user to do shadowing, the user's answering time is long. Therefore, the preset feedback time can be exemplarily set to 10 seconds or 15 seconds.

[0108] When comparing the obtained audio with the user's mouth state, an offline audio can be formed with the audio of the preset feedback time, and the audio data is subjected to sound classification processing to compare whether the audio is silent during the period when the user opens his mouth, so as to confirm whether the user can make a sound.

[0109] It can be seen that the embodiment of the present application can compare the mouth state of the user with the audio data to determine whether the user has language dysfunction, so that when it is recognized that the user is a person with language dysfunction, a face signature method for language-disabled persons can be used to conduct face signature authentication for him, which can improve the face signature efficiency and the user's face signature experience.

[0110] In an optional embodiment, the above-mentioned head movements in the description may include facial movements and head posture change movements. The facial movement is an action that changes the postures of various facial organs, and the head posture change movement is an action that changes the facial orientation. For step S22, the embodiment of the present application provides an implementation manner for recognizing the user's head movement. Please refer to Figure 4 ,Figure 4 The figure is a schematic diagram of the steps for identifying a user's head movement provided by an embodiment of the present application. The steps for identifying a user's head movement may include:

[0111] In step S41, image data of the user completing a facial movement and / or a head posture change movement is received.

[0112] In step S42, the image data is processed based on image classification recognition to obtain a first recognition result of the user completing the facial movement.

[0113] Facial movements during the face signing process may include actions such as opening the mouth, blinking, covering the eyes, and covering the mouth. The methods of image classification recognition may include classification based on the K-Nearest Neighbor (KNN) classification algorithm, support vector machines (SVM), or BP neural network. Specifically, KNN, SVM, and BP neural network can be run using sklearn. First, two different preprocessing functions are defined using the openCV package: the first defines the image feature vector, resizes the image, and then flattens the image into a list of row pixels. The second defines the extraction of the color histogram. The 3D color histogram can be extracted from the HSV color space using cv2.normalize, and then the flattened result is obtained. Then, the parameters to be parsed are constructed, and the accuracy of the above image feature vectors and color histograms is tested. For the entire dataset and subsets with different numbers of labels, the dataset is constructed as parameters parsed into the program. At the same time, the number of neighbors for the k-NN method is constructed as a parsed parameter. Each image feature is extracted from the dataset and placed in an array. Each image can be read using cv2.imread. The two features are extracted using the previously defined functions and appended to the arrays rawImages and features. The dataset is split using the train_test_split function imported from the sklearn package. 85% of the dataset can be used as the training set, and 15% as the test set. Finally, the KNN, SVM, and BP neural network functions are used to evaluate the user's facial data, thereby obtaining the first recognition result.

[0114] It should be understood that the above methods of identifying a user's completed facial movement by KNN, SVM, and BP neural network are merely illustrative. In the specific implementation process, a convolutional neural network (CNN) can also be constructed using TensorFlow to identify the user's facial movement.

[0115] In step S43, an angular value of at least one of the elevation angle, roll angle, and yaw angle of the user's head is recognized based on head pose estimation, and a second recognition result of the user completing the head pose change action is obtained based on the angular value.

[0116] Among them, head pose estimation can be interpreted as inferring the direction of a person's head relative to the camera view or inferring the direction of the head relative to the global coordinate system. However, this subtle difference requires knowledge of the inherent camera device parameters to eliminate the perceptual bias from perspective distortion. Exemplarily, the head movement range of an average adult male includes sagittal flexion and extension from -60.4° to 69.6° (i.e., backward movement from the neck), frontal lateral bending (i.e., the neck bends from right to left) from -40.9° to 36.3°, and horizontal axial rotation (rotation of the head to the left) from -79.8° to 75.3°

[26] . The combination of muscle rotation and relative orientation is an ambiguity that is often overlooked (for example, the contour view of the head does not look exactly the same when the camera views from the front compared to when the camera views from the front and the head turns to the side). Therefore, it is usually assumed that the human head is modeled as a bodiless rigid object. Under this assumption, the pose of the human head is restricted to three degrees of freedom (DOF), including Pitch (elevation angle), Yaw (roll angle), and Roll (yaw angle).

[0117] When estimating the human head pose, the new image of the head can be compared with a set of samples (each sample is labeled with a discrete pose) based on the appearance template method to find the most similar view and thus determine the head pose. A series of head detectors can be trained using the detector array method, each detector being adapted to a specific pose, and a discrete pose is assigned to the detector with the maximum support. The nonlinear regression method can be used to develop a functional mapping from image or feature data to head pose measurement using a nonlinear regression tool. The manifold embedding method can be used to seek a low-dimensional manifold to simulate the continuous change of the head pose. The new image can be embedded into these manifolds and then used for embedded template matching or regression. The flexible model can be used to fit a non-rigid model to the facial structure of each person in the image plane, and the head pose is estimated by feature-level comparison or instantiation of model parameters. The geometric method can be used to determine the pose of the relative configuration using the positions of features such as the eyes, mouth, and tip of the nose. The tracking method can also be used to recover the global pose change of the head from the movement between the observed video frames.

[0118] It should be understood that the above head pose estimation methods are only illustrative, and in the specific implementation process, multiple of the above methods can also be combined for head pose estimation.

[0119] In step S44, face signing authentication is performed on the user based on the first recognition result and the second recognition result.

[0120] Among them, image classification and head pose estimation can be used simultaneously to identify the user, determine whether the user makes a facial movement and / or the head pose change movement, or in the way of preset keywords, when keywords such as "opening the mouth, blinking, covering the mouth, covering the eyes" are detected, facial movement recognition is performed on the user, and when keywords such as "shaking the head, nodding, raising the head" appear, head pose estimation is performed on the user, so as to obtain a first recognition result indicating whether the user makes a facial movement and a second recognition result indicating whether the user makes a head pose change movement, and then determine whether the user gives a face signing feedback according to the instruction based on the first recognition result and the second recognition result.

[0121] It can be seen from this that the embodiments of the present application can combine image classification and head pose estimation to identify the facial movement feedback and head pose change feedback of the user for the face signing instruction, so as to improve the accuracy and flexibility of face signing.

[0122] In an optional embodiment, for step S22, in addition to instructing the user to make facial movements and head pose change movements, the user can also be instructed to make hand movements to complete face signing authentication. Please refer to Figure 5 , Figure 5 which is a schematic diagram of the face signing steps for instructing the user to make hand movements provided by the embodiments of the present application. The face signing steps for instructing the user to make hand movements may include:

[0123] In step S51, image data of the user completing the hand movement is received.

[0124] In step S52, multiple key points of the user's hand are identified from the image data, and a third recognition result of the user completing the hand movement is obtained based on the positional relationship between the multiple key points.

[0125] Exemplarily, the hand movements instructed to the user may include hand postures such as making a fist, making an OK gesture, giving a thumbs up, making a 1 gesture, making a 2 gesture, making a 3 gesture, making a 4 gesture, and making a 5 gesture. The key points of the hand can first establish a hand bone model, and each joint point in the bone model is used as a key point. When recognizing the user's hand movement, the hand posture can be determined according to the relative position of each joint point; multiple joint points of the hand can be preset as key points, and the relative position of each key point is determined during hand movement recognition to determine the hand posture. In addition, key points can be set separately for each action. When instructing the user to make a hand movement, the relative position of each key point corresponding to the action is recognized. A third recognition result indicating whether the user makes a hand movement is obtained according to the relative position between the key points.

[0126] In step S53, face signing authentication is performed on the user based on the third recognition result.

[0127] It should be understood that the above method for recognizing the user's hand movements can be used either as a method for individually conducting an in-person signature verification for the user or in combination with the recognition of the user's facial movements and head posture change movements in the above content to complete the in-person signature authentication for the user.

[0128] As can be seen, the embodiment of the present application adds a method for recognizing hand movements on the basis of recognizing the user's head movements, which can perform diverse and complete movement recognition on the user, enrich the recognition methods, and thus improve the user's in-person signature experience.

[0129] In an optional embodiment, for the comparison of the in-person signature audio with the data segment of the user opening the mouth in the in-person signature video data in step S32, in practical applications, in addition to the user having a language function disorder resulting in inability to perform speech recognition, it is also possible that the user conducts in-person signature authentication using a dialect. The pronunciation of dialects in some regions is quite different from that of Mandarin, and even the word order is different. Therefore, for step S32, after comparing the in-person signature audio with the data segment of the user opening the mouth in the in-person signature video data to obtain a comparison result, the method may further include:

[0130] When the comparison result indicates that the user is not a person with a language function disorder and the audio recognition result of the user does not match the preset answer to the question multiple times, confirm that the user has an accent and / or dialect; and send a first action instruction to the user to specify that the user complete the body movement.

[0131] Specifically, when it is recognized that the user's mouth shape is moving and there is sound, but the result of speech recognition is incorrect multiple times, it can be determined that the user does not understand the question or the user has a dialect resulting in incorrect speech recognition. Among them, a threshold for the number of incorrect questions or a threshold for the proportion of incorrect questions can be preset. When the number of incorrect questions answered by the user is higher than the threshold for the number of incorrect questions or the proportion of the questions answered by the user to the total number of questions is higher than the threshold for the proportion of incorrect questions, it can be determined that the user uses a dialect for in-person signature.

[0132] Exemplarily, the process of recognizing the user's head movements and hand movements can refer to the steps of S41 to S42 and S51 to S52 above, which will not be elaborated here. According to the recognition of the user's head movements and hand movements, a fourth recognition result indicating whether the user feedbacks head movements and / or hand movements is obtained, thereby completing the in-person signature authentication for the user.

[0133] As can be seen, the embodiment of the present application can determine whether the user conducts in-person signature authentication through a dialect or whether there is an accent resulting in an incorrect speech recognition result, and ensure the normal progress of the in-person signature by switching to sending an action instruction to the user, which can improve the user experience.

[0134] In an optional embodiment, the feature information further includes at least one of the user's location information, image text information, scene information, and portrait information; wherein, the location information includes information characterizing the geographical location of the user, the image text information includes text information characterizing the natural images in the face signing video data, the scene information includes the type of scene where the user is located, and the portrait information includes information characterizing the dressing features of the user.

[0135] Step S11 of sending a face signing instruction to the user based on the feature information may include:

[0136] Sending a face signing instruction to the user based on at least one of the location information, image text information, scene information, and portrait information.

[0137] In order to prevent fraud by users during the video face signing process, the applicant has designed a face signing method during the research process that can send spontaneous questions to users during face signing, thereby increasing the randomness.

[0138] Exemplarily, it may be possible to request the user to grant permission to obtain location information, obtain the user's location information in real time, and ask questions to the user based on the location information. For example, "Are you near the Financial City?" The specific implementation solution may be: when the user's location information is obtained, query the specific address where the location is located in the electronic map, and ask questions about the address. Key information in the address can be mainly asked, such as the name of the address and the name of the landmark building nearby.

[0139] It is possible to perform text recognition on the natural images in the user's video stream, classify the recognized text, and ask questions to the user. For example, if words such as "bank" or "mall" appear in the user's background, questions can be asked to the user, such as "Are you handling business at the bank?" or "Are you near the mall?" The implementation solution may be: during the video face signing process, the user can capture the user's image in real time through the camera, perform OCR recognition on the background of the user's image to obtain the text information that appears in the background, analyze the recognized text information, and process it based on preset rules, so as to map the text to questions about location, questions about the scene, or questions about behavior, etc.

[0140] Exemplarily, the analysis and processing of the recognized text information may specifically include: performing word segmentation on the text information, and the word segmentation tool can be the jieba word segmentation tool in Python; filtering out stop words and extracting keywords from all the word segments obtained after word segmentation. Among them, stop word filtering can be based on a stop word dictionary, and all the words included in the stop word dictionary are deleted from the word segments. Keyword extraction can be performed using the TF-IDF calculation method based on frequency or the TextRank calculation method based on graph iteration. Finally, questions are asked to the user based on the extracted keywords. Additionally, since the text information in the background of general user images is relatively less, the word segments filtered based on the stop word dictionary can be directly used as keywords to ask questions to the user.

[0141] Optionally, the image scene captured by the camera can be recognized, and questions can be asked about the scenario of the user's in-person signing. The scenarios mainly include: indoor scenes, outdoor scenes, in-vehicle scenes, office scenes, and shopping mall scenes, etc. The implementation solution can be: recognizing the shooting scene during the video in-person signing process, determining the shooting scene based on an image classification model, and asking questions to the user based on the scene, such as "Are you located inside the office?" Among them, the image classification model can be the VGG model, GoogleNet model, or PReLU-nets model, etc. The above image classification models are only proposed schematically. In actual application, other image classification models such as ResNet or Inceptionv3 can also be selected.

[0142] Optionally, the voice data collected during the video in-person signing process can be further recognized, and based on the voice, it can be determined whether the user is in an environment such as a shopping mall or outdoors, so as to ask questions to the user. The implementation solution can be: classifying the voice data, determining the scene corresponding to the voice, and asking questions to the user based on the voice scene. Among them, methods such as the threshold judgment method, Gaussian mixture model (GMM) method, hidden Markov model (HMM) method, artificial neural network (ANN) method, or support vector machine (SVM) method can be used.

[0143] Specifically, the steps of classifying voice data may include preprocessing the collected sound scene data to obtain audio data samples; performing channel separation and audio cutting on the preprocessed audio data samples to obtain multiple groups of audio data, extracting the corresponding gamma-tone filter cepstral coefficients and Mel spectrum features from each group of data, and calculating the first-order and second-order difference features of the Mel spectrum features to construct multiple groups of different input features; for multiple groups of different input features, multiple CNN models with different structures can be set as weak classifiers and each model can be trained; using a support vector machine as a strong classifier, stacking the output results of multiple models as the input features of the support vector machine, training the new fused model, and using the classification result of the new model as the final result of sound scene classification, so as to determine the current corresponding scene of the user.

[0144] Optionally, it is also possible to detect ornaments on the portrait captured by the camera, such as recognizing objects such as the user's hat, scarf, earrings, etc., so as to determine the portrait features, and questions can be asked to the user according to the portrait features. The implementation solution can be: perform face cropping on the portrait, perform object detection on the portrait based on the face-cropped data, so as to detect whether the portrait wears a hat, a scarf, earrings or has other ornaments, and then questions can be asked based on the user's ornaments, such as "Excuse me, do you wear a hat?"

[0145] It can be seen that the embodiments of the present application can extract various feature information from the user's video interview, and ask questions to the user according to the extracted feature information, which can increase the diversity and randomness of the questions in the video interview, reduce fraud in the video interview process, and thus can improve the safety factor in the unmanned and self-service video interview.

[0146] Based on the same inventive concept, the embodiments of the present application also provide an interview device 60. Please refer to FIG. 6. Figure 6 It is a schematic diagram of an interview device provided by the embodiments of the present application. The interview device 60 may include:

[0147] An acquisition module 61, configured to acquire feature information from the user's video interview data, where the feature information includes the user's mouth state information and audio data.

[0148] An instruction sending module 62, configured to send an interview instruction to the user based on the feature information.

[0149] An identification module 63, configured to receive and identify the user's feedback on the interview instruction, obtain an identification result, and perform interview authentication on the user according to the identification result.

[0150] Optionally, the interview instruction includes a first action instruction, and the body action includes at least one of a head action, a torso action, and a limb action;

[0151] The in-person signing device 60 may further include:

[0152] A judgment module, configured to determine whether the user is a person with language dysfunction based on the user's mouth state and the audio data.

[0153] The instruction sending module 62 may specifically be configured to:

[0154] When the user is a person with language dysfunction, send a first action instruction to the user, specifying that the user complete the limb movement, so as to receive and recognize the feedback of the user on the first action instruction, obtain the recognition result, and perform in-person signing authentication on the user according to the recognition result.

[0155] Optionally, the judgment module may specifically be configured to:

[0156] Recognize the user's mouth state of the in-person signing video data based on a mouth state recognition model to obtain mouth state information; when the mouth state information indicates that the number of words the user opens the mouth reaches a threshold, obtain audio stream data from the in-person signing video data; obtain in-person signing audio from the audio stream data according to a preset feedback time, compare the in-person signing audio with the data segment where the user opens the mouth in the in-person signing video data, and obtain a comparison result; and determine whether the user is a person with language dysfunction based on the comparison result.

[0157] Optionally, the head movement may include a facial movement and a head posture change movement, the facial movement is a movement that changes the postures of various facial organs, and the head posture change movement is a movement that changes the facial orientation;

[0158] The recognition module 63 may specifically be configured to:

[0159] Receive the image data of the user completing the facial movement and / or the head posture change movement; process the image data based on image classification recognition to obtain a first recognition result of the user completing the facial movement; estimate and recognize at least one of the elevation angle, roll angle, and yaw angle of the user's head based on head posture, and obtain a second recognition result of the user completing the head posture change movement based on the angle value; and perform in-person signing authentication on the user based on the first recognition result and the second recognition result.

[0160] Optionally, the recognition module 63 may further specifically be configured to:

[0161] Receive the image data of the user completing the hand movement; identify multiple key points of the user's hand from the image data, and obtain a third recognition result of the user completing the hand movement based on the positional relationship between the multiple key points; and perform in-person signing authentication on the user based on the third recognition result.

[0162] Optionally, the recognition module 63 may further be specifically configured to:

[0163] When the comparison result indicates that the user is not a person with speech function disorder and the audio recognition result of the user does not match the preset answer to the question multiple times, confirm that the user has an accent and / or dialect; and send a first action instruction to the user to specify that the user complete the limb movement.

[0164] Optionally, the feature information further includes at least one of the user's location information, image text information, scene information, and portrait information; wherein, the location information includes information representing the geographical location of the user, the image text information includes the text information representing the natural image in the in-person signing video data, the scene information includes the type of scene where the user is located, and the portrait information includes the information representing the dressing characteristics of the user;

[0165] The instruction sending module 62 may be specifically configured to:

[0166] Send an in-person signing instruction to the user based on at least one of the location information, image text information, scene information, and portrait information.

[0167] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, which includes a memory and a processor. When the processor reads and runs the program instructions stored in the memory, it executes the steps in any of the above implementation manners.

[0168] Among them, the processor may include one or more (only one is shown in the figure), which may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it may also be a dedicated processor, including a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Moreover, when there are multiple processors, a part of them may be general-purpose processors, and another part may be dedicated processors.

[0169] The memory may include one or more, which may be, but are not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0170] The processor and other possible components can access the memory, read and / or write data therein. In particular, one or more computer program instructions may be stored in the memory, and the processor can read and run these computer program instructions to implement the methods provided in the embodiments of the present application.

[0171] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are read and run by a processor, the steps in any of the above implementation manners are executed.

[0172] The computer-readable storage medium may be various media that can store program codes, such as a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the storage medium is used to store a program. After receiving an execution instruction, the processor executes the program. The method executed by the electronic terminal defined by the process disclosed in any embodiment of the present invention can be applied to the processor or implemented by the processor.

[0173] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.

[0174] In addition, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0175] Furthermore, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0176] It is replaceable and can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part.

[0177] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).

[0178] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article, or device comprising the said elements.

[0179] The above are only embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A face-to-face signing method, characterized in that, it includes: Obtain feature information from the user's face-to-face signing video data, where the feature information includes the user's mouth state information and audio data; Determine whether the user is a person with language dysfunction based on the user's mouth state information and the audio data; Send a face-to-face signing instruction to the user based on the feature information; where the face-to-face signing instruction includes a first action instruction, and the first action instruction is an instruction for specifying the user to complete a body movement, and the body movement includes at least one of a head movement, a torso movement, and a limb movement; and Receive and identify the feedback of the user on the face-to-face signing instruction, obtain an identification result, and perform face-to-face signing authentication on the user according to the identification result; Among them, the sending the face-to-face signing instruction to the user based on the feature information includes: when the user is a person with language dysfunction, sending a first action instruction to the user to specify the user to complete the body movement, so as to receive and identify the feedback of the user on the first action instruction, obtain the identification result, and perform face-to-face signing authentication on the user according to the identification result.

2. The method according to claim 1, characterized in that, The determining whether the user is a person with language dysfunction based on the user's mouth state information and the audio data includes: Performing user mouth state recognition on the face-to-face signing video data based on a mouth state recognition model to obtain the mouth state information; When the number of words in the mouth state information indicating that the user opens the mouth reaches a threshold, obtain audio stream data from the face-to-face signing video data; obtain face-to-face signing audio from the audio stream data according to a preset feedback time, compare the face-to-face signing audio with the data segment of the user opening the mouth in the face-to-face signing video data, and obtain a comparison result; and determine whether the user is a person with language dysfunction based on the comparison result.

3. The method according to claim 2, characterized in that, After comparing the face-to-face signing audio with the data segment of the user opening the mouth in the face-to-face signing video data to obtain a comparison result, the method further includes: When the comparison result indicates that the user is not a person with language dysfunction and the audio recognition result of the user does not match the preset answer to the question multiple times, confirm that the user has an accent and / or dialect; and send a first action instruction to the user to specify the user to complete the body movement.

4. The method according to claim 1, characterized in that, wherein, The head movement includes a facial movement and a head posture change movement, the facial movement is a movement that changes the postures of various facial organs, and the head posture change movement is a movement that changes the facial orientation; The specifying the user to complete the body movement to receive and identify the feedback of the user on the first action instruction, obtain the identification result, and perform face-to-face signing authentication on the user according to the identification result includes: Receiving the image data of the user completing the facial movement and / or the head posture change movement; Process the image data based on image classification recognition to obtain a first recognition result of the user completing the facial action; Based on head pose estimation, recognize the angular value of at least one of the elevation angle, roll angle, and yaw angle of the user's head, and obtain a second recognition result of the user completing the head pose change action based on the angular value; and Perform face signing authentication on the user based on the first recognition result and the second recognition result.

5. The method according to claim 1, wherein, wherein, the limb actions include hand actions. Specifying the user to complete the limb action to receive and recognize the user's feedback on the first action instruction, obtaining the recognition result, and performing face signing authentication on the user according to the recognition result includes: Receiving the image data of the user completing the hand action; Recognize multiple key points of the user's hand from the image data, and obtain a third recognition result of the user completing the hand action based on the positional relationship between the multiple key points; and Perform face signing authentication on the user based on the third recognition result.

6. The method according to claim 1, wherein, the characteristic information further includes at least one of the user's location information, image text information, scene information, and portrait information; wherein, the location information includes information representing the user's geographical location, the image text information includes text information representing the natural image in the face signing video data, the scene information includes the type of scene where the user is located, and the portrait information includes information representing the user's wearing characteristics; The sending the face signing instruction to the user based on the characteristic information includes: Sending a face signing instruction to the user based on at least one of the location information, image text information, scene information, and portrait information.

7. A face signing device, wherein, comprising: An acquisition module, configured to acquire characteristic information from the user's face signing video data, wherein the characteristic information includes the user's mouth state information and audio data; A judgment module, configured to determine whether the user is a person with speech dysfunction based on the user's mouth state information and the audio data; An instruction sending module, configured to send a face signing instruction to the user based on the characteristic information; wherein, the face signing instruction includes a first action instruction, and the first action instruction is an instruction for specifying the user to complete a limb action, and the limb action includes at least one of a head action, a torso action, and a limb action; and A recognition module, configured to receive and recognize the user's feedback on the face signing instruction, obtain a recognition result, and perform face signing authentication on the user according to the recognition result; wherein, the sending the face signing instruction to the user based on the characteristic information includes: when the user is a person with speech dysfunction, sending a first action instruction to the user, specifying the user to complete the limb action, to receive and recognize the user's feedback on the first action instruction, obtain the recognition result, and perform face signing authentication on the user according to the recognition result.

8. An electronic device, It is characterized in that including a memory and a processor, wherein computer program instructions are stored in the memory, and when the computer program instructions are read and run by the processor, the method according to any one of claims 1-6 is executed.

9. A computer-readable storage medium It is characterized in that computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are run by a processor, the steps in the method according to any one of claims 1-6 are executed.

Citation Information

Patent Citations

  • Remote video face-to-face signing method and system

    CN106203024A

  • Face-to-face signature verification method and device, computer equipment and storage medium

    CN112288398A