A method for automatic entry of identity information via voice interaction

By using a voice-interactive method for entering name and identity information, combined with idiom word formation and radical character decomposition, the problem of low recognition accuracy in Chinese name entry has been solved, achieving flexible and user-friendly Chinese entity entry, applicable to various scenarios.

CN118447584BActive Publication Date: 2025-12-02NANJING NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410548792.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-12-02
Estimated Expiration
2044-05-06

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy in Chinese name input, especially in speech recognition where 100% accuracy cannot be achieved. Furthermore, traditional methods heavily rely on known data, limiting their application scenarios.

Method used

We employ a voice-interactive method for automatic entry of name and identity information, combined with named entity recognition and seamless facial recognition. We use idioms and radical decomposition to correct Chinese character references and design a multi-turn dialogue algorithm to eliminate ambiguity.

Benefits of technology

It enables flexible and user-friendly Chinese entity input, improving input efficiency and accuracy. It is suitable for various complex scenarios and supports clear reference to named entities such as Chinese names and place names.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447584B_ABST
    Figure CN118447584B_ABST
Patent Text Reader

Abstract

This invention discloses a voice-interactive automatic identity information input method, based on a voice-interactive name identity information automatic input module based on named entity recognition and a seamless input module for accompanying face identity information. It includes the following steps: (1) determining the single-character subject in the text with speech conversion errors; (2) using Chinese character referencing correction methods based on idioms and radical decomposition; (3) designing a multi-turn dialogue algorithm for name correction; (4) determining the focus user; (5) filtering high-quality face templates; and (6) focus user identity recognition and human-computer interaction based on face tracking. This invention fills the gaps in the field of human-computer voice interaction Chinese character referencing correction and accurate input, and also provides a more user-friendly design scheme for identity input systems in welcoming scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of natural language processing and computer vision, and in particular, focuses on an innovative method for automatic voice-interactive identity information entry. This method not only improves the efficiency and accuracy of identity information entry but also greatly optimizes the user experience, providing a brand-new solution for identity information management in the era of intelligent technology. Background Technology

[0002] With the continuous progress of society and the rapid development of information technology, the management and identification of identity information has become increasingly crucial. In many fields such as security checks, access control, meeting check-in, and welcoming services, identity recognition and information entry have become indispensable. However, regrettably, the interactive performance of most intelligent systems in these fields remains relatively low. Faced with unknown data, they often rely on inefficient and rigid manual entry methods, which are not only simplistic but also offer a poor user experience.

[0003] In voice input, due to the inherent ambiguity of speech information, voice recognition technology cannot guarantee the complete accuracy of the recognized text. Therefore, systems that use voice input for identity information are currently rare in the market. This is especially true for Chinese names, which are highly unique, with flexible and varied combinations and unpredictable meanings. Names are entirely determined by the individual; they can be either "Wang Erniu" or "Wang Erniu," with no absolute right or wrong. Given this characteristic, voice recognition cannot achieve 100% accuracy when recognizing Chinese names.

[0004] In today's era of sweeping intelligent technology, traditional data entry methods are proving inadequate and urgently need transformation. An intelligent and user-friendly data entry method not only meets the needs of modern development but is also key to improving user experience and increasing work efficiency. Therefore, exploring and developing more intelligent and convenient data entry technologies has become an important task in the field of information technology.

[0005] Most existing patented methods for correcting personal name entries rely on matching with known databases of names, thus severely limiting their application scenarios. For example, patent CN114818668A calculates the minimum edit distance between the current candidate name and names in a known database and selects the name with the minimum distance as the correction result; while patent CN114138940A performs pronunciation similarity matching and sorting based on names in a known database to achieve name correction. Similarly, patent CN115630632A also relies on existing specific data for correction, mainly depending on domain-specific knowledge graphs and contextual semantics to determine the matching relationship between names and information such as occupation, position, and region. While these methods can solve the name correction problem to some extent, their heavy reliance on known data significantly limits their application scenarios.

[0006] This invention innovatively employs a dialogue-based interactive approach, combining idiom-based word formation with radical-based character decomposition for Chinese character reference correction. This enables interactive voice correction of any Chinese entity, such as names, place names, and organization names. Compared to existing patents, this invention is not only more flexible and versatile but also offers a more user-friendly experience and greater applicability, capable of addressing Chinese entity correction needs in various complex scenarios. Summary of the Invention

[0007] This invention relates to a voice-interactive method for automatically entering identity information. The core technologies involve voice-interactive name and identity information entry technology and non-intrusive facial recognition identity information entry technology. Specifically, it includes the following steps:

[0008] A voice-interactive automatic identity information input method is characterized by: using a voice-interactive automatic name identity information input module based on named entity recognition and a face identity information input module that is used in combination, wherein the user's voice input is based on the voice-interactive automatic name identity information input module based on named entity recognition, and the user's video image input is used in conjunction with the face identity information input module.

[0009] The voice-interactive automatic name and identity information input module based on named entity recognition includes:

[0010] (1) Determining the subject of a single character in a text with incorrect speech conversion: Analyze the Chinese character intended to be corrected from the text with inaccurate speech conversion;

[0011] (2) Chinese character reference correction method based on idioms and radicals: analyze user intent and correct errors using idioms or radicals;

[0012] (3) Design of multi-round dialogue algorithm for name error correction: Design a complete, highly fault-tolerant and diverse dialogue logic to handle various possible situations;

[0013] The无感式伴随人脸身份信息录入模块 includes:

[0014] (4) Focus user determination: Ignore those who communicate unintentionally, screen and focus on the target user;

[0015] (5) High-quality face template screening: For the focus user in the conversation, screen out high-quality face templates and store them in the database for user identity recognition;

[0016] (6) Focus user identity recognition and human-computer interaction based on face tracking: Track the focus user in the conversation to maintain the consistency of the focus user's identity.

[0017] As a further improvement of the present invention, step (1) specifically includes:

[0018] ① Determination of single-word subject based on dependency syntax analysis: Perform dependency syntax analysis on the user message text, find the nominal subject or root word therein, and verify whether the single-word subject in the sentence is successfully found by combining phonetic similarity analysis; <​​​​​​​​​​​​​​​③推定基于已知偏旁部首的汉字:根据偏旁部首列表缩小候选集范围,先筛选出所有带“氵”的,然后再筛选包含“州”的,最终根据读音判断正误,完成纠错。

[0025] As a further improvement of the present invention, step (3) specifically includes:

[0026] ①Definition of dialogue variables: Define multi-turn dialogue-related variables for name error correction, including: the user intention to be recognized in the dialogue, the actions to be taken by the system, the entities to be extracted from the user message text, and the slots to be set during the dialogue process;

[0027] ②Recognition of dialogue state: Recognize the user intention and set slots after extracting entities from the message text, providing a basis for the system to take corresponding actions;

[0028] ③Analysis of name error correction methods: The designed Chinese character reference error correction methods for name error correction include error correction based on idiom word formation and radical character decomposition description, precise correction based on the specified Chinese character method, and the machine asking the user to verify the Chinese character presumption result, etc. Take the single-character subject extracted in step 1: clarify the character to be corrected. For example, if the voice conversion of the user message "宋 is the 宋 in 唐宗宋祖" is incorrect and the text is "送 is the 送 in 唐宗宋祖", the single-character subject is "送". In step 2, combine the single-character subject and determine the correct form of the single-character subject through idiom word formation and radical character decomposition description. For example, determine that "送" is actually "宋" through the idiom word formation "唐宗宋祖". Precise correction based on the specified Chinese character method: In the case where the single-character subject cannot be analyzed, clarify the character to be corrected according to the digital information entity given by the user, such as "the second character is wrong". The machine asks the user to verify the Chinese character presumption result: After completing the error correction, explain each character in the name in terms of radical, idiom word formation, and number of strokes to confirm the correctness with the user;

[0029] ④Decision-making of dialogue behavior: Based on the dialogue strategy and the current dialogue state, decide the next dialogue action and continue or end the dialogue by means of inquiry. With the constructed dialogue strategy as the guidance and decision-making, determine the next action that the system should take based on the current user intention, the extracted entities, the slot state, and the form state.

[0030] As a further improvement of the present invention, step (4) specifically includes:

[0031] ①Capture of potential users based on face detection and attribute intelligent analysis: Perform face detection on the video image and conduct face attribute analysis, including face key points, gender, expression, whether wearing glasses, whether having a beard, and face features;

[0032] ② User filtering based on face size and tilt angle: Determine whether the face is close and facing the user based on the face image size and tilt angle. If the face image is smaller than the threshold and the tilt angle is greater than the threshold, it is considered that there is no intention to communicate, thus filtering out people who pass by and have no intention to communicate.

[0033] ③ Focus user identification based on dwell time: users whose dwell time exceeds a threshold are identified as focus users;

[0034] ④ Determining the communication order among multiple focus users: If multiple focus users appear at the same time, the distance is determined by the size of the face image, and communication is carried out in order from the closest to the farthest.

[0035] ⑤ Focused communication user reminder based on facial attribute description: Add the facial attribute description of the current communication user to the conversation to provide a reminder.

[0036] As a further improvement of the present invention, step (5) specifically includes:

[0037] ① Face quality screening based on explicit information of pose, lighting, and occlusion: filtering faces with excessively large deflection angles, excessively strong or dark lighting, or excessively large occlusion areas;

[0038] ② Intelligent screening of face quality based on implicit comprehensive information: Implicit comprehensive information is covered by constructing high-quality and low-quality sample sets, and modeling and analysis are performed using artificial intelligence models;

[0039] ③ When tracking, the faces are entered into the database based on explicit and implicit information filtering methods.

[0040] As a further improvement of the present invention, step (6) specifically includes:

[0041] ① Identity consistency label determination based on face tracking: Combine tracking algorithms to obtain consistency labels of faces in video footage;

[0042] ②Focus User Identity Recognition Integrating Face Quality Analysis and Consistency Label Constraints: Unlike single-frame face identity recognition, a multi-frame face identity recognition and voting scheme is adopted, and combined with identity consistency label constraints of tracked faces, to obtain stable face identity recognition in time series;

[0043] ③ Maintain the identity information of key users: Maintain the identity information of key users and synchronize and update the identity information with the dialogue system.

[0044] The beneficial effects of this invention are as follows:

[0045] 1. This invention realizes name and identity information input based on voice interaction, which has the advantages of flexibility, user-friendliness and ambiguity elimination.

[0046] 2. This invention enables seamless input of facial identity information. Based on stable facial recognition in video frames and facial image quality analysis, it can provide focus user identification and locking functions for voice-interactive name and identity information input, thereby improving the targeting and effectiveness of name and identity information input.

[0047] 3. The core right of this invention is the identification and correction of single characters based on idioms and radicals. This technology can be used to eliminate ambiguity and clarify the reference of all named entities, not limited to the reference of names, such as the identification of homophones of navigation place names based on voice interaction in driving scenarios. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments are briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating a voice-interactive method for automatically entering identity information.

[0050] Figure 2 This is a flowchart illustrating the method for determining single-character subjects in speech-to-text errors.

[0051] Figure 3 This is a flowchart of the Chinese character reference error correction method based on idiom word formation and radical character decomposition.

[0052] Figure 4 This is a flowchart illustrating the design of a multi-turn dialogue algorithm for name correction.

[0053] Figure 5 This is a flowchart of the dialogue strategy for name correction;

[0054] Figure 6 This is a flowchart of a multi-turn dialogue logic in an embodiment of the present invention;

[0055] Figure 7 This is a flowchart for determining the focus user;

[0056] Figure 8 This is a flowchart for selecting high-quality face templates;

[0057] Figure 9 This is a flowchart of focus-based user identification and human-computer interaction based on face tracking;

[0058] Figure 10 This is a flowchart of multi-frame face recognition and voting to stabilize the recognition results. Detailed Implementation

[0059] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0060] This invention provides a method for automatic entry of identity information via voice interaction, the overall process of which is as follows: Figure 1 As shown, its specific implementation method is as follows: It adopts a combination of a voice interactive automatic name and identity information input module based on named entity recognition and a face identity information input module that is not sensitive to human presence. The user's voice input is based on the voice interactive automatic name and identity information input module based on named entity recognition, and the user's video image input is accompanied by the face identity information input module that is not sensitive to human presence.

[0061] The automatic name and identity information entry module based on named entity recognition for voice interaction includes:

[0062] (1) Determining the subject of a single character in a text with incorrect speech conversion: Analyze the Chinese character intended to be corrected from the text with inaccurate speech conversion;

[0063] (2) Chinese character reference correction method based on idioms and radicals: analyze user intent and correct errors using idioms or radicals;

[0064] (3) Design of a multi-turn dialogue algorithm for name correction: Design a sound, fault-tolerant, and diverse dialogue logic to deal with various possible situations;

[0065] The accompanying facial recognition data entry module includes:

[0066] (4) Focus user identification: Ignore people who have no intention of communicating, and filter and focus on target users;

[0067] (5) High-quality face template filtering: High-quality face templates are filtered out for the focus user in the conversation and stored in the database to facilitate user identification;

[0068] (6) Focus user identification and human-computer interaction based on face tracking: Track the focus user in the conversation and maintain the consistency of the focus user's identity.

[0069] In this embodiment, step (1) involves determining the single-word subject in the text with speech conversion errors, as shown in the flowchart below. Figure 2 As shown, the specific steps are as follows:

[0070] 1) Determination of single-character subject based on dependency syntactic analysis. Perform dependency syntactic analysis on the user message text using a natural language processing library. Preferably, use spacy for dependency syntactic analysis to find the nominal subject or root word therein. After finding it, use the lazy_pinyin library to obtain the Chinese character pinyin and perform string similarity analysis with the name pinyin. If the similarity is higher than the threshold, it is considered that the single-character subject is found;

[0071] 2) Determination of single-character subject based on analysis of the词性 (should be "词性" in Chinese, meaning "word class" in English) before and after demonstrative pronouns. Since there are errors in speech conversion and the single-character subject cannot be directly found through syntactic analysis, based on demonstrative pronouns such as "的" and "是" in the sentence, analyze the word class and pronunciation information of the words before and after them. If it is a single character and the pronunciation similarity with a certain character in the name exceeds the threshold, it is considered that the single-character subject is found;

[0072] 3) Determination of single-character subject based on pronunciation similarity analysis. Preferably, perform pronunciation similarity analysis on the idiom phrase entities extracted by the entity extractor of Rasa and match them with the previously extracted name slot values. If the similarity exceeds the threshold, it is considered that the single-character subject is found.

[0073] In this embodiment, the method for correcting Chinese character reference based on idiom phrase composition and character decomposition by radical in step (2) has a flowchart as Figure 3 shown, and the specific steps are as follows:

[0074] 1) Pronunciation similarity analysis of idioms in the idiom phrase correction method. For the idiom phrase correction method, perform pronunciation similarity analysis on the idiom and the name, and then combine it with the single-character subject to determine whether the three contain a certain character with extremely similar pronunciation at the same time, and thus perform correction;

[0075] 2) Analysis of the description diversity of radicals in the Chinese character reference correction method described by radical decomposition. For the radical correction method, perform analysis of the description diversity of radicals. First, convert the radical entities extracted from the user message text into common names through the constructed radical and its common name library, such as converting "water radical"

[0076] to "three dots of water" i.e. "氵". It is also possible that it is not a radical but a combination, for example, "on the right is [Xuzhou](compound)'s Zhou", then according to the demonstrative word and pronunciation similarity analysis, take out "zhou" and store it in the radical list;

[0077] 3) Deduction of Chinese characters based on known radicals. Reduce the candidate set range according to the radical list and screen in the constructed radical and its Chinese character dictionary. The dictionary is like {"氵": "没, 洗, 滴, 酒"}. First, screen out all the characters with the "氵" radical, and then screen out those containing "zhou". Finally, judge the correctness according to the pronunciation to complete the correction.

[0078] In this embodiment, step (3) is for designing a multi-turn dialogue algorithm for name correction. The algorithm design process is as follows: Figure 4 As shown, the dialogue strategy flow is as follows: Figure 5 As shown, Figure 6 The following is an example, with the specific steps as follows:

[0079] 1) Dialogue Variable Definition: Define multi-turn dialogue-related variables for name correction, including: user intent to be identified in the dialogue, actions the system needs to take, entities to be extracted from user message text, and slots to be set during the dialogue. User intents include user_give_name, user_give_compound, etc.

[0080] user_give_pianpang, as shown in Table 1. Actions include action_handle_multi_intent,

[0081] The action_show_name is shown in Table 2. Entities such as name, compound, and radical are shown in the table.

[0082] As shown in Figure 3. Slots such as isOK and num are detailed in Table 4.

[0083] 2) Dialogue State Recognition: Identify user intent and extract entities from the message text, then assign slots to provide a basis for the system to take corresponding actions. Preferably, use the intent classifier from the open-source task-oriented dialogue system framework Rasa.

[0084] DIETClassifier identifies intent. Preferably, entity extraction is performed using a Rasa-based entity extractor.

[0085] The CRFEntityExtractor and RegexEntityExtractor entity extractors obtain all entities and their types from the user message text and set slots. For example, if the entity type is a radical and the value is "three dots of water", a radical slot is set.

[0086] 3) Analysis of name correction methods: The name correction method includes Chinese character reference correction based on idioms and radicals, precise correction based on Chinese character specification, and machine verification of the user's inference results. (31) Extracting the subject of a single character in step 1: Identify the character that the user wants to correct in a sentence, such as the incorrect text of the user message "Song is the Song of Tang Zong and Song Zu" which is "Song is the Song of Tang Zong and Song Zu", and its subject is "Song". (32) Combining the subject of a single character in step 2, determine the correct form of the subject of a single character through idioms and radicals. For example, determine the correct form of "Song" through the idiom "Tang Zong and Song Zu".

[0087] “Song”. (33) Precise correction based on Chinese character specification: When the subject of a single character cannot be analyzed, it can also identify which character needs to be corrected based on the numerical information entity that the user may provide, such as “the second character is wrong”. (34) The machine verifies the user’s inference of Chinese characters: After the correction is completed, each character in the name is explained using radicals, idioms, number of strokes, etc., so as to confirm the correctness of the name to the user in reverse.

[0088] 4) Dialogue Behavior Decisions: Based on the dialogue strategy and the current dialogue state, determine the next dialogue action, continuing or ending the dialogue through inquiries. Guided by the constructed dialogue strategy, decisions are made based on the current user intent and existing dialogue state.

[0089] The extracted entity, slot status, and form status determine the next action the system should take.

[0090] In this embodiment, step (4) focus user determination, the flowchart is as follows: Figure 7 As shown, the specific steps are as follows:

[0091] 1) Capture potential user faces based on face detection. (11) Collect training samples of faces. (12) Label face bounding boxes and their locations. (13) Preferably, use MTCNN multi-task cascaded convolutional neural network for real-time face detection.

[0092] (14) Train the model to learn the face region in the image. (15) The camera transmits each frame of the image to MTCNN. The model outputs the user's face image F and represents it as a quadruple data (x,y,w,h), which represent (horizontal coordinate, vertical coordinate, width of face rectangle, height of face rectangle), respectively.

[0093] 2) Face attribute analysis based on deep learning. (21) Preferably, the face attribute dataset is CelebA.

[0094] (22) Preferably, AttributeCNN is used for facial attribute prediction. (23) The face image F is passed to the model, and the output face attributes include, but are not limited to, gender, age, beard, glasses, etc.

[0095] 3) Acquisition of facial landmarks and facial features. (31) Preferably, 68 facial landmarks can be obtained quickly and easily through the dlib facial landmark detector; (32) Collect facial training samples, each containing multiple facial images.

[0096] (33) Use MTCNN to process the training samples, retain only the face region and crop it to form a new training dataset. (34) Preferably, use FaceNet deep neural network for face feature extraction. (35) Train the model to learn the Euclidean features of face images, so that the faces of the same person are spatially similar and the faces of different people are separated as much as possible. (36) Pass the face image F to FaceNet to obtain a 512-dimensional feature vector.

[0097] 4) Face deflection angle calculation. (41) Preferably, deep learning-based methods often consume more resources than traditional calculation methods, thus failing to meet the requirements of real-time analysis. Therefore, the face deflection angle is calculated by combining facial key points with geometric transformation. The deflection angle includes the yaw angle, pitch angle, and roll angle. The yaw angle represents the left-right rotation angle, the pitch angle represents the up-down rotation angle, and the roll angle represents the rotation angle around the horizontal axis. (42) Preferably, the affine transformation matrix from the 2D face in the image to the 3D face model is estimated using the solvePnP function of OpenCV.

[0098] (43) The rotation vector and translation vector are obtained by inputting 68 facial key points into solvePnP. (44) Preferably, the rotation vector and translation vector are processed by OpenCV's Rodrigues, hconcat, and...

[0099] The `decomposeProjectionMatrix` function ultimately yields Euler angles, namely Yaw, Pitch, and Roll.

[0100] 5) User filtering based on face size and tilt angle. The yaw angle, pitch angle, and face image F are combined to determine whether the person is close and looking directly at the face. Preferably, if the w or h of the face image F is less than 80° and the absolute value of the tilt angle yaw or pitch is greater than 60°, the person is considered to have no intention to communicate and no further steps are taken, thus filtering out people who are passing by and have no intention to communicate.

[0101] 6) Determining the focus user based on dwell time. (61) Preferably, the DeepSORT tracking algorithm is used for tracking.

[0102] (62) Create a tracker for each detected face. (64) Update the tracker in each frame, handle faces that have not appeared for too long as they are lost, and update the lifespan of tracked users. (63) Determine the user's dwell time based on the tracker's lifespan t. If the dwell time t > 3s, the user is identified as the focus user.

[0103] 7) Determining the order of interaction among multiple focus users. (71) Distance is determined by the size of the face image, where size = w * h. (72) Interactions are conducted sequentially from near to far, with larger sizes indicating closer distances and smaller sizes indicating farther distances. (73) During interaction, the labels of the currently interacting users are maintained, which are timestamps for strangers. (74) When no interaction is being conducted, the labels of the currently interacting users are updated to the labels of the closest users.

[0104] 8) Focused user alerts based on facial attribute descriptions. Add the user's name to the chat during the current interaction.

[0105] Facial attribute descriptions are used for reminders; simply transmit the user's facial attribute analysis information to the Rasa dialogue system. In this embodiment, step (5) involves high-quality facial template screening, and the process is as follows: Figure 8 As shown, the specific steps are as follows:

[0106] 1) Face quality screening based on explicit information of pose, lighting, and occlusion. (11) Obtain lighting and occlusion information based on the face attributes analyzed in step (4). (12) Obtain the calculated face deflection angle based on step (4). (13) Filter out faces with excessive deflection angle, excessively strong or dark lighting, or excessively large occlusion areas;

[0107] 2) Intelligent screening of face quality based on implicit comprehensive information. (21) Collect face training samples. (22) Evaluate the face training samples in terms of contrast, brightness, sharpness, focal length and illumination, and give evaluation scores. (23) Use artificial intelligence models for modeling and analysis. Preferably, use FaceQnet to integrate multiple image features and standardize quality measures to achieve accurate evaluation of face quality. (24) Input the face image F into FaceQnet to obtain the face image quality score S.

[0108] 3) Faces are entered into the database based on explicit and implicit information filtering during tracking. (31) During tracking, the face images of the focus user are filtered based on explicit and implicit information. (32) High-quality faces are added to the tracker. (33) When the identity information maintained by the tracker changes, the user's face image is stored in the database.

[0109] In this embodiment, step (6) is based on focus user identification and human-computer interaction using face tracking, such as... Figure 9 The specific steps are as follows:

[0110] 1) Identity consistency label determination based on face tracking. (11) Preferably, DeepSORT tracking algorithm is used to track the focus user. (12) The face and its features detected in each frame are passed to the tracking algorithm for tracker update, including feature update and trajectory update of the tracked object. (13) The association between different faces in consecutive frames is determined by feature matching and trajectory merging, and the identity consistency label of each focus user is returned;

[0111] 2) Focused user identification integrating face quality analysis and consistency label constraints. (21) Before focused user identification, the face is first evaluated as described in step (5). If the face quality score S is less than 0.4, identification is skipped. (22) The face identity at a certain moment is not simply based on the recognition result of the current frame, but rather a vote is taken on the recognition results of thirty consecutive frames, and the longest continuous and stable recognition result is selected as the face identity. In this way, even if the user's face deflection during the communication process causes an identification error, the face identity recognition is maintained stably over time under the consistency label constraints and voting scheme given by face tracking. The process is as follows. Figure 10 As shown.

[0112] 3) Maintain the identity information of the focus user. (31) Maintain the identity information of the focus user during tracking. The latest user identity information is obtained by the focus user identity recognition. (32) If the identity information changes after the focus user identity recognition, the old name tag and the updated name tag of the focus user are stored as key-value pairs in the cache database Redis, and the Rasa dialogue system is notified to update the information of the communication object synchronously. (33) After receiving the update information, Rasa retrieves the key-value pairs from Redis and updates the current communication user. (34) When Rasa successfully corrects the name, it tells the face recognition module in the same way. After receiving the information, it updates the name information and identity of the communication user in the database.

[0113] In this embodiment, Table 1 represents the intents defined in the multi-turn dialogue; Table 2 represents the entities defined in the multi-turn dialogue; Table 3 represents the actions defined in the multi-turn dialogue; and Table 4 represents the slots defined in the multi-turn dialogue.

[0114] Table 1

[0115]

[0116] Table 2

[0117]

[0118] Table 3

[0119]

[0120] Table 4

[0121]

[0122]

[0123] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatically recording identity information via voice interaction, characterized in that, Adopt a method that combines a voice interactive name identity information automatic input module based on named entity recognition and a non-intrusive input module for accompanying face identity information. Among them, the user's voice input is based on the voice interactive name identity information automatic input module of named entity recognition, and the user's video image input is based on the non-intrusive input module for accompanying face identity information; The voice interactive name identity information automatic input module based on named entity recognition includes: (1) Determination of single-character subject in text with incorrect voice conversion: Analyze a certain Chinese character whose intention is to correct errors from the text with inaccurate voice conversion; (2) Chinese character reference error correction method based on idiom word formation and radical character decomposition description: Analyze the user's intention and perform Chinese character reference error correction in the way of idiom word formation or radical character decomposition description; (3) Design of multi-round dialogue algorithm for name error correction: Design the dialogue logic, select the appropriate name error correction method according to the dialogue state, combining steps 1 and 2; Extract the single-character subject in step 1 and complete the Chinese character reference error correction method based on idiom word formation and radical character decomposition description in step 2; In the case where the single-character subject cannot be analyzed, perform precise correction based on the specified Chinese character method; After the error correction, the machine asks the user to verify the Chinese character presumption result; The non-intrusive input module for accompanying face identity information includes: (4) Focus user determination: Ignore people with unintentional communication, screen and focus on the target user; (5) High-quality face template screening: Screen high-quality face templates for the focus user during the conversation for face database entry and retrieval; (6) Focus user identity recognition and human-computer interaction based on face tracking: Track the focus user during the conversation to maintain the consistency of the focus user's identity; The focus user identity recognition and human-computer interaction based on face tracking synchronizes the information of the focus user to the dialogue system in real time to achieve the synchronization and update of identity information.

2. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, The specific content of step (1) includes: (11) Determination of single-character subject based on dependency syntax analysis: Perform dependency syntax analysis on the user's message text, find the nominal subject or root word therein, and verify whether the single-character subject in the sentence is successfully found by combining the analysis of pronunciation similarity; (12) Determination of single-character subject based on the analysis of the词性 before and after demonstrative pronouns: According to the demonstrative pronouns "的" and "是" in the sentence, combine syntax analysis of the词性 and pronunciation information of the words before and after them to verify whether the single-character subject is successfully found; (13) Determination of single-character subject based on pronunciation similarity analysis: Directly perform pronunciation similarity analysis on the entity extracted by the entity extractor, match it with the extracted name slot value, and if certain conditions are met, it is considered that the single-character subject is found.

3. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, The specific content of step (2) includes: (21) Analysis of pronunciation similarity of idioms in the idiom word formation error correction method: For the Chinese character reference error correction method of idiom word formation, perform pronunciation similarity analysis on idioms and names, and then combine the single-character subject to confirm whether error correction is needed; (22) Analysis of the diversity of radical character descriptions in the Chinese character reference error correction method of radical character decomposition description: For the Chinese character reference error correction method of radical character decomposition description, perform analysis of the diversity of radical character descriptions; (23) Based on known radicals, the candidate set is narrowed down according to the list of radicals, and finally the correctness is judged according to the pronunciation to complete the error correction.

4. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, Step (3) of the multi-turn dialogue algorithm design for name correction specifically includes: (31) Dialogue variable definition: Define multi-turn dialogue-related variables for name correction, including: user intent to be identified in the dialogue, actions to be taken by the system, entities to be extracted from user message text, and slots to be set during the dialogue. (32) Dialogue state recognition: Recognize user intent and extract entities from message text and set slots; (33) Dialogue behavior decision: Based on the dialogue strategy and the current dialogue status, decide the next dialogue action and continue or end the dialogue by asking questions.

5. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, Step (4) specifically includes: (41) Capturing potential users based on face detection and intelligent attribute analysis: performing face detection and face attribute analysis on video images, including facial key points, gender, expression, whether glasses are worn, whether there is a beard, and facial features; (42) User filtering based on face size and deflection angle: Based on the face image size and deflection angle, it is determined whether the person is close and looking directly at the face. If the face image is smaller than the threshold and the deflection angle is greater than the threshold, it is considered that there is no intention to communicate, thus filtering out people who pass by and have no intention to communicate. (43) Focus user determination based on dwell time: users whose dwell time exceeds the threshold are determined to be focus users; (44) Determining the order of communication among multiple focus users: If multiple focus users appear at the same time, the distance is determined by the size of the face image, and communication is carried out in order from near to far. (45) Focused communication user reminder based on facial attribute description: Add the facial attribute description of the current communication user to the conversation to provide a reminder.

6. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, Step (5), the screening of high-quality face templates, specifically includes: (51) Face quality screening based on explicit information of pose, lighting and occlusion: filtering faces with excessive deflection angle, excessively strong or dark lighting, or excessively large occlusion area; (52) Intelligent screening of face quality based on implicit comprehensive information: Implicit comprehensive information is covered by constructing high-quality and low-quality sample sets, and modeling and analysis are performed using artificial intelligence models; (53) Faces are entered into the database based on explicit and implicit information filtering methods during tracking.

7. The method for automatic entry of voice-interactive identity information as described in claim 1, characterized in that, Step (6) based on face tracking for focused user identification and human-computer interaction specifically includes: (61) Identity consistency label determination based on face tracking: obtain the consistency label of the face in the video frame by combining the tracking algorithm; (62) Focused user identity recognition integrating face quality analysis and consistency label constraints: Unlike single-frame face identity recognition, a multi-frame face identity recognition and voting scheme is adopted, and combined with identity consistency label constraints of tracked faces, to obtain stable face identity recognition in time series; (63) Maintain the identity information of the focus user: Maintain the identity information of the focus user and synchronize and update the identity information with the dialogue system.

Citation Information

Patent Citations

  • Speech recognition error correction method and device based on artificial intelligence and storage medium

    CN107220235A

  • Automatic face registration and recognition system and method in monitoring scene

    CN112183162A

  • Intelligent voice office robot system based on RASA

    CN114138940A