A method for obtaining user tags for human-computer interaction

By identifying users through voiceprint features and facial contour samples, and dynamically adjusting the accuracy of the speech recognition model, the system solves the problem of accuracy and efficiency in human-computer interaction systems when faced with different user accents and dialects, achieving efficient, personalized, adaptable, and accurate speech recognition.

CN120766688BActive Publication Date: 2026-01-30BEIJING MEILAN ZHIDA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510610449.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2026-01-30
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Existing human-computer interaction systems struggle to guarantee the accuracy and efficiency of speech recognition when faced with differences in accents, dialects, and speech habits among different users. Furthermore, they lack effective identification and utilization of user identity and personalized characteristics, and cannot adapt to changes in the language environment in real time.

Method used

By extracting the voiceprint features of the wake-up voice to identify the user's identity, the accuracy of the machine speech recognition model is adjusted. The speech recognition trend is determined by combining speech test samples and facial contour samples. The model accuracy is dynamically adjusted to adapt to changes in user accents and dialects. The fuzzy fields are judged by the conversion time and confusion level of text interaction information, and the model accuracy is optimized in real time.

Benefits of technology

It achieves fast and accurate user identification, improves the efficiency and accuracy of speech recognition, adapts to different language environments, maintains efficient and smooth interaction, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766688B_ABST
    Figure CN120766688B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology, and more particularly to a method for acquiring user tags for human-computer interaction, comprising: responding to a wake-up voice, extracting the phonetic features of the wake-up voice to identify the identity information status of the interactive user and determine the model accuracy of the machine speech recognition model; receiving the voice interaction information of the interactive user, and converting the voice interaction information into text interaction information according to the speech recognition model; determining whether the text interaction information contains fuzzy fields based on the conversion time and perplexity of the text interaction information; and adjusting the corresponding model accuracy based on the determination result of the presence of fuzzy fields, combined with the position and proportion of the fuzzy fields. This invention provides a method for acquiring user tags for human-computer interaction that can accurately identify user identity, effectively handle voice differences, and adapt to changes in the language environment in real time, thereby improving the accuracy and efficiency of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for obtaining user tags for human-computer interaction. Background Technology

[0002] In the digital age, human-computer interaction (HCI) technology is widely used in smart speakers, intelligent customer service, smart homes, and many other fields, drastically changing people's lifestyles and work habits. However, existing HCI systems face numerous challenges: on the one hand, significant differences in accents, dialects, and speech habits among users make it difficult to guarantee the accuracy of speech recognition, leading to low interaction efficiency; on the other hand, traditional speech recognition models often lack effective identification and utilization of user identity and personalized characteristics, making it difficult to adaptively adjust to specific user situations; furthermore, when users are in different language environments, their speech characteristics dynamically change, and existing technologies struggle to capture and adapt to these changes in real time. Therefore, there is an urgent need for a user tag acquisition method for HCI that can accurately identify user identity, effectively address speech differences, and adapt to changes in the language environment in real time, in order to improve the accuracy and efficiency of HCI.

[0003] Chinese Patent Publication No. CN111914076B discloses a method, system, terminal, and storage medium for constructing user profiles based on human-computer dialogue. The method includes: acquiring dialogue data during human-computer interaction; inputting the dialogue data into an attribute classification model, whereby the attribute classification model encodes the dialogue data using a dialogue encoder, classifies the encoded dialogue data by attribute type using an attribute classifier, extracts the subject and object corresponding to each attribute type using an entity generator, concatenates the subject and object with the corresponding attribute type, and outputs a triplet consisting of the subject, attribute type, and object; and constructs a user profile based on the triplet information. This invention can automatically extract explicit or implicit user attribute information of different attribute types in human-computer dialogue within a unified framework, improving the flexibility and accuracy of user attribute information extraction. Therefore, the existing technology has the following problems:

[0004] Different language environments can cause changes in users' dialects and accents, resulting in reduced efficiency in voice interaction. Summary of the Invention

[0005] To address this issue, the present invention provides a method for obtaining user tags for human-computer interaction, thereby overcoming the problem in the prior art where different language environments can cause changes in user dialects and accents, resulting in reduced voice interaction efficiency.

[0006] To achieve the above objectives, the present invention provides a method for obtaining user tags for human-computer interaction, comprising:

[0007] In response to a wake-up voice, the phonetic features of the wake-up voice are extracted;

[0008] Based on the aforementioned voiceprint features, the identity information and status of the interactive user are identified, and the model accuracy of the machine speech recognition model is determined, including:

[0009] In response to users with unknown identities, the accuracy of the machine speech recognition model is set using the user's voice test samples and facial contour samples.

[0010] In response to an interactive user with a known identity status, the machine speech recognition model is adjusted to the corresponding model accuracy based on the user's speech recognition status.

[0011] Receive voice interaction information from interactive users, and convert the voice interaction information into text interaction information according to the voice recognition model;

[0012] Determine whether the text interaction information contains ambiguous fields based on the conversion time and the perplexity of the text interaction information.

[0013] The model accuracy is adjusted based on the determination of the presence of fuzzy fields, combined with the location and proportion of the fuzzy fields.

[0014] Furthermore, in response to an interactive user with an unknown identity, the process of adjusting the machine speech recognition model to the corresponding model accuracy based on the user's speech recognition status includes,

[0015] In response to an interactive user with an unknown identity, obtain the user's voice test sample and facial contour sample;

[0016] Based on the voice test samples and the facial contour samples, the corresponding interactive user's voice-image association and geographic affiliation are determined to determine their voice recognition trend.

[0017] The speech recognition state is determined based on the speech test samples and the speech recognition trend in order to set the model accuracy of the machine speech recognition model.

[0018] Furthermore, the process of identifying the identity information status of the interactive user based on the aforementioned voiceprint features includes,

[0019] In response to the wake-up voice, the machine's identity information database and identity comparison function are invoked;

[0020] Determine the similarity between the voiceprint features of the wake-up voice and the features of each user tag in the identity information database;

[0021] The identity information status of the interacting user is determined based on the aforementioned similarity scores, wherein...

[0022] If any of the aforementioned similarities is greater than or equal to a preset similarity, then the identity information status of the interacting user is determined to be a known identity status.

[0023] If all the aforementioned similarities are less than the preset similarity, then the identity information status of the interacting user is determined to be unknown.

[0024] The user tag feature is the voiceprint feature of the corresponding interactive user's wake-up voice.

[0025] Furthermore, the process of determining the geographic location of the corresponding interactive user's voice-image association based on the voice test samples and the facial contour samples to determine its voice recognition trend includes,

[0026] Based on the speech test samples, determine the regional attribution of speech and the trend of accent representation;

[0027] The facial geographic affiliation is determined based on the facial contour sample.

[0028] The voice-image associated geographic affiliation of the corresponding interactive user is determined based on the voice geographic affiliation and the facial geographic affiliation.

[0029] The speech recognition trend is determined based on the geographical affiliation of the audio-visual association.

[0030] The accent representation trend includes a weak accent representation trend and a strong accent representation trend, and the speech recognition trend includes a strong speech recognition trend, a weak speech recognition trend, and a difficult speech recognition trend.

[0031] Furthermore, the process of determining speech recognition trends based on the geographical affiliation of the audio-visual association includes,

[0032] The geographical affiliation of the audio-visual association is compared with the pre-stored local speech recognition database;

[0033] Determine the speech recognition trend based on the comparison results;

[0034] The local speech recognition database is a mapping relationship between regional affiliation and speech recognition trends.

[0035] Furthermore, the speech recognition state is determined based on the speech test samples and the speech recognition trend in order to set the model accuracy of the machine speech recognition model;

[0036] Obtain the accent representation trend corresponding to the speech test sample;

[0037] Based on the accent representation trend, determine whether to correct the speech recognition trend to determine the speech recognition state, wherein...

[0038] If the accent representation trend is a weak accent representation trend, then the speech recognition trend is corrected and the speech recognition state is determined based on the corrected speech recognition trend.

[0039] If the accent representation trend is a strong accent representation trend, then it is determined that the speech recognition trend will not be corrected and the speech recognition state will be determined based on the speech recognition trend.

[0040] The model accuracy of the machine speech recognition model is set according to the speech recognition status;

[0041] The speech recognition states include strong speech recognition state, weak speech recognition state, and difficult speech recognition state.

[0042] Furthermore, the model accuracy of the machine speech recognition model is set according to the speech recognition status, including:

[0043] If the speech recognition state is a strong speech recognition state, then the model accuracy is a low model accuracy.

[0044] If the speech recognition state is a difficult speech recognition state, then the model accuracy is a medium model accuracy.

[0045] If the speech recognition state is a weak speech recognition state, then the model accuracy is a high model accuracy.

[0046] Furthermore, the process of determining whether the text interaction information contains fuzzy fields and their locations based on the conversion time and the perplexity of the text interaction information includes:

[0047] Obtain the conversion time of the text interaction information;

[0048] Determine the perplexity of text interaction information based on language models;

[0049] The presence of ambiguous fields in the corresponding text interaction information is determined based on the conversion time and the level of confusion.

[0050] Further, based on the conversion time and the confusion level, it is determined whether the corresponding text interaction information contains ambiguous fields, including:

[0051] If the conversion time is greater than the preset time and the confusion level is less than the preset confusion level, then it is determined that the corresponding text interaction information contains a fuzzy field.

[0052] If the conversion time is less than or equal to the preset time, and / or the confusion level is greater than or equal to the preset confusion level, then it is determined that the corresponding text interaction information does not contain any ambiguous fields.

[0053] Furthermore, the process of adjusting the model accuracy based on the determination of the presence of fuzzy fields, combined with the location and proportion of the fuzzy fields, includes:

[0054] The conversion time of each character in the text interaction information is determined based on the judgment result of the existence of fuzzy fields.

[0055] The blur position and blur ratio of the blur field are determined based on the conversion time of each character;

[0056] The model accuracy is adjusted according to the fuzzy position and fuzzy ratio of the fuzzy field, wherein,

[0057] If the blurred position is in the middle and the blurred ratio is greater than the preset ratio, then it is determined that the model accuracy should be adjusted.

[0058] If the blurred position is at the edge and / or the blurred ratio is less than the preset ratio, then it is determined that the model accuracy will not be adjusted.

[0059] Compared with existing technologies, the beneficial effects of this invention are as follows: The user tag acquisition method for human-computer interaction provided by this invention can quickly and accurately identify the identity information status of the interactive user based on the voiceprint features of the wake-up voice. For new users with unknown identity status, the accuracy of the machine voice recognition model is set through voice test samples and facial contour samples to create a tailored model for new users. For existing users with known identity status, the model is adjusted to the corresponding accuracy based on the user's voice recognition status to achieve personalized service and improve recognition efficiency and interaction accuracy. Secondly, after receiving user voice interaction information, this invention can quickly convert it into text interaction information with the help of an adapted voice recognition model, ensuring the smooth progress of the interaction process, reducing information transmission delay, and improving user experience. Thirdly, by considering the conversion time and confusion of text interaction information, it effectively determines whether there are fuzzy fields in the text. Once a fuzzy field is detected, the model accuracy is adjusted in real time based on its position and proportion, thereby adapting to changes in user accents and dialects in different language environments, avoiding the reduction in voice interaction efficiency due to these factors, and always maintaining high interaction efficiency and accuracy.

[0060] Furthermore, for interactive users with unknown identities, by acquiring voice test samples and facial contour samples, multimodal information fusion analysis is used to determine the geographical affiliation of the voice-image association, thereby clarifying the voice recognition trend. Then, by combining the voice test samples and the recognition trend, the voice recognition status is determined, and the accuracy of the machine voice recognition model is set accordingly. This achieves the following: from accurately collecting user feature data from multiple dimensions, to predicting the voice recognition difficulty based on regional characteristics, and then dynamically adjusting the model accuracy according to the actual situation, it greatly improves the adaptability and accuracy of the voice recognition model to the voices of new users, laying a solid foundation for efficient human-computer interaction in the future.

[0061] Furthermore, by comprehensively considering the conversion time and perplexity of text interaction information, the system can accurately determine whether there are ambiguous fields in the text. The conversion time can intuitively reflect the time consumption of the speech-to-text conversion process. Combined with the perplexity calculated using language models such as BERT and GPT, the system analyzes the conversion efficiency and text logic quality. When the conversion time is greater than the preset time and the perplexity is less than the preset perplexity, the system determines that there are ambiguous fields; otherwise, it determines that they do not exist. This method effectively integrates multi-dimensional indicators to accurately locate possible ambiguous information, providing strong support for improving the accuracy and efficiency of human-computer interaction in the future. It also helps to promptly identify and resolve comprehension biases that may occur in voice interaction.

[0062] Furthermore, based on the determination of whether there are fuzzy fields in the text interaction information, the position and proportion of fuzzy fields are accurately located by determining the conversion time of each character, thereby deciding whether to adjust the model precision. This process cleverly uses the conversion time to judge fuzzy characters, connects them into fuzzy fields, calculates the precise fuzzy proportion, and clearly defines the edge and middle positions: when the fuzzy position is in the middle and the fuzzy proportion is greater than the preset proportion, the model precision is adjusted, otherwise it is not adjusted. This targeted strategy can not only efficiently identify the situations that really need optimization and avoid unnecessary model adjustments, but also effectively improve the speech recognition model's ability to process complex text by optimizing the model precision at critical times, improve the accuracy and stability of human-computer interaction, and greatly enhance the interactive experience. Attached Figure Description

[0063] Figure 1 This is a flowchart illustrating the user tag acquisition method for human-computer interaction according to an embodiment of the present invention.

[0064] Figure 2 This is a process diagram illustrating the identification of the user's identity information status based on the voiceprint features according to an embodiment of the present invention.

[0065] Figure 3 A flowchart for setting the model accuracy of the machine speech recognition model in an embodiment of the present invention;

[0066] Figure 4 This is a flowchart illustrating the process of adjusting model accuracy in an embodiment of the present invention. Detailed Implementation

[0067] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0068] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0069] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0070] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0071] Please see Figure 1 The diagram shown illustrates the process of a user tag acquisition method for human-computer interaction according to an embodiment of the present invention. This embodiment provides a user tag acquisition method for human-computer interaction, including:

[0072] Step S1: In response to the wake-up voice, extract the voiceprint features of the wake-up voice; it is understood that the wake-up voice is usually a fixed, short phrase related to the robot's brand, such as: Hello brand name, Hello brand name, etc. This is existing technology and will not be described in detail here.

[0073] Step S2, based on the voiceprint features, identifies the identity information status of the interactive user and determines the model accuracy of the machine speech recognition model, including,

[0074] In response to users with unknown identities, the accuracy of the machine speech recognition model is set using the user's voice test samples and facial contour samples. It is understood that if it is a new user (i.e., an user with unknown identities), the user's information needs to be input to generate its corresponding user label and then determine the user's voice interaction model.

[0075] In response to an interactive user with a known identity, the machine speech recognition model is adjusted to the corresponding model accuracy based on the user's speech recognition status. It can be understood that if it is an old user (i.e., an interactive user with a known identity), there is no need to input the user's information. The machine stores the user's user tag internally, and the voice interaction of the user can be adjusted according to the user tag.

[0076] Step S3: Receive voice interaction information from the user and convert the voice interaction information into text interaction information according to the voice recognition model;

[0077] Step S4: Determine whether the text interaction information contains fuzzy fields based on the conversion time and the perplexity of the text interaction information.

[0078] Step S5 involves adjusting the model accuracy based on the determination result of the presence of fuzzy fields, combined with the position and proportion of the fuzzy fields. It can be understood that steps S3 to S5 take into account the user's specific situation on that day and adjust the model accuracy in real time according to the user's current speech. In practice, interactive users may improve their dialect and accent in a Mandarin-speaking environment, but the possibility of their accent and dialect will increase in an environment where everyone around them speaks with an accent. Therefore, these three steps can avoid the problem of low voice interaction efficiency when this situation occurs.

[0079] Specifically, in step S2, the process of adjusting the machine speech recognition model to the corresponding model accuracy based on the user's speech recognition state in response to an interactive user with an unknown identity includes:

[0080] Step S21: In response to an interactive user with an unknown identity, obtain the voice test sample and facial contour sample of the interactive user.

[0081] Step S22: Determine the geographic affiliation of the corresponding interactive user's voice-image association based on the voice test sample and the facial contour sample to determine its voice recognition trend;

[0082] Step S23: Determine the speech recognition state based on the speech test samples and the speech recognition trend to set the model accuracy of the machine speech recognition model.

[0083] Please see Figure 2 The diagram illustrates the process of identifying the identity information status of an interactive user based on the voiceprint features according to an embodiment of the present invention. Specifically, in step S1, the process of identifying the identity information status of an interactive user based on the voiceprint features includes...

[0084] Step S11: In response to the wake-up voice, the machine's identity information database and identity comparison function are invoked. It can be understood that the interactive machine has a pre-stored identity information database, which includes the identity tags of all interactive users with a consistent identity status. Each identity tag forms an identity tag folder, which includes an identity tag feature (i.e., the voiceprint feature of the wake-up voice) and the user's voice geographic affiliation, facial geographic affiliation, accent representation trend, speech recognition trend, speech recognition status, model accuracy, and the time for adjusting the model accuracy, i.e., the result of adjusting the model.

[0085] Step S12: Determine the similarity between the voiceprint features of the wake-up voice and the features of each user tag in the identity information database using voiceprint recognition technology. It is understood that each person's voiceprint features are unique, and by comparing the similarity of voiceprint features, it can be determined whether the interactive user who issued the wake-up voice is in a known identity state. In practice, voiceprint recognition technology is an existing technology and will not be described in detail.

[0086] Step S13: Determine the identity information status of the interacting user based on the aforementioned similarities, wherein...

[0087] If any of the aforementioned similarities is greater than or equal to a preset similarity, then the identity information status of the interacting user is determined to be a known identity status.

[0088] If all the aforementioned similarities are less than the preset similarity, then the identity information status of the interacting user is determined to be unknown.

[0089] The user tag feature is the voiceprint feature of the wake-up voice of the corresponding interactive user;

[0090] In practice, the higher the preset similarity setting, the higher the recognition accuracy; when the preset similarity is set above 90%, the recognition accuracy is at a high level; within this range, it can more accurately identify voiceprints that highly match the features in the identity information database, greatly reducing the possibility of misjudgment and improving the accuracy and reliability of recognition.

[0091] Specifically, in step S22, the process of determining the geographic location of the corresponding interactive user's voice-image association based on the voice test samples and the facial contour samples to determine its voice recognition trend includes,

[0092] Step S221: Determine the regional attribution and accent representation trend of the speech based on the speech test samples; in practice, inputting the speech test samples into a language model and machine learning algorithm can determine the regional attribution and accent representation trend of the speech.

[0093] Understandably, language models and machine learning algorithms include: (1) Hidden Markov Model (HMM), which can model the state transition and observation probability of speech, and identify speech patterns with dialect characteristics by analyzing the probability distribution of speech feature sequences; (2) Trained deep neural networks (DNN), such as convolutional neural networks (CNN) and recurrent neural networks (RNN) and their variants Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), etc., which can automatically learn deep features in speech, model and classify the complex features of different accents and dialects; through a large amount of labeled speech data training, the model can learn the speech characteristics and rules of different regional dialects, thereby accurately identifying and judging; (3) Support Vector Machine (SVM), which can take the acoustic features and linguistic features of speech as input, and classify different accents and dialects by constructing a hyperplane; during the training process, SVM will find the feature boundary that best distinguishes different categories of dialects, thereby realizing the dialect classification of unknown speech samples;

[0094] Step S222: Determine the facial geographic affiliation based on the facial contour sample. In practice, key facial feature points, such as the position, shape, and relative distance of the eyes, nose, mouth, and cheekbones, are extracted using facial recognition technology to construct a facial feature model. People from different regions may have some statistical differences in facial morphology. The facial recognition system can infer the possible geographic affiliation of an individual by analyzing and comparing these features. Then, the extracted key facial feature points are compared with an anthropological facial feature database to determine the possible geographic affiliation again. The final facial geographic affiliation is determined based on the two possible geographic affiliations.

[0095] It is understandable that facial recognition technology is an existing technology and will not be elaborated further. In addition, through long-term research, anthropologists have established databases containing facial features of different regions and ethnic groups (i.e., anthropological facial feature databases). These databases record a large amount of facial shape, skin color, hair and other feature information, and are associated with information such as region and ethnicity. By comparing the user's facial contour features with the data in the database, the possible regional affiliation can be inferred by referring to the information in the database.

[0096] Step S223: Determine the voice-image associated regional affiliation of the corresponding interactive user based on the voice regional affiliation and the facial regional affiliation; it is understood that both the voice regional affiliation and the facial regional affiliation may have more than one possible regional affiliation, and the same regional affiliation is selected as the voice-image associated regional affiliation.

[0097] Step S224: Determine the speech recognition trend based on the geographical affiliation of the audio-visual association;

[0098] The accent representation trend includes a weak accent representation trend and a strong accent representation trend, and the speech recognition trend includes a strong speech recognition trend, a weak speech recognition trend, and a difficult speech recognition trend.

[0099] Specifically, in step S224, the process of determining the speech recognition trend based on the geographic affiliation of the audio-visual association includes,

[0100] Step S2241: Compare the location of the audio-visual association with the pre-stored local speech recognition database;

[0101] Step S2242: Determine the speech recognition trend based on the comparison results;

[0102] The local speech recognition database is a mapping relationship between regional affiliation and speech recognition trends.

[0103] In one implementation, 34 provincial-level administrative regions across the country were selected as the data collection areas, covering different dialect regions (including Beijing and Northeast China in the Northern dialect region, Shanghai and Zhejiang in the Wu dialect region, Guangdong in the Yue dialect region, and Fujian in the Min dialect region, etc.). Volunteers of different ages, genders, and occupations were recruited, and at least 1,000 speech test samples were collected from each region (the same as the speech test samples in step S221). Professional linguists, annotation teams, or AI were organized to meticulously annotate the collected speech samples (the annotation content included regional information, dialect type, speech characteristics such as pronunciation, intonation, and linking, as well as vocabulary and grammatical features). The mapping relationship between regional affiliation and speech recognition trend was determined based on the average value of machine recognition of these speech test samples, and the trends were divided into strong speech recognition trend, weak speech recognition trend, and difficult speech recognition trend.

[0104] Please see Figure 3 The diagram shows a flowchart of setting the model accuracy of a machine speech recognition model according to an embodiment of the present invention. Specifically, in step S23, the speech recognition state is determined based on the speech test samples and the speech recognition trend to set the model accuracy of the machine speech recognition model;

[0105] Step S231: Obtain the accent representation trend corresponding to the speech test sample;

[0106] Step S232: Determine whether to correct the speech recognition trend based on the accent representation trend to determine the speech recognition state, wherein,

[0107] If the accent representation trend is a weak accent representation trend, then the speech recognition trend is corrected and the speech recognition state is determined based on the corrected speech recognition trend. It is understood that a weak accent representation trend indicates that the user's accent is not obvious and is closer to standard Mandarin, making it easier for the machine to recognize; therefore, the speech recognition state needs to be corrected. That is, when the speech recognition trend is a strong speech recognition model or a difficult-to-recognize model, the speech recognition state is a strong speech recognition state; when the speech recognition trend is a weak speech recognition model, the speech recognition state is a difficult-to-recognize speech recognition state. In other words, the corrected speech recognition state is stronger than the speech recognition trend.

[0108] If the accent representation trend is a strong accent representation trend, then it is determined that the speech recognition trend will not be corrected and the speech recognition state will be determined based on the speech recognition trend. It can be understood that in the case of a strong accent representation trend, the speech recognition state corresponds to the speech recognition trend, that is: if the speech recognition trend is a strong speech recognition model, then the speech recognition state is a strong speech recognition state; if the speech recognition trend is a weak speech recognition model, then the speech recognition state is a weak speech recognition state; if the speech recognition trend is a difficult speech recognition model, then the speech recognition state is a difficult speech recognition state.

[0109] Step S233: Set the model accuracy of the machine speech recognition model according to the speech recognition status;

[0110] The speech recognition states include strong speech recognition state, weak speech recognition state, and difficult speech recognition state.

[0111] Specifically, in step S233, the model accuracy of the machine speech recognition model is set according to the speech recognition status, including:

[0112] If the speech recognition state is a strong speech recognition state, it means that the voice of the interactive user is easy to recognize, therefore the model accuracy is a low model accuracy.

[0113] If the speech recognition status is "difficult to recognize speech", it means that the speech of the interactive user is not easy to recognize, so the model accuracy is medium.

[0114] If the speech recognition status is weak, it means that the voice of the user is difficult to recognize, therefore the model accuracy is high.

[0115] Please see Figure 4 The diagram shows a flowchart of adjusting model accuracy according to an embodiment of the present invention. Specifically, in step S4, the process of determining whether the text interaction information contains fuzzy fields and their locations based on the conversion time and the perplexity of the text interaction information includes...

[0116] Step S41: Obtain the conversion time of the text interaction information; it is understood that text interaction conversion usually begins when the voice is received, and the conversion time is determined based on the time difference between the start and end of the conversion.

[0117] Step S42: Determine the perplexity of the text interaction information based on the language model. In practice, language models include BERT, GPT, etc., which can calculate the probability score or perplexity of the text. It can be understood that the higher the score or the lower the perplexity, the more the text conforms to the language pattern learned by the language model, and the more fluent and logical it may be, making it easier for AI / interactive machines to understand.

[0118] Step S43: Determine whether there are ambiguous fields in the corresponding text interaction information based on the conversion time and the confusion level.

[0119] Specifically, in step S43, determining whether the corresponding text interaction information contains ambiguous fields based on the conversion time and the confusion level includes,

[0120] If the conversion time is greater than the preset time and the confusion level is less than the preset confusion level, then it is determined that the corresponding text interaction information contains a fuzzy field.

[0121] If the conversion time is less than or equal to the preset time, and / or the confusion level is greater than or equal to the preset confusion level, then it is determined that the corresponding text interaction information does not contain any ambiguous fields.

[0122] In practice, the preset duration is usually 1.2 to 1.5 times the voice input duration, with 1.3a being the preferred setting. That is, if the duration of the user's voice input is 'a', then the corresponding preset duration is 1.3a. If the conversion duration is greater than 1.3a, it is determined that the conversion duration is greater than the preset duration. It can be understood that the shorter the conversion time (i.e., the closer it is to the voice input duration), the faster the machine recognizes the voice, which means that the voice recognition state is stronger.

[0123] In practice, perplexity is an important indicator for measuring the performance of a language model. Theoretically, its value range is [1, +∞), but in practical applications, it usually falls within a relatively narrow range, generally greater than or equal to 1 and usually not exceeding a few hundred: (1) In an ideal situation, when the language model can perfectly predict all test data, that is, when the prediction probability of each word is 1, the perplexity is 1. This is the theoretical optimal situation, indicating that the model completely conforms to the language pattern of the training data, and the understanding and generation of text has reached the ultimate accuracy; (2) In reality, due to the complexity and diversity of language, as well as the limitations of the model itself, it is difficult for the perplexity to reach 1; Generally speaking, a well-trained language model may have a perplexity between tens and hundreds on a suitable dataset; for example, in some common natural language processing tasks, such as text classification and machine translation, the perplexity of a well-performing model may be between 20 and 200; if the perplexity of the model is too high, it usually means that the model's performance is poor and it may not be able to capture the rules and logic of language well, and there will be a large deviation in the understanding and generation of text;

[0124] In practice, the perplexity threshold varies depending on the specific task, dataset, and model. Generally, it may be around 20 to 50: (1) In some relatively simple tasks with specific domains and fixed language patterns, such as intelligent customer service scenarios in certain specific domains, the perplexity threshold may be set to around 20 to 30 after the model has been fully trained; (2) For general natural language processing tasks, such as open-domain dialogue systems or text generation tasks, the perplexity threshold may be appropriately relaxed to 30 to 50 due to the high diversity and complexity of languages. Taking a general intelligent chatbot as an example, it needs to handle dialogue content of various topics and styles. When the perplexity is below 50, it can often be considered that the generated voice text is acceptable in terms of logic and fluency, and the AI ​​can extract effective information and make reasonable responses. Therefore, in practice, the preset perplexity is preferably set to 50.

[0125] In practice, the probability score of text interaction information can also be determined: if the probability score is higher than or equal to the preset score, it is determined that there is no fuzzy field; if the probability score is lower than the preset score, it is determined that there is a fuzzy field.

[0126] Specifically, in step S5, the process of adjusting the corresponding model accuracy based on the determination result of the existence of fuzzy fields, combined with the position and proportion of the fuzzy fields, includes the following:

[0127] Step S51: Determine the conversion time of each character in the text interaction information based on the determination result of the existence of fuzzy fields;

[0128] Step S52: Determine the fuzzy position and fuzzy ratio of the fuzzy field based on the conversion time of each character. In practice, if the conversion time of a single character is greater than the preset conversion time, the character is determined to be a fuzzy character, and adjacent fuzzy characters are linked together as a fuzzy field. In practice, the preset conversion time = voice input time ÷ number of characters in the text interaction information × b, where b is determined based on the preset time. If the preset time is 1.3a, then b = 1.3.

[0129] Step S53: Adjust the corresponding model precision according to the fuzzy position and fuzzy ratio of the fuzzy field, wherein...

[0130] If the fuzzy position is in the middle position and the fuzzy ratio is greater than the preset ratio, then it is determined to adjust the model accuracy; in practice, the model accuracy is increased by one level, such as: (1) adjusting the low model accuracy to the medium model accuracy; (2) adjusting the medium model accuracy to the high model accuracy; (3) maintaining the high model accuracy and calling other computing resources to perform speech recognition for human-computer interaction;

[0131] If the blurred position is at the edge and / or the blurred ratio is less than the preset ratio, then it is determined that the model accuracy will not be adjusted.

[0132] In implementation, the fuzzy ratio = the sum of the number of characters in the fuzzy field ÷ the number of characters in the text interaction information, with a preset ratio ∈ [0.3, 0.5]. The edge position is the first and last 10% of the characters in the text interaction information. In an implementation, the text interaction information has a total of 50 characters, so the positions of the first five characters and the last five characters of the text interaction information are the edge positions, and the remaining positions are the middle positions.

[0133] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0134] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A user tag acquisition method for human-computer interaction, characterized in that, The method comprises: extracting a voiceprint feature of the wake-up voice in response to the wake-up voice; identifying an identity information state of an interactive user and determining a model precision of a machine voice recognition model based on the voiceprint feature, wherein the process of identifying the identity information state of the interactive user based on the voiceprint feature comprises calling a machine identity information library and an identity comparison function in response to the wake-up voice; determining the similarity between the voiceprint feature of the wake-up voice and each user label feature in the identity information library; determining the identity information state of the interactive user according to each similarity, wherein if any of the similarities is greater than or equal to a preset similarity, it is determined that the identity information state of the interactive user is a known identity state; if each of the similarities is less than the preset similarity, it is determined that the identity information state of the interactive user is an unknown identity state, wherein the user label feature is the voiceprint feature of the wake-up voice of the corresponding interactive user; in response to the interactive user in the unknown identity state, obtaining a voice test sample and a face contour sample of the interactive user, determining a voice-image correlation regional attribution of the corresponding interactive user to determine a voice recognition trend according to the voice test sample and the face contour sample, and determining a voice recognition state according to the voice test sample and the voice recognition trend to set the model precision of the machine voice recognition model; in response to the interactive user in the known identity state, adjusting the machine voice recognition model to a corresponding model precision according to the voice recognition state of the user; receiving voice interaction information of an interactive user, and converting the voice interaction information into text interaction information according to the voice recognition model; determining whether there is an ambiguous field in the text interaction information according to the conversion time length of the text interaction information and the ambiguity degree of the text interaction information; adjusting the corresponding model precision according to the determination result of the existence of the ambiguous field, the position of the ambiguous field and the proportion of the ambiguous field.

2. The user tag acquisition method for human-computer interaction according to claim 1, characterized in that, The process of determining the voice-image correlation regional attribution of the corresponding interactive user to determine the voice recognition trend according to the voice test sample and the face contour sample comprises determining a voice regional attribution and a pronunciation representation trend according to the voice test sample; determining a face regional attribution according to the face contour sample; determining the voice-image correlation regional attribution of the corresponding interactive user according to the voice regional attribution and the face regional attribution; determining the voice recognition trend according to the voice-image correlation regional attribution; wherein the pronunciation representation trend comprises a weak pronunciation representation trend and a strong pronunciation representation trend, and the voice recognition trend comprises a strong voice recognition trend, a weak voice recognition trend and a not easy voice recognition trend.

3. The user tag acquisition method for human-computer interaction according to claim 2, characterized in that, The process of determining the voice recognition trend according to the voice-image correlation regional attribution comprises comparing the voice-image correlation regional attribution with a pre-stored regional voice recognition library; determining the voice recognition trend according to the comparison result; wherein the regional voice recognition library is a mapping relationship between each regional attribution and a voice recognition trend.

4. The user tag acquisition method for human-computer interaction according to claim 3, characterized in that, determining a voice recognition state according to the voice test sample and the voice recognition trend to set the model precision of the machine voice recognition model; obtaining a pronunciation representation trend corresponding to the voice test sample; According to the accent characteristic trend, whether to correct the speech recognition trend is determined to determine a speech recognition state, wherein, if the accent characteristic trend is a weak accent characteristic trend, it is determined to correct the speech recognition trend and determine the speech recognition state according to the corrected speech recognition trend; if the accent characteristic trend is a strong accent characteristic trend, it is determined not to correct the speech recognition trend and determine the speech recognition state according to the speech recognition trend; According to the speech recognition state, the model accuracy of the machine speech recognition model is set; wherein the speech recognition state includes a strong speech recognition state, a weak speech recognition state and a non-easy speech recognition state.

5. The user tag acquisition method for human-computer interaction according to claim 4, characterized in that, According to the speech recognition state, the model accuracy of the machine speech recognition model is set, including, if the speech recognition state is a strong speech recognition state, the model accuracy is a low model accuracy; if the speech recognition state is a non-easy speech recognition state, the model accuracy is a medium model accuracy; if the speech recognition state is a weak speech recognition state, the model accuracy is a high model accuracy.

6. The user tag acquisition method for human-computer interaction according to claim 1, wherein, According to the conversion time length of the text interaction information and the perplexity of the text interaction information, the process of determining whether there is a fuzzy field in the text interaction information and its position includes, obtaining the conversion time length of the text interaction information; determining the perplexity of the text interaction information according to the language model; determining whether the corresponding text interaction information has a fuzzy field according to the conversion time length and the perplexity.

7. The user tag acquisition method for human-computer interaction according to claim 6, characterized in that, According to the conversion time length and the perplexity, whether the corresponding text interaction information has a fuzzy field, including, if the conversion time length is greater than a preset time length and the perplexity is less than a preset perplexity, it is determined that the corresponding text interaction information has a fuzzy field; if the conversion time length is less than or equal to a preset time length, and / or, the perplexity is greater than or equal to a preset perplexity, it is determined that the corresponding text interaction information does not have a fuzzy field.

8. The user tag acquisition method for human-computer interaction according to claim 1, wherein, According to the determination result of the existence of the fuzzy field, the process of adjusting the corresponding model accuracy combined with the position of the fuzzy field and the proportion of the fuzzy field includes, determining the conversion time length of each character in the text interaction information based on the determination result of the existence of the fuzzy field; determining the fuzzy position of the fuzzy field and the fuzzy proportion of the fuzzy field according to the conversion time length of each character; According to the fuzzy position of the fuzzy field and the fuzzy proportion of the fuzzy field, the corresponding model accuracy is adjusted, wherein, if the fuzzy position is in the middle position and the fuzzy proportion is greater than a preset proportion, it is determined to adjust the model accuracy; if the fuzzy position is in the edge position and / or the fuzzy proportion is less than a preset proportion, it is determined not to adjust the model accuracy.

Citation Information

Patent Citations

  • A method, system, terminal, and storage medium for constructing user profiles based on human-computer dialogue.

    CN111914076B

  • Training method and device for speech recognition model

    CN109119071A

  • Method and apparatus for using self-service and an electronic device

    CN109343817A