Information processing method and equipment for man-machine practice, and medium
Through real-time voice recognition and emotion recognition, combined with quality inspection rules and knowledge graphs, personalized response problems are generated, and the problem of insufficient accuracy and comprehensiveness of voice recognition in the existing technology is solved, real-time feedback and rich scenes of the training process are achieved, and users' practical capabilities are improved.
Patent Information
- Application Number
- CN202510340451.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
The existing speech recognition technology has limitations in terms of accuracy and comprehensiveness, and it is impossible to achieve effective real-time feedback and evaluation, and it is difficult to generate corresponding response problems for students' personalized replies, making the training session a single scenario and making it difficult to improve practical capabilities.
By obtaining user interactive voice in real time, performing speech recognition, emotion, mute and speech speed recognition, combining quality inspection rules to evaluate the training process in real time, and generating personalized responses to problems based on key semantic information, using knowledge graphs iteratively to update the associated questioning paths, realizing multiple rounds of dialogue and scene simulation.
Real-time feedback and evaluation of the training process is achieved, skill weaknesses are captured, training session scenarios are enriched, and users' practical ability and the accuracy of evaluation results are improved.
Smart Images

Figure CN120183397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to an information processing method, device and medium for human-computer training. Background Art
[0002] At present, the training system mainly relies on the "listening and watching" of the traditional manual training model, which is inefficient, lacks actual combat scenarios, and has slow skill improvement. Secondly, although the existing voice recognition and keywords can help trainees complete basic modular training and scoring, they often have limitations in accuracy and comprehensiveness. With the rapid development of artificial intelligence technology, human-computer dialogue systems are increasingly used in various fields, especially in education and training, customer service and other scenarios. In the field of sales training, human-computer training can effectively improve the professional skills and adaptability of sales personnel, creating greater value for the company.
[0003] At present, human-computer interaction learning methods based on speech recognition have made certain progress. For example, CN112365892A proposes a human-computer dialogue method, which receives the user's current round of dialogue voice and pre-processes the dialogue voice to obtain text information; processes the text information through a preset semantic analysis model to obtain intent information; obtains historical response information, and determines the current round of dialogue status based on the historical response information and intent information; configures the response information corresponding to the dialogue status according to the preset response configuration model, and generates a response voice corresponding to the response information. This method improves the dialogue efficiency and dialogue effect to a certain extent. However, although this technology takes into account the dialogue status and intent information, it lacks personalized practical scenario simulation for students, and it is difficult to generate corresponding response questions for the personalized replies of students, resulting in a single scene in the practice session and difficulty in improving practical ability.
[0004] In addition, although the existing intelligent question-and-answer system can conduct basic question-and-answer interactions, it has limitations in judging and scoring the content of the responses, and is unable to fully tell the interlocutor whether what is said during the conversation complies with standard process specifications or whether any problems arise.
[0005] Secondly, existing speech recognition technology has limitations in terms of accuracy and comprehensiveness, especially in terms of emotion, speech speed, etc., which leads to increased scoring errors without contextual semantic understanding. Although CN116049360A mentions emotion recognition, it is performed to determine feedback words, not for evaluation.
[0006] Third, existing technology integration solutions lack the real-time feedback and evaluation mechanisms provided by cutting-edge technologies, which prevents users from quickly correcting errors during interactions.
[0007] Therefore, there is an urgent need for a human-machine practice information processing method that can score in real time according to the user interaction process and provide personalized question responses, so as to improve the efficiency and effect of training. Summary of the Invention
[0008] The purpose of the present invention is to provide an information processing method, device and medium for human-machine practice, so as to solve the limitations of existing speech recognition technologies in terms of accuracy and comprehensiveness, which cannot achieve effective real-time feedback and evaluation, and it is difficult to generate corresponding response questions for the personalized responses of trainees, resulting in a single scenario in the practice session and difficulty in improving actual combat capabilities.
[0009] The purpose of the present invention can be achieved through the following technical solutions:
[0010] According to the first aspect of the present invention, an information processing method for human-machine practice is provided, and the method includes the following steps:
[0011] S1, obtaining the interactive voice of the user's response to the question in real time;
[0012] S2, performing real-time speech recognition on the interactive voice and converting it into text information with time stamps;
[0013] S3, preprocessing the text information and extracting key semantic information, and performing emotion recognition, silence recognition and speech rate recognition;
[0014] S4, based on the extracted key semantic information, emotion recognition, silence recognition and speech rate recognition results, evaluating the practice process in real time according to the quality inspection rules, and judging whether it is necessary to interrupt the interaction according to the evaluation results. If interruption is required, immediately execute step S5; otherwise, execute step S5 after completing the complete round of interactive speech recognition and processing;
[0015] S5, performing intention recognition based on the extracted key semantic information, mapping the recognized intention to the corresponding nodes of the knowledge graph, iteratively updating the associated follow-up question path according to the knowledge graph, generating response questions in combination with scene labels and process rules, and sending them to the user;
[0016] S6, obtaining the interactive voice of the user's response to the response question, and returning to step S2 for the next round of human-machine practice interaction.
[0017] As a preferred technical solution, the iterative update of the associated follow-up question path according to the knowledge graph and the generation of response questions in combination with scene labels and process rules are specifically:
[0018] Enumerating keywords through part-of-speech tagging, synonym expansion and industry-related words, and constructing a keyword knowledge base under different scene labels and process rules;
[0019] Perform keyword recognition based on the keyword knowledge base and the extracted key semantic information;
[0020] Based on the recognized intent, search for matching nodes and relationship paths in the knowledge graph. Based on the recognized relationship paths and the user's historical interaction records, select the optimal follow-up path and determine the keywords corresponding to the optimal follow-up path;
[0021] Combine the keywords obtained from keyword recognition and the keywords obtained according to the optimal path, match the corresponding scenario tags and process rules, and extract the question with the highest relevance to the keywords in this round from the pre-configured Q&A library of the scenario tags and process rules as the response question.
[0022] As a preferred technical solution, the method for extracting the question with the highest relevance to the keywords in this round from the pre-configured Q&A library of the scenario tags and process rules is as follows:
[0023] Record the question set in the Q&A library as Q = {q1, q2,... q K}, where K is the number of questions. For each question q k , calculate its relevance R rel with the interactive voice in this round:
[0024] R rel (q k ) = αK(q k ) + βS(q k , C) + γH(q k )
[0025] Where α, β, and γ are weights, K(q k ) represents the normalized matching score between question q k and all keywords, S(q k ) represents the similarity between question q k and the extracted semantic information C, and H(q k ) represents the probability that question q k has been mentioned historically under the scenario tags and process rules.
[0026] As a preferred technical solution, the quality inspection rules include rule name, category, corresponding score, and violation conditions, where the violation conditions are constructed by selecting at least one of semantic tags, keywords, and regular expressions and combining with a preset rule threshold.
[0027] As a preferred technical solution, the semantic tags include user semantic tags and machine semantic tags. The user semantic tags are key business nodes involved in the interaction process of the user. The key business nodes are set according to scenario tags and process rules, and multiple non-default global rules are bound to each key business node. The machine semantic tags are used to configure non-default global rules involved in the interaction process of the machine and / or generate guiding rules for answering questions.
[0028] As a preferred technical solution, in step S4, the tolerance degree of semantic deviation for different scenarios is determined according to the scenario tags, and the threshold of the quality inspection rule is adjusted based on the tolerance degree to score the user's practice process.
[0029] As a preferred technical solution, the emotion recognition, mute recognition, and speech rate recognition are specifically as follows:
[0030] Based on the text information with timestamps, the number of words per minute is counted for speech rate recognition;
[0031] Based on the text information with timestamps, the length of the time interval without valid words is counted for mute recognition;
[0032] According to the grammar, semantic structure, and context information of the text information, combined with the results of speech rate recognition, emotion recognition is performed.
[0033] As a preferred technical solution, the first-level evaluation dimensions for real-time evaluation of the practice process according to the quality inspection rules include semantic ability, speech rate control ability, adaptability, accuracy rate, proficiency, and scenario integrity. Among them, the second-level evaluation dimensions of the semantic ability include term accuracy, process compliance, and logical rigor. The speech rate control ability is scored according to the deviation value of the speech rate recognition result from the preset speech rate reference value. The second-level evaluation dimensions of the adaptability include response timeliness, solution innovation, and anti-interference ability. The second-level evaluation dimensions of the accuracy rate include data accuracy, compliance, and user information accuracy. The second-level evaluation dimensions of the proficiency include task completion time and error recurrence rate. The second-level evaluation dimensions of the scenario integrity include process coverage rate and branch processing integrity. Corresponding scoring / deducting scores are set according to the preset quality inspection rules in each second-level evaluation dimension.
[0034] According to the second aspect of the present invention, an electronic device is provided, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.
[0035] According to the third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) Through semantic understanding and in-depth voice quality inspection technology, the present invention uses a knowledge graph and multi-round conversations to simulate intelligent real scenarios, enabling diverse scenario drills and clearances for trainees; during the drill process, the quality inspection results are provided in real time, and weak skill points are captured, enabling users to enter the actual combat state at any time, understand their own shortfalls in conversation skills, and effectively solve the problems of lack of actual combat scenarios and real-time feedback in traditional training modes.
[0038] (2) The present invention iteratively updates the associated follow-up question paths according to the knowledge graph, and can generate response questions personalized according to the user's replies and historical information, combined with scenario tags and process rules, enriching the scenarios in the drill session and enhancing the user's actual combat ability.
[0039] (3) The present invention provides a more comprehensive ability assessment through multi-dimensional evaluation indicators. Combining quality inspection rules, it conducts quality inspections on processes, emotions, speech rates, and semantics, improving the accuracy of the evaluation results, expanding the evaluation dimensions, enabling users to recognize their own shortfalls from various dimensions, and enhancing the drill effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the method of the present invention;
[0041] Figure 2 is a schematic diagram of sensitive word configuration in an embodiment;
[0042] Figure 3 is a schematic diagram of the implementation logic of the present invention in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0044] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connect", "be connected", "couple" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0045] Embodiment 1
[0046] This embodiment provides an information processing method for human-computer sparring, as Figure 1 shown, the method includes the following steps:
[0047] S1. Real-time obtain the interactive voice of the user's reply to the question.
[0048] In this step, the voice signal of the user during the sparring process is collected in real time through a microphone or other audio input device. In this embodiment, the voice signal is digitally processed with a sampling rate of 16 kHz and a quantization accuracy of 16 bits to ensure the quality of the voice signal. At the same time, the collected voice signal can be denoised to improve the accuracy of subsequent speech recognition.
[0049] S2. Perform real-time speech recognition on the interactive voice and convert it into text information with time stamps.
[0050] In this step, a deep learning model is used to perform real-time recognition on the collected speech signals. Specifically, the system first preprocesses the speech signals, including operations such as framing, windowing, and feature extraction. The length of each frame is 25 ms, the frame shift is 10 ms, and the extracted features include Mel Frequency Cepstral Coefficients (MFCC) and acoustic features. Then, a speech recognition model based on the Transformer architecture is used to convert the speech features into text information. This model contains 12 encoder layers and 6 decoder layers, with a hidden layer dimension of 512, a feed-forward network dimension of 2048, and 8 attention heads.
[0051] During the speech recognition process, timestamps are added to each recognized word to record the start time and end time of the word in the original speech, accurate to the millisecond level. For example, for the recognized sentence "Hello, I'm Xiao Wang, a salesperson", the following timestamped text information is generated:
[0052] {
[0053] "text":"Hello, I'm Xiao Wang, a salesperson",
[0054] "words":{"word":"Hello","start_time":0.0,"end_time":0.8},
[0055] {"word":",","start_time":0.8,"end_time":1.0},
[0056] {"word":"I'm","start_time":1.0,"end_time":1.5},
[0057] {"word":"salesperson","start_time":1.5,"end_time":2.0},
[0058] {"word":"Xiao Wang","start_time":2.0,"end_time":2.5}
[0059] }
[0060] S3, preprocess the text information and extract key semantic information, and perform emotion recognition, silence recognition, and speech rate recognition.
[0061] In this step, first, preprocess the recognized text information, including operations such as word segmentation, removing stop words and punctuation marks. Then, use a semantic understanding model based on BERT to extract the key semantic information in the text. This model adopts the method of pre-training plus fine-tuning. In the pre-training stage, a large-scale general corpus is used, and in the fine-tuning stage, a corpus of a specific domain is used. The input of the model is the preprocessed text, and the output is the semantic representation vector of the text and the key semantic information. The role of key semantic information extraction is to improve the accuracy of subsequent intent recognition. Key semantic information extraction can first configure the model, precipitate the rule base, and continuously verify and optimize the key semantic information extraction model through positive inspection, false inspection, missed inspection annotation, and weight value annotation. According to the semantic annotation of the original material of the question, determine the semantic weight value and the semantic deviation degree value. By quantifying the semantic similarity (for example, the similarity between "return to the original" and "cash value" is 35%), synonyms or colloquial expressions can be recognized, avoiding misjudgment in traditional key information matching.
[0062] In this embodiment, the speech rate recognition is specifically as follows: Based on the text information with timestamps, count the number of words per minute (Words Per Minute, WPM) for speech rate recognition. Specifically, calculate the total effective speech duration according to the timestamps of each word in the text, and then calculate the number of words per unit time. For example, if a 30-second speech contains 60 words, the speech rate is 120 WPM.
[0063] In this embodiment, the silence recognition is specifically as follows: Based on the text information with timestamps, count the length of the time interval without valid words for silence recognition. For example, if the silent time interval > 5 seconds or there are only filler words within 5 seconds, a prompt is triggered. Record the start time, end time, and duration of the silence for subsequent analysis.
[0064] In this embodiment, the emotion recognition is specifically as follows: According to the grammar, semantic structure, and context information of the text information, combined with the result of speech rate recognition, perform emotion recognition. In one embodiment, a multi-modal emotion recognition model can be used, which fuses text features and speech features. The text features are extracted by the BERT model, and the speech features include fundamental frequency, energy, and spectral features, etc. The model outputs the emotion category (such as positive, negative, neutral) and the emotion intensity.
[0065] S4. Based on the extracted key semantic information, emotion recognition, silence recognition, and speech rate recognition results, evaluate the pair practice process in real time according to the quality inspection rules, and judge whether it is necessary to interrupt the interaction according to the evaluation result. If it is necessary to interrupt, immediately execute step S5; otherwise, execute step S5 after completing the full-round interactive speech recognition and processing.
[0066] In this step, the user's practice process is evaluated in real time according to preset quality inspection rules. The quality inspection rules include rule name, category, corresponding score, and violation conditions. Among them, the violation conditions are constructed by selecting at least one of semantic tags, keywords, and regular expressions and combining with preset rule thresholds. For example, in the scenario of requirement mining, a communication etiquette standard detection rule is formulated, classified as service standard category, minor violation (deduct 5 points), violation condition: (number of interruptions > 2 times / minute) OR (mute duration > 5 seconds and no buffer words are used).
[0067] Semantic tags include user semantic tags and machine semantic tags.
[0068] User semantic tags are the key business nodes involved in the user's interaction process. The key business nodes are set according to scenario tags and process rules, and each key business node is bound with multiple non-default global rules. The main purposes of setting user semantic features are business orientation and rule triggering ability. For example, in the product promotion scenario, the label "product income promise" is defined. When words such as "guaranteed income" and "zero risk" are used during the user's practice process, points will be deducted, and the income knowledge points required by the supervision will be pushed for learning. After watching, the next process will be carried out for practice; in addition, a global rule prohibiting sensitive words is configured for this label, and points will be deducted when sensitive words are detected. In the risk disclosure scenario, the label "risk prompt confirmation" is defined. When the user does not actively confirm the product risk, points will be deducted and a prompt will be given, and a process blocking rule will be configured. The next link is not allowed to enter until it is confirmed. In addition, a global rule of "customer portrait verification" is configured for this label, and the frequency of risk prompt confirmation is increased for elderly customers.
[0069] Machine semantic tags are used to configure non-default global rules involved in the machine's interaction process and / or generate guiding rules for responding to questions. The method of configuring global rules can refer to the global rule configuration of the aforementioned user semantic tags, and this embodiment will not elaborate here. For the guiding rules for generating responses to questions, the following is an example: Define the semantic tag for compliance processing in wealth management sales. When this tag is detected in the conversation, the guiding tag for compliance processing in wealth management sales will be automatically matched, and this guiding tag is used to guide the generation of subsequent responses to questions.
[0070] Such as Figure 2 Figure 14 shows a sensitive word processing method in an embodiment. By constructing a knowledge base of sensitive words, shielding words, and derivative words, and associating and configuring process rules and quality inspection rules, the scoring and process management of sensitive words involved in the human-machine practice process are realized.
[0071] For example, in the scenario of compliance training for insurance sales scripts, the following quality inspection rules and corresponding score configurations are set:
[0072] 1. Compliance detection: The key semantic information involved includes "principal and interest guaranteed", "rigid payment", "absolutely safe", "the same as deposits", etc. The set weight is 40%, and the deduction rule is: for each detected relevant semantic information, 10 points will be deducted;
[0073] 2. Process accuracy: The key semantic information involved includes "health notification", "exemption clause description", "surrender loss description", etc. The set weight is 60%, and the deduction rule is: if the detected process sequence is inconsistent, 20 points will be deducted.
[0074] In one embodiment, the tolerance degree of semantic deviation for different scenarios is determined according to the scenario label, and the threshold of the quality inspection rule is adjusted based on the tolerance degree to score the user's practice process. For example, corresponding speech rate thresholds can be set for different practice scenarios. For example, the speech rate of the key paragraphs of product introduction should not be higher than 180 words per minute. If the speech rate exceeds 180 words per minute, the higher the deviation degree, the lower the score. In the risk confirmation scenario, the speech rate needs to be reduced to prompt the customer of relevant risks, so the speech rate threshold can be set to 150 words per minute. In the annuity insurance sales scenario, the semantic weight of "possible loss of principal" is set to 0.9 (high risk), while the weight of "historical return" is set to 0.3 (low risk), so as to distinguish the importance of different semantics.
[0075] In this embodiment, the first-level evaluation dimensions for real-time evaluation of the practice process according to the quality inspection rules include semantic ability, speech rate control ability, strain ability, accuracy rate, proficiency, and scenario integrity. The weight settings for each dimension are 25%, 10%, 20%, 25%, 10%, and 10% respectively, and are detailed as follows:
[0076] 1. Semantic ability
[0077] The secondary evaluation dimensions of semantic ability include:
[0078] (1) Term accuracy: Use industry-standard terms (such as R1-R5 risk levels). If used incorrectly, 2 points will be deducted each time. The corresponding score for this part is 0-10 points.
[0079] (2) Process compliance: It is necessary to completely trigger the preset semantic labels. If a certain process is not triggered, the corresponding score will be deducted. The corresponding score for this part is 0-8 points.
[0080] (3) Logical rigor: The answer needs to be non-contradictory and the logical chain needs to be complete. If a logical flaw is found, 3 points will be deducted for each case. The corresponding score for this part is 0-7 points.
[0081] 2. Speech rate control ability
[0082] The speech rate control ability is scored according to the deviation between the speech rate recognition result and the preset speech rate reference value. In this embodiment, the speech rate reference value is set to 180WPM±10%, with a full score of 10 points. 1 point will be deducted for every 5WPM exceeding it, and 1 point will be deducted for every 10WPM below it. The lower limit is 0 points.
[0083] 3. Resilience
[0084] The secondary evaluation dimensions of resilience include:
[0085] (1) Response time: 8 points if the response time is ≤ 30 seconds. 2 points will be deducted for every 5 seconds of delay. The corresponding score for this part is 0-8 points.
[0086] (2) Innovation of solution: 6 points will be awarded for proposing more than two solutions, and 3 points will be awarded for only following the process. The corresponding score for this part is 0-6 points.
[0087] (3) Anti-interference ability: 6 points will be awarded for effectively resolving the machine's intentionally misleading questions, and 3 points will be awarded for requiring guidance. The corresponding score for this part is 0-6 points.
[0088] 4. Accuracy
[0089] The secondary evaluation dimensions of accuracy include:
[0090] (1) Data accuracy: If there is ≤1 error in the product yield rate / term validity period, etc., 12 points will be awarded. If this condition is not met, the corresponding points will be deducted. The corresponding score for this part is 0-12 points.
[0091] (2) Compliance: 5 points will be deducted for each use of prohibited words (such as "guaranteed principal and interest"), and the corresponding score for this part is 0-5 points.
[0092] (3) User information accuracy: If the customer profile information is repeated incorrectly ≤ 1 time, 5 points will be awarded. If it is not met, the corresponding points will be deducted. The corresponding score for this part is 0-5 points.
[0093] 5. Proficiency
[0094] The secondary assessment dimensions of proficiency include:
[0095] (1) Task completion time: If the standard process time is ≤ 80% of the benchmark time, 6 points will be awarded. If it is not met, the corresponding points will be deducted. The corresponding score for this part is 0-6 points.
[0096] (2) Error recurrence rate: 4 points will be awarded if the same type of error is repeated ≤ 1 time. If this condition is not met, the corresponding points will be deducted. The corresponding score for this part is 0-4 points.
[0097] 6. Scene completeness
[0098] The secondary evaluation dimensions of scene completeness include:
[0099] (1) Process coverage rate: Completing all preset nodes (such as the three stages of health notification) earns 5 points, and if not completed, the corresponding score will be deducted. The corresponding score for this part is 0 - 5 points.
[0100] (2) Completeness of branch processing: If the processing rate of the abnormal path (such as the customer rejecting the plan) reaches 100%, 5 points will be awarded, and if not completed, the corresponding score will be deducted. The corresponding score for this part is 0 - 5 points.
[0101] Based on the above scoring results, the ability grading as shown in Table 1 below is set.
[0102] Table 1 Ability Grading
[0103]
[0104] S5. Perform intent recognition based on the extracted key semantic information, map the recognized intent to the corresponding nodes in the knowledge graph, iteratively update the associated follow-up paths according to the knowledge graph, generate responses to questions in combination with scenario tags and process rules, and send them to the user.
[0105] In this step, first perform intent recognition based on the extracted key semantic information. The system uses an intent recognition model based on BERT. This model maps the user's answer to predefined intent categories through semantic understanding of the text. For example, for the user's answer "I want to know about your product functions", the intent of "product consultation" may be recognized.
[0106] Then, map the recognized intent to the corresponding nodes in the knowledge graph. The knowledge graph is a structured knowledge representation method that contains entities, attributes, and relationships. The knowledge graph contains key concepts, process nodes, and their relationships in various scenarios. For example, the intent of "product consultation" may be mapped to the "product" node in the knowledge graph, and this node has associated relationships with nodes such as "function", "price", and "usage method".
[0107] Next, according to the recognized intent, search for matching nodes and relationship paths in the knowledge graph, select the optimal follow-up path based on the recognized relationship path and the user's historical interaction records, and determine the keywords corresponding to the optimal follow-up path.
[0108] In addition, keywords can also be enumerated through part-of-speech tagging, synonym expansion, and industry-related words (such as "initial premium" associated with "payment term" and "rate floating") to build a keyword knowledge base under different scenario tags and process rules; then, perform keyword recognition according to the keyword knowledge base and the extracted key semantic information.
[0109] Combine the keywords obtained from keyword recognition and the keywords obtained according to the optimal path, match the corresponding scenario tags and process rules, and extract the question with the highest correlation with the keywords in this round from the Q&A library pre-configured with the scenario tags and process rules as the response question.
[0110] The form of a single-round response question can be: "What does the {keyword} you mentioned specifically refer to"; an example of multi-round follow-up questions is: after detecting "cash value", additional advanced questions such as calculation logic and influencing factors are added. The way to manage the follow-up path through knowledge graph iterative update can be: add an association between "health notification" and specific disease history description.
[0111] S6. Obtain the interactive voice of the user's response to the response question, and return to step S2 to perform the next round of human-machine practice interaction.
[0112] In this step, the interactive voice of the user's response to the response question is obtained through a microphone or other audio input device. Then, return to step S2 to perform real-time speech recognition on the new interactive voice, convert it into text information with timestamps, and perform subsequent processing to form a complete human-machine practice interaction loop.
[0113] As Figure 3 Shown is a schematic diagram of the actual application of the information processing method of this embodiment. The user can initiate human-machine practice through the APP, which is embedded with the code to implement the above information processing method, and realizes real-time human-machine practice and user ability evaluation through speech recognition, scenario dialogue, and ability scoring.
[0114] Embodiment 2
[0115] This embodiment provides an information processing method for human-machine practice. Based on Embodiment 1, this method elaborates on "According to the knowledge graph iterative update to associate the follow-up path, and generate a response question by combining the scenario tag and the process rule" in step S5. Specifically as follows:
[0116] In this step, first, the keywords are exhausted through part-of-speech tagging, synonym expansion, and industry-related words to construct a keyword knowledge base under different scenario tags and process rules.
[0117] For part-of-speech tagging, a part-of-speech tagging model based on BiLSTM-CRF is used. The input of this model is the text after word segmentation, and the output is the part-of-speech tag of each word. The hidden layer dimension of the model is 256. Bidirectional LSTM is used to capture context information, and the CRF layer is used for sequence annotation. This embodiment mainly focuses on parts of speech such as nouns, verbs, adjectives, and adverbs. Words of these parts of speech usually contain important semantic information.
[0118] For synonym expansion, a pre-trained word vector model is used to calculate the semantic similarity between words and find words that are semantically similar to the keyword. The word vector model used is trained based on the Word2Vec algorithm, with a word vector dimension of 300, and the training corpus contains text data in a specific domain. The similarity threshold is set to 0.8, that is, when the cosine similarity between two words is greater than 0.8, they are considered synonyms.
[0119] For industry-related words, a predefined industry dictionary and association rules are used to find words that are associated with the keyword in a specific industry. For example, in the financial industry, "interest rate" is associated with words such as "deposit", "loan", and "wealth management". The industry dictionary contains professional terms and common words in multiple industries, and the association rules are constructed based on industry knowledge and statistical analysis.
[0120] Through the above methods, a keyword knowledge base is constructed for each scenario label and process rule. For example, for the "product consultation" scenario, the keyword knowledge base may contain keywords such as "product", "function", "price", "usage method", etc., as well as their synonyms and related words.
[0121] Then, keyword recognition is performed based on the keyword knowledge base and the extracted key semantic information. In this embodiment, a keyword recognition model based on the attention mechanism is used. The input of this model is the semantic representation vector of the text, and the output is the keywords contained in the text and their importance scores. The dimension of the attention layer of the model is 128, and the multi-head attention mechanism is used to capture key information at different positions.
[0122] Next, according to the recognized intent, matching nodes and relationship paths are searched for in the knowledge graph. Each node in the knowledge graph represents a concept or entity, and the edges between the nodes represent the relationships between them. A graph traversal algorithm is used to search for nodes and relationship paths related to the recognized intent in the knowledge graph. For example, if the recognized intent is "product consultation", nodes and relationship paths related to "product" will be searched for in the knowledge graph.
[0123] Based on the recognized relationship paths and the user's historical interaction records, the optimal follow-up path is selected, and the keywords corresponding to the optimal follow-up path are determined. In one embodiment, a reinforcement learning algorithm can be used to optimize the selection of the follow-up path. The state space of this algorithm is the current dialogue state, the action space is the possible follow-up paths, and the reward function is designed based on the user feedback and the completion of the dialogue goal. The optimal follow-up strategy is learned through multiple rounds of interaction to maximize the long-term reward.
[0124] Finally, combine the keywords obtained from keyword recognition and the keywords obtained according to the optimal path, match the corresponding scenario tags and process rules, and extract the question with the highest correlation degree with the keywords in this round from the question and answer library pre-configured with the scenario tags and process rules as the response question.
[0125] In a preferred embodiment, denote the set of questions in the question and answer library as Q = {q1, q2,... q K}, where K is the number of questions. For each question q k , calculate its correlation degree R rel with the interactive voice in this round:
[0126] R rel (q k ) = αK(q k ) + βS(q k , C) + γH(q k )
[0127] Where α, β, and γ are weights. K(q k ) represents the normalized matching score between question q k and all keywords. S(q k ) represents the similarity between question q k and the extracted semantic information C. H(q k ) represents the probability that question q k has been mentioned historically under the scenario tags and process rules.
[0128] Specifically, the calculation formula for K(q k ) is:
[0129] K(q k ) = count(q k ∩ keywords) / count(keywords)
[0130] Where count(q k ∩ keywords) represents the number of keywords included in question q k , and count(keywords) represents the total number of keywords.
[0131] The calculation formula for S(q k , C) is:
[0132] S(q k , C) = cos(vec(q k ), vec(C))
[0133] Where vec(q k ) represents the vector of question q kThe semantic vector of vec(C) represents the semantic vector of the extracted semantic information C, and cos represents the cosine similarity function.
[0134] H(q k ) is calculated as follows:
[0135] H(q k ) = count(q k in history) / count(history)
[0136] Among them, count(q k in history) represents the number of times the question q k is mentioned in the historical interaction, and count(history) represents the total number of historical interactions.
[0137] Then, select the question with the highest degree of association as the response question and send it to the user.
[0138] Embodiment 3
[0139] This embodiment provides an information processing method for human-computer sparring. Based on Embodiment 1, the emotion recognition, mute recognition, and speech rate recognition in step S3 are described in detail as follows:
[0140] 1. Speech rate recognition
[0141] In this step, based on the timestamped text information, count the number of words per minute for speech rate recognition. Specifically, calculate the total effective speech duration according to the timestamps of each word in the text, and then calculate the number of words per unit time. The speech rate calculation formula used is:
[0142] WPM = (word_count / speech_duration)*60
[0143] Among them, word_count represents the number of words, speech_duration represents the effective speech duration (unit: second), and 60 is the coefficient for converting seconds to minutes.
[0144] For example, if a 30-second speech contains 60 words, the speech rate is:
[0145] WPM = (60 / 30)*60 = 120
[0146] Compare the calculated speech rate with the preset speech rate range to determine whether the user's speech rate is appropriate.
[0147] 2. Mute recognition
[0148] For silence recognition, based on the timestamped text information, the length of the time interval without valid words is statistically counted. The time interval between two consecutive valid words is compared with a preset threshold. If it exceeds the threshold, silence is considered to exist. For example, filler words such as "ah" and "um" are not regarded as valid words.
[0149] The following silence recognition algorithm can be used:
[0150] Initialize the silence list silence_list as empty;
[0151] Traverse all the words in the text. For two adjacent valid words word_i and word_i+1:
[0152] a. Calculate the time interval interval = word_i+1.start_time - word_i.end_time
[0153] b. If interval > threshold (the threshold), add the silence record to silence_list:
[0154] silence_list.append({
[0155] "start_time": word_i.end_time,
[0156] "end_time": word_i+1.start_time,
[0157] "duration": interval
[0158] })
[0159] Calculate the total silence duration total_silence_duration = sum(silence.duration for silence in silence_list);
[0160] Calculate the silence ratio silence_ratio = total_silence_duration / total_duration, where total_duration is the total duration of the entire speech;
[0161] Use the calculated silence duration and silence ratio for subsequent evaluation and analysis. For example, if the single silence time exceeds 5 seconds or the total silence time ratio exceeds 30%, it may be considered that the user's adaptability is insufficient.
[0162] 3. Emotion recognition
[0163] For emotion recognition, based on the grammar, semantic structure, and context information of the text information, combined with the result of speech rate recognition, emotion recognition is performed. In this embodiment, a multi-modal emotion recognition model is adopted, which fuses text features and speech features.
[0164] Among them, the text features are extracted by the BERT model, and the specific steps are as follows:
[0165] 11) Input the text into the pre-trained BERT model;
[0166] 12) Extract the hidden state of the last layer of the BERT model as the semantic representation of the text;
[0167] 13) Use the attention mechanism to weight the semantic representation to obtain the emotional features of the text.
[0168] The speech features include fundamental frequency, energy, spectral features, etc., and the specific steps are as follows:
[0169] 21) Extract features such as fundamental frequency (F0), energy, and Mel-frequency cepstral coefficients (MFCC) from the original speech signal;
[0170] 22) Calculate the statistics of these features, such as mean, standard deviation, maximum value, minimum value, etc.;
[0171] 23) Combine these statistics into the emotional features of the speech.
[0172] After fusing the text features and speech features, input them into the emotion classifier to obtain the emotion category (such as positive, negative, neutral) and emotion intensity. The emotion classifier is a fully connected neural network, including two hidden layers, with 128 neurons in each layer, using the ReLU activation function, and the output layer uses the Softmax activation function for multi-classification.
[0173] Use the recognized emotion information for subsequent evaluation and analysis. For example, if it is recognized that the user's emotion is "angry" and the emotion intensity is high, it may be considered that the user's emotion control ability is insufficient.
[0174] Embodiment 4
[0175] The electronic device of the present invention includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0176] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0177] The processing unit executes the various methods and processes described above, such as methods S1 to S6. For example, in some embodiments, methods S1 to S6 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S6 described above can be executed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S6 by any other suitable means (e.g., by means of firmware).
[0178] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0179] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0180] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0181] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An information processing method for human-computer training, characterized in that: The method comprises the following steps: S1, real-time acquisition of interactive voice responses from users to questions; S2, performs real-time speech recognition on interactive speech and converts it into text information with timestamp; S3, preprocesses text information and extracts key semantic information, and performs emotion recognition, silence recognition, and speech rate recognition; S4, based on the extracted key semantic information, emotion recognition, silence recognition and speech speed recognition results, the practice process is evaluated in real time according to the quality inspection rules, and it is determined whether the interaction needs to be interrupted according to the evaluation results. If interruption is required, step S5 is executed immediately. Otherwise, step S5 is executed after completing the complete interactive speech recognition and processing of this round; S5, performs intent recognition based on the extracted key semantic information, maps the recognized intent to the corresponding node of the knowledge graph, iteratively updates the associated question path based on the knowledge graph, generates response questions based on the scenario labels and process rules, and sends them to the user; S6, obtaining the interactive voice of the user in response to the question, and returning to step S2 to perform the next round of human-computer interaction.
2. The information processing method for human-computer training according to claim 1, characterized in that: The iterative updating of the associated question path based on the knowledge graph and the generation of response questions in combination with the scenario labels and process rules are specifically as follows: Through part-of-speech tagging, synonym expansion and industry-related words, the keywords are exhaustively listed to build a keyword knowledge base under different scenario labels and process rules; Perform keyword recognition based on the keyword knowledge base and the extracted key semantic information; According to the identified intent, search for matching nodes and relationship paths in the knowledge graph, select the optimal inquiry path based on the identified relationship path and user historical interaction records, and determine the keywords corresponding to the optimal inquiry path; Combine the keywords obtained by keyword recognition and the keywords obtained according to the optimal path, match the corresponding scenario tags and process rules, and extract the questions with the highest correlation with the keywords of this round from the question and answer library pre-configured by the scenario tags and process rules as the response questions.
3. The information processing method for human-computer training according to claim 2, characterized in that: The method for extracting the question with the highest correlation with the keyword of this round from the question-answer library pre-configured by the scenario label and process rule is as follows: The question set in the question-answering database is Q = {q1, q2, ...q K }, where K is the number of questions. For each question q k , calculate its correlation R with the current round of interactive speech rel : R rel (q k )=αK(q k )+βS(q k ,C)+γH(q k ) Among them, α, β, γ are weights, K(q k ) represents question q k and the normalized matching score of all keywords, S(q k ) represents question q k The similarity between the extracted semantic information C and H(q k ) represents question q k The probability of history being mentioned under the scenario label and process rules.
4. The information processing method for human-computer training according to claim 1, characterized in that: The quality inspection rules include a rule name, a category, a corresponding score and a violation condition, wherein the violation condition is constructed by selecting at least one of a semantic tag, a keyword and a regular expression and combining it with a preset rule threshold.
5. The information processing method for human-computer training according to claim 4, characterized in that: The semantic tags include user semantic tags and machine semantic tags. The user semantic tags are key business nodes involved in the user's interaction process. The key business nodes are set according to scenario tags and process rules, and each key business node is bound to multiple non-default global rules; the machine semantic tags are used to configure the non-default global rules involved in the machine's interaction process and / or generate guidance rules for responding to questions.
6. The information processing method for human-computer training according to claim 1, characterized in that: In step S4, the tolerance of different scenarios to semantic deviation is determined according to the scenario labels, the threshold of the quality inspection rule is adjusted based on the tolerance, and the user's practice process is scored.
7. The information processing method for human-computer training according to claim 1, characterized in that: The emotion recognition, silence recognition and speech speed recognition are specifically as follows: Based on the text information with timestamps, the number of words per minute is counted to perform speech speed recognition; Based on the text information with timestamp, the length of time interval without valid words is counted to perform silence recognition; Emotion recognition is performed based on the grammatical, semantic structure and contextual information of the text information and the results of speech rate recognition.
8. The information processing method for human-computer training according to claim 1, characterized in that: The first-level evaluation dimensions for evaluating the practice process in real time according to the quality inspection rules include semantic ability, speech rate control ability, adaptability, accuracy, proficiency and scene completeness, wherein the second-level evaluation dimensions of the semantic ability include terminology accuracy, process compliance and logical rigor, the speech rate control ability is scored according to the deviation value between the speech rate recognition result and the preset speech rate benchmark value, the second-level evaluation dimensions of the adaptability include response timeliness, solution innovation and anti-interference ability, the second-level evaluation dimensions of the accuracy rate include data accuracy, compliance and user information accuracy, the second-level evaluation dimensions of the proficiency include task completion time and error reproduction rate, the second-level evaluation dimensions of the scene completeness include process coverage and branch processing completeness, and in each second-level evaluation dimension, the corresponding scoring / deduction points are set according to the preset quality inspection rules.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Man-machine conversation method and device, electronic device and storage medium
CN112365892A
Intelligent voice dialogue scene verbal skill intervention method and system based on customer portrait
CN116049360A
Cited By
Insurance intention identification method and system
CN121034295A
Intelligent sales practice method based on three-dimensional dynamic scoring and strategy chain continuity
CN121660611A