Dialogue analysis method, device, glasses, equipment and readable storage medium

By converting the conversation audio into text in real time and labeling the function type of the conversation wheel, combining sequence detection rules and diagnostic network to generate feedback, the problem of inability to instant feedback in traditional dialogue analysis methods is solved, and dialogue efficiency and quality are improved.

CN119646145BActive Publication Date: 2025-07-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411688696.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-07-22
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

Traditional dialogue analysis methods cannot achieve immediate feedback on dialogue situations, resulting in the inability to timely optimize the dialogue process in scenarios where dialogue direction needs to be flexibly adjusted, affecting teaching effect and interaction quality.

Method used

By collecting conversation audio in real time to convert it into text, using the first text network to label the call wheel function type, combining preset sequence detection rules and the second text network diagnosis to generate feedback information, providing dialogue leaders with immediate adjustment suggestions.

Benefits of technology

Realize instant feedback and flexible adjustments during the dialogue process, improve the efficiency and quality of dialogue, promote effective communication and interaction, and improve participants' satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646145B_ABST
    Figure CN119646145B_ABST
Patent Text Reader

Abstract

The present application provides a dialogue analysis method, apparatus, glasses, device and readable storage medium. By collecting dialogue audio in real time and converting it into dialogue text, and annotating the dialogue text based on the discourse function type of the turn, it can provide in-depth information support for the understanding of dialogue types. Using the sequence detection rules of preset dialogue types and machine learning algorithms to accurately identify key information and structural features in the dialogue, and then instantly generating feedback information and displaying it, enabling the dialogue guide to quickly understand the real-time situation of the dialogue during the dialogue process, facilitating quick dialogue adjustment, optimizing the dialogue process, and improving the efficiency and effect of the dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a dialogue analysis method, apparatus, glasses, device, and readable storage medium. Background Art

[0002] Currently, in traditional dialogue analysis scenarios, dialogue records are usually relied on means such as recording and video recording, and after the entire dialogue ends, a large language model is used to conduct a detailed analysis of the dialogue content to identify problems existing in the dialogue process. However, this analysis method cannot meet the dialogue scenarios that require efficient feedback on the dialogue situation to facilitate flexible adjustment of the dialogue direction, such as classroom teaching, customer interviews, hosting scenarios, etc. Taking the classroom teaching scenario as an example, teachers need to wait until after class to obtain the dialogue situation through classroom dialogue analysis, and then retrospectively analyze the teaching situation in the classroom for subsequent teaching adjustment, missing the important opportunity to immediately adjust teaching strategies and interaction methods in the classroom, which affects the teaching effect.

[0003] Therefore, the traditional dialogue analysis method has obvious limitations in scenarios that require efficient dialogue feedback to flexibly adjust the dialogue direction, and there is an urgent need for a dialogue analysis method that can achieve efficient feedback on the dialogue situation to facilitate flexible adjustment of the dialogue direction. Summary of the Invention

[0004] In view of this, to solve the above technical problems, the present application provides a dialogue analysis method, apparatus, glasses, device, and readable storage medium.

[0005] Specifically, the present application is implemented through the following technical solutions:

[0006] According to the first aspect of the embodiments of the present application, a dialogue analysis method is provided, and the method includes:

[0007] Convert the real-time received dialogue audio into dialogue text;

[0008] Use a first text network to annotate each turn in the dialogue text to obtain an annotated text with an annotation result; the annotation result is used to represent the discourse function type of the turn;

[0009] According to the order and annotation result of the turns in the annotated text, obtain an annotation result sequence;

[0010] According to the sequence detection rules of a preset dialogue type, determine the sequence recognition result of the annotation result sequence; the sequence recognition result includes the dialogue type to which the annotation result sequence belongs; the preset dialogue type is determined respectively according to different dialogue goals in the dialogue scenario to which the dialogue audio belongs;

[0011] Based on the sequence recognition result, use a second text network to diagnose the annotated text and generate feedback information; the feedback information is at least used to prompt the conversation situation to the conversation leader in real time to optimize the conversation.

[0012] Optionally, the step of using the first text network to annotate each turn in the conversation text includes:

[0013] Obtain the classification conditions for each preset annotation label; the preset annotation labels correspond one-to-one with the discourse function types of turns in the conversation context; construct a prompt text applicable to prompt engineering according to the classification conditions; use the first text network and the prompt text to identify and annotate the annotation labels to which each turn in the conversation text belongs.

[0014] Optionally, the step of using the first text network to annotate each turn in the conversation text includes:

[0015] Input the text content in the turn into the first text network to obtain the annotation result output by the first text network; wherein, multiple classification labels are set in the output layer of the first text network, and the classification labels correspond one-to-one with the discourse function types of turns in the conversation context; the annotation result is obtained by annotating the classification label greater than the set threshold in the probability distribution of the multiple classification labels generated by the first text network.

[0016] Optionally, the step of determining the sequence recognition result of the annotation result sequence according to the sequence detection rules of the preset conversation type includes:

[0017] When the annotation result sequence meets any of the sequence detection rules, determine the preset conversation type corresponding to the satisfied sequence detection rule as the conversation type to which the annotation result sequence belongs, and generate a first recognition result;

[0018] When the annotation result sequence does not meet the sequence detection rules of all preset conversation types, determine that the conversation type it belongs to indicates an unknown conversation type, and respectively determine the rule information that the annotation result sequence does not meet for each sequence detection rule, and generate a second recognition result.

[0019] Optionally, the step of diagnosing the annotated text using the second text network according to the sequence recognition result includes:

[0020] When the sequence recognition result is the first recognition result, obtain the conversation goal of the preset conversation type corresponding to the satisfied sequence detection rule; according to the conversation goal, use the second text network to diagnose the annotated text.

[0021] Optionally, diagnosing the labeled text using a second text network according to the sequence recognition result includes:

[0022] When the sequence recognition result is the second recognition result, obtaining the rule information included in the second recognition result; indicating an unknown dialogue type according to the rule information and the dialogue type to which it belongs, and diagnosing the labeled text using the second text network.

[0023] Optionally, the method further includes:

[0024] Obtaining video data collected simultaneously with the dialogue audio; aligning the video data with the labeled text according to timestamps; determining the participation level of the dialogue participants according to the video data and the labeled text; determining detection parameters matching the participation level of the dialogue participants according to a preset corresponding relationship; the corresponding relationship includes preset participation levels and detection parameters corresponding to each participation level;

[0025] Diagnosing the labeled text using a second text network according to the sequence recognition result includes: diagnosing the labeled text using the second text network according to the sequence recognition result and the detection parameters.

[0026] Optionally, the method further includes:

[0027] Outputting the feedback information in a structured mode; the feedback information may at least include at least one of dialogue participation situation, interaction prompt, and dialogue direction suggestion;

[0028] Converting the feedback information into a virtual page according to a preset page layout and projecting it for display within the user's line of sight.

[0029] According to a second aspect of the embodiments of the present application, a dialogue analysis device is provided, and the device includes:

[0030] An audio conversion module for converting real-time received dialogue audio into dialogue text;

[0031] A labeling module for labeling each turn in the dialogue text using a first text network to obtain a labeled text with a labeling result; the labeling result is used to represent the discourse function type of the turn;

[0032] A sequence acquisition module for obtaining a labeling result sequence according to the order and labeling result of the turns in the labeled text;

[0033] A sequence recognition module, configured to determine the sequence recognition result of the labeled result sequence according to the sequence detection rules of a preset dialogue type; the sequence recognition result includes the dialogue type to which the labeled result sequence belongs; the preset dialogue type is determined respectively according to different dialogue objectives in the dialogue scenario to which the dialogue audio belongs;

[0034] A feedback information generation module, configured to diagnose the labeled text by using a second text network according to the sequence recognition result, and generate feedback information; the feedback information is at least used to prompt the dialogue situation to the dialogue leader in real time to optimize the dialogue.

[0035] Optionally, the labeling module is specifically configured to:

[0036] Obtain the classification conditions of each preset labeling tag; the preset labeling tags are in one-to-one correspondence with the discourse function types of turns in the dialogue scenario; construct a prompt text applicable to prompt engineering according to the classification conditions; use the first text network and the prompt text to identify the labeling tags to which each turn in the dialogue text belongs and perform labeling.

[0037] Optionally, the labeling module is specifically configured to:

[0038] Input the text content in the turn into the first text network to obtain the labeling result output by the first text network; wherein, multiple classification tags are set in the output layer of the first text network, and the classification tags are in one-to-one correspondence with the discourse function types of turns in the dialogue scenario; the labeling result is obtained by labeling the classification tags greater than a set threshold in the probability distribution of the multiple classification tags generated by the first text network.

[0039] Optionally, the sequence recognition module is specifically configured to:

[0040] When the labeled result sequence satisfies any of the sequence detection rules, determine the preset dialogue type corresponding to the satisfied sequence detection rule as the dialogue type to which the labeled result sequence belongs, and generate a first recognition result;

[0041] When the labeled result sequence does not satisfy the sequence detection rules of all preset dialogue types, determine that the indicated dialogue type belongs to an unknown dialogue type, and respectively determine the rule information that the labeled result sequence does not satisfy for each sequence detection rule, and generate a second recognition result.

[0042] Optionally, the feedback information generation module is specifically configured to:

[0043] When the sequence recognition result is the first recognition result, obtain the dialogue target of the preset dialogue type corresponding to the satisfied sequence detection rule; according to the dialogue target, use the second text network to diagnose the annotated text.

[0044] Optionally, the feedback information generation module is specifically configured to:

[0045] When the sequence recognition result is the second recognition result, obtain the rule information included in the second recognition result; according to the rule information and the indicated unknown dialogue type of the dialogue type to which it belongs, use the second text network to diagnose the annotated text.

[0046] Optionally, the apparatus further includes:

[0047] Obtain video data collected simultaneously with the dialogue audio; align the video data with the annotated text according to timestamps; determine the dialogue participant engagement level according to the video data and the annotated text; according to a preset correspondence, determine detection parameters matching the dialogue participant engagement level; the correspondence includes preset engagement levels and detection parameters corresponding to each engagement level;

[0048] The feedback information generation module is specifically configured to: according to the sequence recognition result and the detection parameters, use the second text network to diagnose the annotated text.

[0049] Optionally, the apparatus further includes:

[0050] Output the feedback information in a structured mode; the feedback information may at least include at least one of dialogue participation status, interaction prompts, and dialogue direction suggestions; according to a preset page layout, convert the feedback information into a virtual page and project it for display within the user's line of sight.

[0051] According to a third aspect of the embodiments of the present application, there is provided an electronic device, the electronic device includes: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the above-mentioned dialogue analysis method by calling the computer program.

[0052] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned dialogue analysis method is implemented.

[0053] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0054] In the technical solution provided by the present application above, by collecting dialogue audio in real time and converting it into dialogue text, and annotating the dialogue text based on the discourse function type of the turn, it is possible to provide in-depth information support for the understanding of dialogue types, accurately identify key information and structural features in the dialogue by using the sequence detection rules of the preset dialogue type and machine learning algorithms, thereby instantly generating feedback information and displaying it, enabling the dialogue leader to quickly understand the real-time situation of the dialogue during the dialogue process, facilitating quick dialogue adjustment, optimizing the dialogue process, improving the efficiency and effect of the dialogue, promoting effective communication and interaction among dialogue participants, and enhancing the overall quality of the dialogue and the satisfaction of participants.

[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. In addition, any embodiment in the present application does not need to achieve all the above effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0057] Figure 1A is an overall flowchart of a dialogue analysis method shown in an exemplary embodiment of the present application;

[0058] Figure 1B is an example of the discourse function type of a preset turn in a classroom teaching scenario shown in an exemplary embodiment of the present application;

[0059] Figure 1C is an example of a preset dialogue type divided based on different teaching objectives in a classroom teaching scenario shown in an exemplary embodiment of the present application;

[0060] Figure 2 is a step flowchart of obtaining the sequence recognition result of the annotation result sequence corresponding to the annotated text based on the sequence detection rule shown in an exemplary embodiment of the present application;

[0061] Figure 3 is a step flowchart of a second text network diagnosis according to the sequence recognition result shown in an exemplary embodiment of the present application;

[0062] Figure 4 is another step flowchart of a second text network diagnosis according to the sequence recognition result shown in an exemplary embodiment of the present application;

[0063] Figure 5 is a schematic diagram of the structure of glasses applying the dialogue analysis method shown in an exemplary embodiment of the present application;

[0064] Figure 6It is a schematic structural diagram of a dialogue analysis device shown in an exemplary embodiment of the present application;

[0065] Figure 7 It is a hardware schematic diagram of an electronic device shown in an exemplary embodiment of the present application. Detailed implementation manners

[0066] Next, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0067] It should be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. In the present application, terms such as first, second, and third may be used to describe various information, but such information should not be limited to these terms, and these terms are only used to distinguish the same type of information from each other.

[0068] In the current dialogue analysis scenario, traditional dialogue recording and analysis methods mainly rely on media such as audio and video recordings to capture and save dialogue content in order to completely save the entire process of the dialogue. After the entire dialogue process ends, a large language model is used to conduct a detailed analysis and interpretation of the recorded dialogue content to identify possible problems or deficiencies in the dialogue process.

[0069] With the increasing requirements for the quality of dialogue interaction in various industries, this traditional dialogue analysis method has obvious limitations. That is, to a certain extent, this dialogue analysis method can achieve a comprehensive review and evaluation of dialogue content, but its inherent lag limits its application effect in dialogue scenarios that require immediate feedback and adjustment.

[0070] Taking the classroom teaching scenario as an example, since traditional dialogue analysis methods need to analyze and evaluate dialogue content after class, teachers can often only obtain the feedback results of classroom dialogue after class, thus missing the important opportunity to immediately adjust teaching strategies and interaction methods in class, affecting teachers' precise control of the teaching process, restricting the improvement of teaching effects, and unable to meet the needs of teachers in this scenario for efficient feedback on dialogue situations and flexible adjustment of teaching strategies and interaction methods according to real-time interaction situations with students.

[0071] In addition, in other dialogue scenarios that require immediate feedback and adjustment, such as customer interviews and hosting scenarios, traditional dialogue analysis methods also have the same limitations. Therefore, there is an urgent need for a dialogue analysis method that can achieve efficient feedback on dialogue situations and facilitate flexible adjustment of the dialogue direction.

[0072] In view of the above technical problems, the present application proposes a dialogue analysis method. The execution subject of this dialogue analysis method can be a server, an intelligent terminal device such as a mobile phone or a computer, or a device or apparatus such as dedicated hardware and chips that can provide the computing power required for dialogue analysis. The present application does not limit this. This method can capture and analyze the dialogue content in real time, and provide instant feedback and adjustment suggestions on the dialogue situation, thereby effectively improving the quality and efficiency of the dialogue.

[0073] This dialogue analysis method can be widely applied to various dialogue scenarios that require instant feedback and flexible adjustment of the dialogue situation, such as meeting hosting scenarios, classroom teaching scenarios, psychological counseling scenarios, customer interview scenarios, etc. Using this dialogue analysis method can instantly provide the evaluation results of the specific situation of the current dialogue content for the dialogue leaders (such as the host, teacher, psychologist, etc.) in the dialogue scenario, prompt the deficiencies of the dialogue content and dialogue adjustment suggestions, so as to improve the quality of dialogue interaction, optimize the dialogue process, and promote effective communication and interaction among dialogue participants.

[0074] See Figure 1A The step flow chart shown by way of example. This method can at least include the following steps:

[0075] S101, convert the dialog audio received in real time into dialog text;

[0076] The dialog audio received in real time refers to the digital audio signal captured and transmitted in real time through a microphone or other audio input devices during the dialogue process. For example, in the classroom teaching scenario, the dialog audio is the classroom dialogue collected in real time, including the dialogue interaction between the teacher and the students. When the execution subject of the method and the audio input device for collecting the dialog audio are located inside the same device, the dialog audio can be transmitted through short-distance and fast transmission methods such as Bluetooth, Wi-Fi direct connection or wired connection; when the execution subject of the method and the audio input device are not located inside the same device, such as when the dialog audio collected by the microphone needs to be transmitted to the cloud server for processing in real time, the dialog audio can be transmitted through networks such as the Internet and local area network.

[0077] The dialog text refers to the content that converts the speech information in the dialog audio into a readable text form, and can be stored in the form of turns. A turn is a basic unit in a dialogue, referring to a continuous speech in the dialogue, spoken by one participant until interrupted or taken over by the speech of another participant. In the dialog text, each turn can be clearly identified and is usually distinguished by line breaks, punctuation marks or other delimiters. The dialog text can include speaker identification, speech content, timestamp information indicating the time point when the turn occurs, and other information such as emotion tags, semantic role annotation, etc.

[0078] To reduce background noise and improve the clarity of the speech signal, before converting the speech content of the conversation audio into conversation text, the received conversation audio can be preprocessed, including operations such as noise reduction, gain control, echo cancellation, and endpoint detection, thereby improving the accuracy of speech recognition.

[0079] Regarding converting the conversation audio into conversation text, audio features can be extracted from the digital audio signal through a pre-trained speech recognition model, and the extracted audio features can be converted into text content. The speech recognition model can be a neural network model based on deep learning, such as a recurrent neural network, a long short-term memory network, or a Transformer, etc. These models can learn the mapping relationship between the speech signal and the text content through the training of a large amount of speech data.

[0080] S102, using the first text network to annotate each turn in the conversation text to obtain an annotated text with annotation results; the annotation results are used to represent the discourse function type of the turn.

[0081] A turn refers to a continuous speech of a speaker in a conversation until another speaker starts to speak, and it is the basic unit of a conversation. Each turn contains a complete semantic unit, which can be a sentence or multiple sentences. For example, if A says a paragraph and B responds with a sentence, then A's paragraph belongs to one turn and B's sentence belongs to one turn. For example, in the conversation text example: Teacher: "Please describe the picture you saw yesterday." Student: "The picture is a forest with many green trees.", then the teacher's speech belongs to one turn and the student's speech belongs to another turn, that is, this conversation text example contains two turns.

[0082] The discourse function type is used to represent the role or function of a turn in a conversation, such as question, answer, statement, supplementary explanation, etc., which is convenient for understanding the structure and context of the conversation and for subsequent conversation analysis and processing. The discourse function type is related to the conversation context to which the received conversation audio belongs. Based on different conversation contexts, there are unique interaction purposes, social norms, and information flow patterns, and this factor will affect the function assumed by the turn. Therefore, in order to more accurately reflect the conversation characteristics and requirements in this conversation context, the classification of the discourse function type can be set separately for different conversation contexts.

[0083] For example, if the collected conversation is in the classroom teaching scenario, then the conversation context to which the conversation audio belongs is the classroom teaching scenario, see Figure 1BAs shown, the discourse function types of the turn in this dialogue context may include but are not limited to: invitation to elaborate, elaboration, invitation to reason, explanation of reasons, invitation to coordinate ideas, approval, questioning, review and citation, extensive reference, other invitations, silence, etc. Different discourse function types cooperate with each other in classroom teaching scenarios to jointly construct rich classroom dialogues, promote the transfer of knowledge, the cultivation of thinking and the conduct of interaction. Similarly, in other dialogue scenarios, such as the host scenario, the discourse function types may include but are not limited to opening greetings, guiding topics, summarizing, inviting speeches, adjusting the atmosphere, etc., forming a unique discourse function type system according to the characteristics of the dialogue context to meet different interaction needs.

[0084] In text processing, annotation refers to the process of adding additional information (such as discourse function type or emotion label) to each unit (such as sentence or turn) in the text. The annotation result is the information obtained after annotation. In this embodiment, it represents the annotation information assigned to each turn in the dialogue text, which represents the discourse function type of the turn.

[0085] The first text network is used to identify and annotate the speech function type of each turn in the dialogue text, and the basis for the turn annotation is the speech function types of multiple turns preset for the dialogue context to which the collected dialogue audio belongs. The first text network may include a deep learning network, such as a convolutional neural network, a recurrent neural network, a long short-term memory network, a Transformer, etc., which can be used as the first text network after being pre-trained or fine-tuned according to the annotation task of the present application. In this embodiment, for the acquired dialogue text, each turn can be identified and segmented by detecting pauses, line breaks or other separators in the dialogue text, and then each turn is passed as input to the first text network, so that the first text network identifies the speech function type of each turn and outputs the annotation result representing the speech function type of the turn or the turn with the annotation result. By identifying and annotating the speech function type of each turn (i.e., a continuous speech of each speaker) in the dialogue text, a text with annotation information is finally generated.

[0086] In the process of annotating using the first text network dialogue round, in order to ensure the consistency and readability of the annotation results and improve the efficiency of subsequent analysis and processing of the annotation results, the discourse function type can be represented by coding, that is, a unique code is set for each discourse function type, so that when annotating in the dialogue round, the corresponding discourse function type can be quickly identified and the corresponding coding annotation can be added. For example, see Figure 1B The coding example in , sets a unique code for each discourse function type in the aforementioned example classroom teaching scenario.

[0087] For example, suppose there is a conversation text M1 as follows:

[0088] Teacher: "Please describe the painting you saw yesterday."

[0089] Student: "The painting is of a forest with many green trees."

[0090] Teacher: "Can you tell me why this painting impressed you deeply?"

[0091] Student: "Because the forest looks very peaceful and makes me feel relaxed."

[0092] The labeled text M2 after the first text network annotation can be exemplified as follows:

[0093] Teacher: "Please describe the painting you saw yesterday." [Invite for elaboration / ELI]

[0094] Student: "The painting is of a forest with many green trees." [Elaboration / EL]

[0095] Teacher: "Can you tell me why this painting impressed you deeply?" [Invite for reasoning / REI]

[0096] Student: "Because the forest looks very peaceful and makes me feel relaxed." [Reason explanation / RE]

[0097] It should be understood that based on one turn representing a continuous speech of a speaker and containing at least one sentence, therefore, one turn corresponds to one or more discourse function types. In the process of labeling the turn using the first text network, in the case where the turn corresponds to multiple discourse function types, the labels representing these discourse function types can be added to the same turn in sequence according to the order of the text content in the turn, and the labeling result representing multiple discourse function types is used as the labeling result of the turn. That is, the same turn can include multiple labels representing discourse function types, and each label corresponds to a specific part of the text content in the turn.

[0098] For example, taking the discourse function types in the previous classroom teaching scenario as an example, for the turn example "According to the Pythagorean theorem we learned in the previous class, can you explain why here...?", the text content "According to the Pythagorean theorem we learned in the previous class" in this turn represents review and reference / RB, and the text content "Can you explain why" represents asking a question / Q. That is, this turn corresponds to two discourse function types, and the labeling result of this turn can be expressed as [Review and reference, Asking a question] or [RB, Q].

[0099] S103. Obtain a labeled result sequence according to the order and labeling result of the turns in the labeled text;

[0100] The labeled result sequence is an abstract representation of the dialogue text, removing the specific dialogue content in the dialogue text and retaining the functional types and sequential relationships of the turns. It contains the labeled results indicating the discourse functional types of each turn in the dialogue text and is arranged in the order of the turns appearing in the dialogue text.

[0101] For example, for the example labeled text M2 mentioned above, taking the labeled result in encoded form as an example, the labeled result sequence corresponding to this labeled text can be represented as [ELI, EL, REI, RE].

[0102] S104, according to the sequence detection rules of the preset dialogue type, determine the sequence recognition result of the labeled result sequence; the sequence recognition result includes the dialogue type to which the labeled result sequence belongs; the preset dialogue type is determined respectively according to different dialogue objectives in the dialogue situation to which the dialogue audio belongs;

[0103] The preset dialogue type is a dialogue type set respectively according to different dialogue objectives in the dialogue situation to which the collected dialogue audio belongs. Different dialogue types are used to indicate that the dialogue in this dialogue situation has different orientations and characteristics. For example, taking the dialogue situation of classroom teaching as an example, since different teaching objectives in the classroom are reflected through the interaction methods and topic contents between teachers and students, such as Figure 1C As shown, classroom dialogue can be divided into four categories according to teaching objectives: critical inquiry, collaborative knowledge construction, teaching and supportive dialogue, and reflection and metacognitive dialogue. Each dialogue type corresponds to different educational objectives shown in the figure; among them, critical inquiry dialogue is used to stimulate students' in-depth thinking and questioning ability about the knowledge learned; collaborative knowledge construction dialogue means constructing knowledge by coordinating multiple viewpoints, jointly solving problems, sharing resources and ideas; teaching and supportive dialogue is used to indicate that teachers provide knowledge explanations and provide teaching examples for argument support during the dialogue process with students; reflection and metacognitive dialogue is used to indicate guiding dialogue participants to reflect on the knowledge learned or the methods and strategies adopted in history.

[0104] Based on different dialogue types having different labeled result sequence patterns, in this embodiment, sequence detection rules are set for each dialogue type. The sequence detection rules include a set of predefined rules, which are set based on the dialogue objectives corresponding to the dialogue type, as well as information such as the characteristics and structure of this dialogue type. Among them, the rules are set based on the discourse functional types of the turns in the dialogue situation to which the dialogue audio belongs, considering factors such as the order of the turns in the dialogue, the relationship between the discourse functional types of the turns, the occurrence frequency or combination method of specific turns, etc., and are used to identify the dialogue type to which the labeled text belongs according to the labeled result sequence corresponding to the labeled text.

[0105] For example, taking the dialogue type - critical inquiry in the aforementioned classroom teaching as an example, its exemplary sequence detection rule can be expressed as: [meeting the following three conditions simultaneously in the same topic: ① including at least one of invitation to elaborate (ELI), invitation to reason (REI), and coordination invitation (CI); ② including at least one of elaboration (EL) and reason explanation (RE); ③ including raising questions (Q)]. Based on this exemplary sequence detection rule, for example, the labeled result sequence patterns that conform to this sequence detection rule may include, but are not limited to, [REI, RE, Q], [Q, RE, REI], [CI, Q, RE], [ELI, Q, RE], etc.

[0106] In this embodiment, according to the preset sequence detection rules for different dialogue types, the labeled result sequence obtained in step S103 is matched or compared with the preset sequence detection rules for each dialogue type, and the matching degree between the functional types and sequential relationships of the turns in the labeled result sequence and the sequence detection rules is analyzed. It is determined which dialogue type the labeled result sequence belongs to based on whether the labeled result sequence corresponding to the labeled text meets any of the sequence detection rules, and a sequence recognition result is generated. This sequence recognition result is used to represent the matching situation between the labeled result sequence and the sequence detection rules, providing the overall structure and type information of the dialogue audio. Since this embodiment is for labeling the dialogue text received in real time to obtain the labeled result sequence, there are two types of matching situations for the labeled result sequence corresponding to this, namely meeting the sequence detection rules of one of the preset dialogue types and not meeting the sequence detection rules of all preset dialogue types. For the matching situation of not meeting the sequence detection rules of all preset dialogue types, it can be determined that the dialogue type to which the labeled result sequence belongs is an unknown type.

[0107] S105, according to the sequence recognition result, use the second text network to diagnose the labeled text and generate feedback information; the feedback information is at least used to prompt the dialogue situation to the dialogue leader in real time to optimize the dialogue.

[0108] This second text network is used to further analyze, diagnose, or evaluate the labeled text. It can be a pre-trained deep learning model such as a convolutional neural network, a recurrent neural network, or its variants, etc. Or, it can also be other types of machine learning models that are fine-tuned for text diagnosis tasks. There are functional differences compared to the first text network. The first text network described in the previous steps is mainly used for labeling dialogue text, while this second text network focuses on analyzing and diagnosing the labeled text, so as to output feedback information for guiding and adjusting the dialogue, achieving the effects of instantaneously feedbacking the dialogue situation and flexibly adjusting the dialogue based on the feedback information.

[0109] In this embodiment, the second text network is used to perform a more in-depth analysis of the annotated text according to the dialogue target corresponding to the dialogue type included in the sequence recognition result, obtain the problems or improvement points in the dialogue process represented by the annotated text, and generate feedback information to optimize the current dialogue process. The diagnosis of the second text network may include, but is not limited to, data analysis and processing in aspects such as discourse function diagnosis, turn structure analysis, and analysis of engagement-related elements.

[0110] The feedback information is the result generated after the second text network diagnoses the annotated text, and is used to provide information or suggestions about the dialogue situation to the dialogue leader, and may include, but is not limited to, dialogue quality assessment, dialogue improvement suggestions, dialogue precautions, etc. For example, assessments in aspects such as the fluency, engagement, and topic concentration of the dialogue, whether the progress of the dialogue meets expectations, and whether there are any deficiencies or overuses of certain discourse function types, so that the dialogue leader can understand the current state of the dialogue, adjust the dialogue strategy in a timely manner, and improve the dialogue effect. Among them, the dialogue leader refers to the speaker who plays a guiding role in the dialogue, and is responsible for proposing topics, guiding discussions, or ensuring the smooth progress of the topics. For example, teachers in classroom teaching scenarios, hosts in hosting scenarios, and psychologists in psychological counseling scenarios.

[0111] Regarding the diagnosis of the annotated text using the second text network according to the sequence recognition result, after obtaining the sequence recognition result, according to the dialogue type to which the annotation result sequence included in the sequence recognition result belongs, obtain the preset prompt template corresponding to the dialogue type, and fill the prompt template with the information carried in the sequence recognition result to obtain a diagnostic prompt text, and then input the diagnostic prompt text and the annotated text into the trained second text network to obtain the feedback information output by the second text network based on the diagnosis result.

[0112] For example, in a classroom teaching scenario, after processing through the above steps to obtain a sequence recognition result, construct a diagnostic prompt text based on the sequence recognition result, and use the diagnostic prompt text and the annotated text as the input of the second text network; the second text network extracts high-level text features from the input based on the parameters learned in the pre-training stage, and generates corresponding feedback information through the output layer that generates text from features. For example, the feedback information may include teacher guidance suggestions generated according to the classroom situation, such as "Please avoid long explanations", "It is recommended to switch to a question-and-answer mode", "It is recommended to further explain the current topic", to guide the adjustment of teaching strategies.

[0113] In the embodiments of the present disclosure, by quickly converting the real-time received conversation audio into conversation text and using the first text network to construct an annotated text with annotation results, the discourse function type of each turn in the conversation text is accurately identified, which is convenient for better and deeper understanding of the structure and content of the conversation. Based on the order of turns and the annotation results in the annotated text, an annotation result sequence is generated, and multiple sequence detection rules for different conversation types set based on the conversation goal in the current conversation context are used to identify the conversation type to which the annotation result sequence belongs and generate a sequence recognition result. Then, combined with the sequence recognition result and machine learning, the annotated text is deeply diagnosed to identify the problems or deficiencies in the conversation and generate feedback information to be conveyed to the conversation leader, realizing instant conversation analysis and conversation situation feedback during the conversation process, enabling the conversation leader to understand the conversation situation in real time during the conversation and flexibly adjust the conversation rhythm and direction according to the feedback, which helps to guide the conversation to develop in a more efficient and in-depth direction and improve the quality and effect of the conversation.

[0114] Based on the conversation analysis method provided in the foregoing embodiments, taking the classroom teaching scenario as an example, during the classroom teaching conversation process, the teacher can obtain feedback information instantaneously through real-time conversation analysis to understand the students' participation, understanding, and interaction quality; for example, if the feedback information shows that some students have a low participation rate in a certain topic, the teacher can immediately adjust the teaching strategy, increase the interaction session, and improve the students' participation rate and learning effect. Taking the hosting scenario as an example, by collecting the conversation during the hosting process in real time and analyzing it using this method, the host can obtain feedback information instantaneously during the program to understand the audience's reaction and interaction situation; for example, if the feedback information shows that the audience has little interest in a certain topic, the host can immediately adjust the topic, increase the interaction session, and improve the attraction of the program and the audience participation rate.

[0115] After explaining the overall process of the conversation analysis method provided in this application through the foregoing embodiments, further explanations and descriptions will be made on the implementation details of each process link.

[0116] In some embodiments, regarding the conversion of the conversation audio received in real time into conversation text described in the foregoing step S101, it can be achieved through automatic speech recognition technology, i.e., ASR (Automatic Speech Recognition). The ASR technology can convert the speech signal into text through steps such as speech preprocessing, feature extraction, acoustic model matching, and language model decoding. Alternatively, it is also possible to train and optimize the model through technologies such as machine learning and natural language processing, and use the trained model to process the input speech content and output text. For example, the acoustic model is trained with a large amount of speech data to improve the recognition ability for different pronunciations and accents, and the converted text is verified and corrected for grammar and semantics through natural language processing technology.

[0117] Taking the automatic speech recognition technology as an example, the conversation text can be obtained through the following steps: for the conversation audio, perform real-time text transcription on the conversation audio through speech recognition technology to output the initial text in the form of a conversation; perform data preprocessing on the initial text to obtain the conversation text; the data preprocessing includes at least one of data cleaning, syntactic analysis, and text normalization.

[0118] That is to say, when obtaining the initial text of the speech conversion text through speech recognition technology, data preprocessing can be performed on the converted initial text to further improve the accuracy and fluency of the conversation text, enhance the text quality and readability. Among them, data cleaning is used to remove noise and irrelevant information in the initial text, such as incorrect punctuation marks, redundant spaces, repeated words, etc., to ensure the neatness and consistency of the text; syntactic analysis is used to parse the grammatical structure of the initial text, perform text processing such as part-of-speech tagging that is beneficial to understanding the sentence, help identify and correct grammar errors in the sentence, and improve the logic and readability of the text; text normalization is to convert non-standard terms, abbreviations, slang, etc. in the initial text into standard language forms to ensure that the text conforms to general language norms and is easy to understand and use. Through this data preprocessing step, the quality of the conversation text output by the automatic speech recognition technology can be significantly improved, making it more accurate, fluent, and easy to understand, providing a more reliable basis for subsequent tasks such as conversation text analysis and information extraction.

[0119] In some embodiments, for the annotation of each turn in the dialogue text using the first text network in step S102 above, an annotated text with annotation results is obtained. The added annotation results can exist in the form of text labels, digital codes, or other metadata. The annotation results are used to enrich the content of the dialogue text, so as to fully understand the structure and content of the dialogue text and accurately generate feedback information. During the process of annotating each turn in the dialogue text using the first text network, the first text network can first determine the dialogue scenario to which the collected dialogue audio belongs, and obtain multiple discourse function types that a turn may have, which are predefined for this dialogue scenario; next, for each turn in the dialogue text, use the first text network to identify whether the text content in this turn conforms to one or more of the predefined multiple discourse function types, and annotate the turn according to the order of the text content in the turn and the discourse function type that the text content conforms to, to obtain the annotation result of this turn.

[0120] Regarding determining the discourse function type that a turn conforms to, it can be implemented by setting the first text network as a large language model and combining with prompt engineering, or by setting the first text network as a pre-trained classification model. Further, when the discourse function type of this turn is determined, the first text network can also perform annotation processing synchronously, that is, input the dialogue text into the first text network to obtain an annotated text composed of turns with annotation results output by the first text network.

[0121] Taking the first text network including a large language model and combining with prompt engineering as an example, the steps of annotating each turn in the dialogue text using the first text network can be implemented in the following way: obtain the classification conditions of each preset annotation label; the preset annotation labels correspond one-to-one with the discourse function types of turns in the dialogue scenario; construct a prompt text applicable to prompt engineering according to the classification conditions; use the first text network and the prompt text to identify the annotation labels to which each turn in the dialogue text belongs and perform annotation.

[0122] Among them, the classification conditions are used to define the discourse function type corresponding to the preset annotation label, and can be set from aspects such as the text content, tone, vocabulary, and phrases in the turn. For example, for the discourse function type "invite elaboration", its classification conditions are used to define or explain what kind of text content conforms to the category of inviting elaboration.

[0123] Prompt text constructed according to the classification conditions and applicable to prompt engineering, which includes descriptions of each preset annotation label, is used to guide the first text network to more accurately identify the discourse function type that the turn conforms to, that is, the annotation label to which the turn belongs. The first text network uses the prompt text as auxiliary information to annotate each turn in the dialogue text, that is, according to the content of the prompt text and the dialogue text, it can identify the annotation label to which each turn belongs and perform corresponding annotations.

[0124] In this annotation method, the classification conditions provide clear definitions and judgment criteria for the preset annotation labels, ensuring the consistency and accuracy of the annotation. Based on the classification conditions, constructing prompt text as auxiliary information can guide the first text network to more efficiently utilize context information, reduce the dependence on additional annotation resources, more accurately identify the discourse function type of the turn, reduce the cases of mislabeling and missing labeling, achieve automated annotation, and improve the accuracy and efficiency of annotation.

[0125] Taking the example that the first text network includes a pre-trained classification model, the steps of using the first text network to annotate each turn in the dialogue text can also be achieved in the following way: input the text content in the turn into the first text network to obtain the annotation result output by the first text network; among them, multiple classification labels are set in the output layer of the first text network, and the classification labels correspond one-to-one with the discourse function types of the turns in the dialogue context; the annotation result is obtained by annotating the classification label greater than the set threshold in the probability distribution of the multiple classification labels generated by the first text network.

[0126] Among them, in the case where the first text network includes a classification model, the training samples used in the pre-training process of the first text network include multiple sample turns, and each sample turn has a known annotation representing the discourse function type; during the pre-training process, by inputting the sample turn into the first text network, obtaining the predicted label output by the first text network, calculating the loss between the predicted label and the corresponding known annotation of the sample turn and minimizing the loss, and iteratively training the first text network until the training convergence condition is reached, obtaining the trained first text network, which can use the knowledge learned during the training process to understand and analyze the input turn and complete the annotation task based on the turn function type.

[0127] In this annotation method, by using the pre-trained first text network to automatically annotate each turn in the dialogue text, the labor cost is saved and the annotation efficiency can be improved.

[0128] In some embodiments, based on the present application, accurate annotation is performed on the dialogue text converted from the real-time received dialogue audio to obtain an annotation result sequence. The dialogue text will be continuously updated as the dialogue progresses. Therefore, the turn annotation in the dialogue text and the recognition process according to the sequence detection rules after obtaining the annotation result sequence are both real-time. To avoid the computational resource consumption and redundant recognition results caused by frequent turn annotation and sequence recognition, a corresponding annotation result sequence can be obtained for the annotation text to be analyzed at a fixed period, and the sequence recognition result of the annotation result sequence can be obtained according to the sequence detection rules.

[0129] Among them, the annotation text to be analyzed can be determined in any of the following ways: The dialogue annotation text before the current time and within a preset duration (such as 5 minutes) can be used as the annotation text to be analyzed. The current time refers to the time when the annotation result sequence is obtained. This method can ensure that the dialogue analysis takes into account the dialogue content in the recent period, facilitating the identification of ongoing dialogue types or trends. Alternatively, when there is a last detection in the historical detection and there is an annotation result sequence that meets the sequence detection rules of any preset dialogue type, all the annotation texts after the annotation result sequence that meets the sequence detection rules can be used as the text to be analyzed. This method can accurately locate the node where the dialogue type changes and use the dialogue text after this node as the text range to be analyzed, reducing redundant analysis and improving processing efficiency.

[0130] During the progress of the dialogue, the real-time received dialogue text is in a dynamically updated state. Since the dialogue has not ended, the dialogue text has the characteristic of incomplete information. Based on this, when the annotation result sequence obtained after annotating the turns in the dialogue text does not fully conform to the sequence detection rules of any preset dialogue type, that is, before the dialogue text reaches a complete state, it is impossible to accurately determine the dialogue type to which the annotation result sequence belongs. Therefore, regarding the determination of the sequence recognition result of the annotation result sequence according to the sequence detection rules of the preset dialogue type described in the foregoing step S104, as Figure 2 shown, the sequence recognition result can be determined at least in the following ways:

[0131] S201, compare the annotation result sequence corresponding to the annotation text with the sequence detection rules of each preset dialogue type respectively;

[0132] That is, for the sequence detection rules of each predefined conversation type, the labeled result sequence is compared with all the conditions included in the sequence detection rule. If the labeled result sequence meets all the conditions, it is determined that the labeled result sequence meets the sequence detection rule; otherwise, if the labeled result sequence does not meet at least one of the conditions in the sequence detection rule, it is determined that the labeled result sequence does not meet the sequence detection rule, and the labeled result sequence is continuously compared with other sequence detection rules one by one.

[0133] For example, the conversation types include A, B, and C, and the corresponding sequence detection rules are Z1, Z2, and Z3 respectively. Assume that Z1 includes: [meeting the following three conditions in the same topic: including at least one of invitation to elaborate in detail (ELI), invitation to reason (REI), and coordination invitation (CI); including at least one of elaboration in detail (EL) and explanation of reasons (RE); including asking questions (Q)]. Then, when the labeled result sequence meets all these three conditions, it is determined that the labeled result sequence meets the sequence detection rule Z1 of conversation type A.

[0134] S202, when the labeled result sequence meets any of the sequence detection rules, the predefined conversation type corresponding to the sequence detection rule that is met is determined as the conversation type to which the labeled result sequence belongs, and a first recognition result is generated.

[0135] For the sequence detection rules of multiple predefined conversation types, when the labeled result sequence meets any one of the sequence detection rules, the conversation type to which the sequence detection rule that the labeled result sequence meets belongs can be used as the conversation type to which the labeled result sequence belongs. Next, a first recognition result can be generated according to the conversation type to which the labeled result sequence belongs. In addition, when it is determined that the labeled result sequence meets the sequence detection rule, the comparison between the labeled result sequence and the sequence detection rules of other conversation types can be stopped to avoid wasting resources.

[0136] The first recognition result may include the conversation type or the identifier of the conversation type to which the labeled result sequence belongs, and may also include other additional information, such as the timestamp information of the start and end of the conversation corresponding to the labeled result sequence, the conversation goal corresponding to the conversation type to which the labeled result sequence belongs, etc.

[0137] S203, when the labeled result sequence does not meet the sequence detection rules of all predefined conversation types, it is determined that the conversation type to which it belongs indicates an unknown conversation type, and for each sequence detection rule, the rule information that the labeled result sequence does not meet is determined respectively, and a second recognition result is generated; wherein, the second recognition result may at least include the rule information that the labeled result sequence does not meet.

[0138] When the labeled result sequence fails to meet any of the sequence detection rules of the preset conversation types, it can be determined that the conversation type to which the labeled result sequence belongs is an unknown conversation type, which is used to identify the conversation type that cannot be classified, indicating that the current labeled result sequence has characteristics or patterns different from any preset conversation type.

[0139] Meanwhile, in order to deeply analyze why the labeled result sequence fails to match any preset conversation type, for each sequence detection rule of the preset conversation type, it is also possible to determine the rule information that the labeled result sequence does not meet, that is, to clarify the specific reason why the labeled result sequence does not meet the sequence detection rule. For example, for the sequence detection rule Z1 corresponding to the aforementioned conversation type A, if the labeled result sequence does not meet the condition "including asking question Q" in Z1, the determined rule information that the labeled result sequence does not meet is "Conversation type A: does not meet including asking question Q", which is used to clearly identify the specific condition that the labeled result sequence fails to meet under the sequence detection rule Z1 of conversation type A, providing data support for the diagnosis of the second text network and generating feedback information subsequently.

[0140] Based on the determined unknown conversation type and the rule information that the labeled result sequence does not meet under each conversation type, a second recognition result is constructed. The second recognition result cooperates with the first recognition result to fully cover and process the labeled result sequences that cannot be classified into the preset conversation types and the labeled result sequences that are classified into the known conversation types, increasing the comprehensiveness and pertinence of the feedback information for the subsequent processing of diagnosing the labeled text according to the sequence recognition result to generate feedback information, and realizing the comprehensive understanding of the conversation content and optimization suggestions.

[0141] In the embodiments of the present disclosure, through the sequence detection rules of the preset conversation types, the conversation type to which the labeled result sequence belongs is accurately judged. When the labeled result sequence meets a certain sequence detection rule, its belonging preset conversation type is determined and the first recognition result is generated. For the labeled result sequences that do not meet all the sequence detection rules of the preset conversation types, it is determined that they belong to the unknown conversation type, and the rule information that is not met is respectively determined for each sequence detection rule to generate the second recognition result. The sequence recognition result generated in the above manner provides a detailed diagnosis basis for the second text network to diagnose the labeled text, facilitating the second text network to deeply analyze the characteristics and patterns of the labeled result sequence, and improving the pertinence and accuracy of the generated feedback information.

[0142] Based on the sequence recognition result including the first recognition result and the second recognition result described in the above embodiments, in some embodiments, when the sequence recognition result is the first recognition result, for the diagnosis of the labeled text using the second text network according to the sequence recognition result in the aforementioned step S105, such as Figure 3The flowchart of the second text network diagnosis steps shown can achieve the diagnosis of the annotation result sequence belonging to the preset dialogue type through the following steps:

[0143] S301, when the sequence recognition result is the first recognition result, obtain the dialogue target of the preset dialogue type corresponding to the sequence detection rule satisfied;

[0144] That is, based on the first recognition result generated including the dialogue type to which the annotation result sequence belongs, the known dialogue type refers to the dialogue type to which the sequence detection rule satisfied by the annotation result sequence belongs, that is, classified as the preset dialogue type. Therefore, according to the corresponding relationship between the preset dialogue type and the dialogue target, the dialogue target of the dialogue type to which the annotation result belongs can be determined.

[0145] S302, according to the dialogue target, use the second text network to diagnose the annotated text.

[0146] The dialogue target can be used as auxiliary information and introduced into the diagnosis process of the second text network to guide the second text network to diagnose the annotated text, so that the second text network can evaluate whether the content, structure, grammar, etc. of the annotated text meet the expectations according to the dialogue target, deeply analyze the possible problems in the annotated text, and thus generate suggestions on adjusting the dialogue strategy to better meet the dialogue target, or other information as the feedback information generated after the second text network is diagnosed.

[0147] For example, taking the dialogue type - critical inquiry in the classroom teaching scenario as an example, when the annotation result sequence belongs to this critical inquiry type of dialogue, according to its dialogue target - cultivating students' critical thinking, the dialogue target can be introduced into the second text network in the form of a prompt text or other forms. The diagnosis of the annotated text by the second text network can include but is not limited to checking whether the key elements of critical inquiry are fully reflected in the dialogue, whether the dialogue has sufficient discourse functions such as questioning and reasoning, etc.

[0148] In the embodiments of the present disclosure, when the annotation result sequence is classified into the preset dialogue type, by introducing the dialogue target of the dialogue type to which the annotation result sequence belongs as auxiliary information into the diagnosis process of the second text network, the diagnosis becomes more targeted, so as to generate more dialogue adjustment suggestions or information that meet the dialogue target, making the dialogue more in line with the expectations and needs of users.

[0149] For the case where the sequence recognition result is the second recognition result, this embodiment provides a diagnosis method for indicating an unknown dialogue type for the dialogue type to which the annotation result sequence belongs, that is, for the diagnosis of the annotated text using the second text network according to the sequence recognition result described in the foregoing step S105, asFigure 4 The flowchart of the second text network diagnosis step shown can be implemented in the following manner when diagnosing the annotated text based on the second recognition result:

[0150] S401, when the sequence recognition result is the second recognition result, obtain the rule information included in the second recognition result;

[0151] S402, according to the rule information and the indicated unknown dialogue type of the belonging dialogue type, use the second text network to diagnose the annotated text.

[0152] That is, for the case where the annotation result sequence does not belong to the preset dialogue type, when generating the second recognition result, for each sequence detection rule, determine the rule information that the annotation result sequence does not satisfy. Based on this, in the process of diagnosing the annotated text according to the second recognition result, the rule information that the annotation result sequence does not satisfy can be introduced into the second text network as a prompt message, so that the second text network can give targeted dialogue adjustment suggestions as feedback information based on this prompt message.

[0153] In the process of using the second text network to diagnose the annotated text according to the rule information and the indicated unknown dialogue type of the belonging dialogue type, it is also possible to first determine the matching degree of each dialogue target in the dialogue context to which the annotated text and the dialogue audio belong through content analysis by a large language network according to the text content of the annotated text and the annotation result sequence; next, use the dialogue target that satisfies the highest matching degree, or the dialogue target that satisfies the highest matching degree and is greater than the set threshold as the designated target, and generate a prompt text as auxiliary information to guide the second text network to diagnose according to the rule information that the annotation result sequence does not satisfy under the preset dialogue type corresponding to the designated target. The finally generated feedback information can specifically prompt the dialogue guide to adjust along the direction of the preset dialogue type corresponding to the designated target.

[0154] For example, taking the preset dialogue types - critical inquiry, collaborative knowledge construction, teaching and instructional dialogue as examples, each preset dialogue type corresponds to a dialogue goal. In the case where the annotation result sequence corresponding to the annotated text does not conform to the sequence detection rules of each of the three preset dialogue types, a second recognition result is generated. When diagnosing the annotated text using the second text network, in-depth analysis of the text content of the annotated text and the annotation result sequence as text can be carried out to predict the dialogue goal that the annotated text most conforms to. For example, if the annotated text is most matched with the dialogue goal of cultivating critical thinking, then according to the rule information that is not satisfied under the sequence detection rules of critical inquiry type dialogues for the annotation result sequence, prompt information can be generated based on this rule information and combined with the annotated text as the input to the second text network, so as to facilitate the diagnosis of the second text network and the generation of feedback information.

[0155] In the embodiments of the present disclosure, by introducing the rule information that the annotation result sequence does not satisfy for each sequence detection rule as auxiliary information into the second text network and using the second text network to perform diagnosis in combination with the rule information, a more fine-grained diagnosis method is provided, making the diagnosis result more in line with user expectations and the generated feedback information more accurate.

[0156] In some embodiments, for the diagnosis of the annotated text using the second text network to generate feedback information according to the sequence recognition result in step S105 described above, it can also be implemented in the following manner: determining a target diagnosis method that matches the sequence recognition result according to a pre-constructed mapping relationship; the mapping relationship includes different recognition results and their corresponding diagnosis methods; the different recognition results include a result indicating that the annotation result sequence belongs to one of the preset dialogue types, and a result indicating that the annotation result sequence does not belong to any preset dialogue type; through the machine learning, performing sequence diagnosis on the annotated text according to the target diagnosis method.

[0157] In the embodiments of the present disclosure, by pre-constructing a mapping relationship to associate different recognition results with their corresponding diagnosis methods, flexible configuration of the diagnosis method is achieved. When it is necessary to add a new dialogue type or modify the diagnosis logic of an existing dialogue type, only the mapping relationship needs to be updated, thereby improving the scalability. At the same time, selecting the corresponding target diagnosis method according to the recognition result can ensure that the diagnosis process is more targeted, reduce unnecessary calculations and resource consumption, and improve the accuracy of diagnosis and feedback information generation.

[0158] In some embodiments, to further improve the detail and practicality of diagnosis, this embodiment introduces an innovative diagnostic dimension, namely the engagement level of the conversation participants, to enrich the diagnostic dimensions and methods of the second text network for the annotated text. Based on this diagnostic dimension, the foregoing conversation analysis method may further include the following steps of determining detection parameters, that is, obtaining video data collected simultaneously with the conversation audio; aligning the video data with the annotated text according to the time stamps; determining the engagement level of the conversation participants based on the video data and the annotated text; determining the detection parameters matching the engagement level of the conversation participants according to the pre-set corresponding relationship; the corresponding relationship includes the pre-set engagement levels and the detection parameters corresponding to each engagement level.

[0159] That is to say, while processing the conversation audio, video data synchronized with the audio can be collected by a video capture device such as a camera. By precisely aligning the video data with the annotated text according to the time stamp information, the consistency and accuracy of the analysis can be ensured. Based on this, video analysis technology and natural language processing technology can be used to comprehensively evaluate the behavior, expressions, and speech content of the conversation participants, evaluate the engagement of the conversation participants in the conversation context to which the conversation audio belongs, so as to determine the engagement level of the conversation participants. Next, based on the pre-set corresponding relationship, the detection parameters corresponding to each engagement level are clarified. The detection parameters are used to capture the key information related to the engagement of the conversation participants, and are used to analyze the reasons for the occurrence of this engagement level of the conversation participants. Different detection parameters are set for different engagement levels of the conversation participants, indicating different diagnostic focuses of the annotated text.

[0160] For example, taking the classroom teaching scenario as an example, the engagement level of classroom conversation participants can be divided into three levels: low engagement, medium engagement, and high engagement. Taking the low engagement level as an example, the detection parameters set may include, but are not limited to, whether the teacher explains for a long time, the silence mode and duration of the students, whether the teacher asks questions too frequently, etc.

[0161] Based on the detection parameters matching the engagement level of the conversation participants determined in the above manner, for the diagnosis of the annotated text using the second text network in step S105 above, a prompt message can be generated based on the sequence recognition result and the detection parameters. When using the second text network to diagnose the annotated text, the second text network can be guided by this prompt message.

[0162] In the embodiments of the present disclosure, by combining video data and annotated text, information during the conversation is comprehensively captured, the engagement level of the conversation participants is evaluated, and based on this engagement level, detection parameters matching this engagement level are introduced as the diagnostic dimensions for the second text network to diagnose the annotated text, thereby improving the accuracy of the diagnostic results and further generating feedback information that better meets the user's needs.

[0163] In some embodiments, after generating the feedback information in the foregoing step S105, the method may further include the following display step. Through this display step, the generated feedback information can be converted into a visual page for the user to view, so that the user can quickly grasp the key information of the conversation and make corresponding adjustments, thereby enhancing the user experience and interaction efficiency. This display step can be implemented in the following manner: outputting the feedback information in a structured mode; the feedback information may at least include at least one of the conversation participation situation, interaction prompts, and conversation direction suggestions; according to the preset page layout, converting the feedback information into a virtual page and projecting it for display within the user's line of sight.

[0164] Among them, the structured mode output is used to represent organizing the feedback information in a clear and orderly manner, and structured data in the form of tables, lists, or other forms can be used. In this embodiment, the feedback information is suggestions and information for guiding the conversation or improvement. The conversation participation situation can at least be used to describe information such as the activity level and contribution degree of each participant in the conversation. The interaction prompts may include guiding prompts or suggestions for enhancing the interactivity of the conversation, and the conversation direction suggestions are used to represent conversation direction prompts for the continuation of the conversation, such as asking more questions and interacting regarding topic A.

[0165] The preset page layout refers to the design layout of the page or interface for displaying information in a virtual environment, and this page layout can be customized according to user needs and system characteristics. Virtual page projection display refers to the technology of projecting information in the form of a page into the user's line of sight using virtual reality, augmented reality (AR), or other display technologies. After obtaining the structured feedback data, this feedback data can be filled into the corresponding positions of the virtual page to generate a complete virtual page and project it for display within the user's line of sight.

[0166] During the process of displaying the feedback information, the feedback information can be presented in the form of simple icons and text to enhance the readability of the information. Different colors or icons can also be used to indicate different feedback states, engagement levels, and interaction modes, providing conversation feedback without disturbing the conversation rhythm of the conversation leader.

[0167] In the embodiments of the present disclosure, in order to improve the user experience and interaction efficiency, after analyzing the real-time received dialogue audio to generate feedback information, the feedback information is converted into a virtual page according to a preset page layout and immediately projected and displayed within the user's line of sight, providing an intuitive and easy-to-read information display platform, reducing the difficulty for the user to understand the information, and enabling the user to receive the feedback information in real time while the dialogue is in progress and quickly adjust the dialogue strategy based on the feedback information to improve the dialogue quality.

[0168] Based on the dialogue analysis method provided in the foregoing embodiments, the present application also provides a pair of glasses capable of displaying feedback information in real time, which can collect dialogue audio in real time and generate feedback information for the dialogue audio through the foregoing dialogue analysis method, and use the glasses to display the feedback information in the user's line of sight, facilitating the user to quickly grasp the dialogue situation and adjust the dialogue.

[0169] As Figure 5 Exemplarily shown is a schematic diagram of the glasses structure. The glasses may at least include an audio collection module 501, a power management module 502, a data processing module 503, a wireless transmission module 504, and a display module 505. It should be understood that the positions and settings of the modules in the figure are only for exemplary display, and the present application does not make any limitations in this regard. For the glasses structure in the figure, wherein, the audio collection module is used to collect dialogue audio in real time; the power management module is used to provide the power required for the operation of the glasses; the wireless transmission module is used to transmit the dialogue audio collected by the audio collection module to the data processing module in real time; the data processing module is used to generate feedback information for the real-time received dialogue audio by using the dialogue analysis method provided in the foregoing embodiments, and send the feedback information to the display module through the wireless transmission module; the display module is used to project the feedback information within the user's line of sight based on augmented reality technology, and the display template can be embedded in the glasses lenses.

[0170] The audio collection module may include at least one of a built-in camera and a microphone. The camera is used to capture the dialogue environment in real time, and the microphone is used to collect the dialogue sound in real time. The audio collection module may further include a preprocessing module for preliminarily processing the collected audio and video data.

[0171] The display template may include a head-up display, which is respectively embedded in the two lenses of the glasses. The head-up display is used to display the feedback information within the user's line of sight through transparent projection technology.

[0172] Regarding the data processing module generating feedback information by using the dialogue analysis method provided in the foregoing embodiments, when the computing resources of the data processing module itself support running the dialogue analysis method, the data processing module runs the dialogue analysis method to process the real-time received dialogue audio, so as to generate feedback information; when the real-time received dialogue audio is processed by a server or other devices or apparatuses disposed at a remote end that support running the dialogue analysis method, a communication connection between the data processing module and the server or other devices or apparatuses disposed at the remote end that support running the dialogue analysis method can be established in advance, and the data processing module receives the returned feedback information and transmits it to the display module.

[0173] In the embodiments of the present disclosure, by integrating audio acquisition, power management, data processing, wireless transmission, and augmented reality display technologies into a pair of glasses, a highly integrated and real-time interactive dialogue feedback system and augmented reality glasses are provided for users. The built-in audio acquisition module of the glasses can capture dialogue audio in real time and quickly transmit it to the data processing module for analysis through the wireless transmission module. This real-time processing ability ensures that users can obtain instant feedback information of the dialogue without waiting or relying on external devices; the feedback information is instantaneously displayed through the display module embedded in the glasses lens, enabling users to directly obtain key information within their line of sight without being distracted to view other devices, greatly improving the convenience of use. Additionally, by using augmented reality technology, the display module can project the feedback information in the form of a virtual image within the user's line of sight, reducing the risk of sensitive information leakage. The display of the feedback information can seamlessly integrate with the user's real environment. The immersive visual experience not only enhances the intuitiveness of the information but also increases the user's sense of participation and interactivity. Based on the fact that all functions are integrated in the glasses, users can obtain the required feedback information without holding a device or performing complex operations, bringing convenience to user operations.

[0174] Corresponding to the embodiments of the foregoing dialogue analysis method, refer to Figure 6 As shown, the present application also provides an embodiment of a dialogue analysis device, and the device includes:

[0175] An audio conversion module 601, configured to convert real-time received dialogue audio into dialogue text;

[0176] A marking module 602, configured to mark each turn in the dialogue text by using a first text network to obtain a marked text with a marking result; the marking result is used to represent the discourse function type of the turn;

[0177] A sequence obtaining module 603, configured to obtain a marking result sequence according to the order and marking result of the turns in the marked text;

[0178] A sequence recognition module 604, configured to determine a sequence recognition result of the labeled result sequence according to a sequence detection rule of a preset conversation type; the sequence recognition result includes the conversation type to which the labeled result sequence belongs; the preset conversation type is determined according to different conversation objectives in the conversation scenario to which the conversation audio belongs;

[0179] A feedback information generation module 605, configured to diagnose the labeled text by using a second text network according to the sequence recognition result, and generate feedback information; the feedback information is at least used to prompt the conversation situation to the conversation leader in real time to optimize the conversation.

[0180] In some embodiments, the labeling module is specifically configured to:

[0181] Obtain classification conditions of each preset labeling tag; the preset labeling tag corresponds to the discourse function type of the turn in the conversation scenario one by one; construct a prompt text applicable to prompt engineering according to the classification conditions; use the first text network and the prompt text to identify and label the labeling tag to which each turn in the conversation text belongs.

[0182] In some embodiments, the labeling module is specifically configured to:

[0183] Input the text content in the turn into the first text network to obtain a labeling result output by the first text network; wherein, a plurality of classification tags are set in the output layer of the first text network, and the classification tags correspond to the discourse function type of the turn in the conversation scenario one by one; the labeling result is labeled according to the classification tag greater than a set threshold in the probability distribution of the plurality of classification tags generated by the first text network.

[0184] In some embodiments, the sequence recognition module is specifically configured to:

[0185] When the labeled result sequence satisfies any one of the sequence detection rules, determine the preset conversation type corresponding to the satisfied sequence detection rule as the conversation type to which the labeled result sequence belongs, and generate a first recognition result;

[0186] When the labeled result sequence does not satisfy the sequence detection rules of all preset conversation types, determine that the conversation type to which it belongs indicates an unknown conversation type, and respectively determine the rule information that the labeled result sequence does not satisfy for each sequence detection rule, and generate a second recognition result.

[0187] In some embodiments, the feedback information generation module is specifically configured to:

[0188] When the sequence recognition result is the first recognition result, obtain the dialogue target of the preset dialogue type corresponding to the satisfied sequence detection rule; according to the dialogue target, use the second text network to diagnose the annotated text.

[0189] In some embodiments, the feedback information generation module is specifically configured to:

[0190] When the sequence recognition result is the second recognition result, obtain the rule information included in the second recognition result; according to the rule information and the indicated unknown dialogue type of the dialogue type to which it belongs, use the second text network to diagnose the annotated text.

[0191] In some embodiments, the apparatus further includes:

[0192] Obtain video data collected simultaneously with the dialogue audio; align the video data and the annotated text according to timestamps; determine the dialogue participant engagement level according to the video data and the annotated text; according to the preset corresponding relationship, determine the detection parameters matching the dialogue participant engagement level; the corresponding relationship includes preset engagement levels and the detection parameters corresponding to each engagement level;

[0193] The feedback information generation module is specifically configured to: according to the sequence recognition result and the detection parameters, use the second text network to diagnose the annotated text.

[0194] In some embodiments, the apparatus further includes:

[0195] Output the feedback information in a structured mode; the feedback information may at least include at least one of dialogue participation status, interaction prompts, and dialogue direction suggestions; according to the preset page layout, convert the feedback information into a virtual page and project it for display within the user's line of sight.

[0196] The specific implementation processes of the functions and roles of each unit in the above apparatus are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative effort.

[0198] An embodiment of the present application further provides an electronic device. The schematic structural diagram of the electronic device is as shown in Figure 7 Figure Figure 7 . The electronic device 700 includes at least one processor 701, a memory 702, and a bus 703. At least one processor 701 is electrically connected to the memory 702. The memory 702 is configured to store at least one computer-executable instruction, and the processor 701 is configured to execute the at least one computer-executable instruction, so as to perform the steps of any one of the dialogue analysis methods provided in any one of the embodiments or any alternative implementation manners of the present application.

[0199] Further, the processor 701 may be an FPGA (Field-Programmable Gate Array), or other devices with logical processing capabilities, such as an MCU (Microcontroller Unit) or a CPU (Central Process Unit).

[0200] An embodiment of the present application further provides another readable storage medium storing a computer program, which is used to implement the steps of any one of the dialogue analysis methods provided in any one of the embodiments or any alternative implementation manners of the present application when being executed by a processor.

[0201] The readable storage medium provided by the embodiment of the present application includes, but is not limited to, any type of disk (including a floppy disk, a hard disk, an optical disk, a CD-ROM, and a magneto-optical disk), a ROM (Read-Only Memory), a RAM (Random Access Memory), an EPROM (Erasable Programmable Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a flash memory, a magnetic card, or an optical card. That is, the readable storage medium includes any medium that stores or transmits information in a form readable by a device (such as a computer).

[0202] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims may be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily the specific order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0203] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for dialogue analysis, characterized in that, The method includes: Converting the real-time received dialogue audio into dialogue text; Using a first text network to annotate each turn in the dialogue text to obtain an annotated text with annotation results; the annotation results are used to represent the discourse function type of the turn; Obtaining an annotation result sequence according to the order and annotation results of the turns in the annotated text; Determining a sequence recognition result of the annotation result sequence according to a sequence detection rule of a preset dialogue type; the sequence recognition result includes the dialogue type to which the annotation result sequence belongs; the preset dialogue type is determined respectively according to different dialogue goals in the dialogue scenario to which the dialogue audio belongs; Diagnosing the annotated text using a second text network according to the sequence recognition result to generate feedback information; the feedback information is at least used to prompt the dialogue situation to the dialogue leader in real time to optimize the dialogue.

2. The method according to claim 1, characterized in that, The using a first text network to annotate each turn in the dialogue text includes: Obtaining classification conditions for each preset annotation label; the preset annotation labels correspond one-to-one with the discourse function types of turns in the dialogue scenario; Constructing a prompt text applicable to prompt engineering according to the classification conditions; Using the first text network and the prompt text to identify and annotate the annotation labels to which each turn in the dialogue text belongs.

3. The method according to claim 1, characterized in that, The using a first text network to annotate each turn in the dialogue text includes: Inputting the text content in the turn into the first text network to obtain an annotation result output by the first text network; Wherein, multiple classification labels are set in the output layer of the first text network, and the classification labels correspond one-to-one with the discourse function types of turns in the dialogue scenario; the annotation result is obtained by annotating the classification label greater than a set threshold in the probability distribution of the multiple classification labels generated by the first text network.

4. The method according to claim 1, wherein The determining a sequence recognition result of the annotation result sequence according to a sequence detection rule of a preset dialogue type includes: When the annotation result sequence satisfies any one of the sequence detection rules, determining the preset dialogue type corresponding to the satisfied sequence detection rule as the dialogue type to which the annotation result sequence belongs, and generating a first recognition result; When the annotation result sequence does not satisfy the sequence detection rules of all preset dialogue types, determining that the indicated dialogue type belongs to an unknown dialogue type, and respectively determining the rule information that the annotation result sequence does not satisfy for each sequence detection rule, and generating a second recognition result.

5. The method according to claim 4, characterized in that, The diagnosing the annotated text using a second text network according to the sequence recognition result includes: When the sequence recognition result is the first recognition result, obtaining the dialogue goal of the preset dialogue type corresponding to the satisfied sequence detection rule; Diagnosing the annotated text using the second text network according to the dialogue goal.

6. The method according to claim 4, wherein The diagnosing the annotated text using a second text network according to the sequence recognition result includes: When the sequence recognition result is the second recognition result, obtaining the rule information included in the second recognition result; Indicating an unknown dialogue type according to the rule information and the dialogue type to which it belongs, and diagnosing the annotated text by using the second text network.

7. The method according to claim 1, characterized in that, The method further includes: Obtaining video data collected simultaneously with the dialogue audio; Aligning the video data and the annotated text according to time stamps; Determining the participation level of the dialogue participants according to the video data and the annotated text; Determining a detection parameter matching the participation level of the dialogue participants according to a preset corresponding relationship; the corresponding relationship includes preset participation levels and detection parameters corresponding to each participation level; The diagnosing the annotated text by using the second text network according to the sequence recognition result includes: Diagnosing the annotated text by using the second text network according to the sequence recognition result and the detection parameter.

8. The method according to claim 1, characterized in that, The method further includes: Outputting the feedback information in a structured mode; the feedback information can at least include at least one of dialogue participation situation, interaction prompt, and dialogue direction suggestion; Converting the feedback information into a virtual page according to a preset page layout and projecting it within the user's line of sight.

9. A pair of glasses, characterized in that, The glasses at least include an audio collection module, a power management module, a data processing module, a wireless transmission module, and a display module; The audio collection module is used to collect dialogue audio in real time; The power management module is used to provide power required for the operation of the glasses; The wireless transmission module is used to transmit the dialogue audio collected by the audio collection module to the data processing module in real time; The data processing module is used to generate feedback information for the dialogue audio received in real time by using the dialogue analysis method according to any one of claims 1-8, and send the feedback information to the display module through the wireless transmission module; The display module is used to project the feedback information within the user's line of sight based on augmented reality technology.

10. A dialogue analysis device, characterized in that, The device includes: An audio conversion module, configured to convert the dialogue audio received in real time into dialogue text; A marking module, configured to mark each turn in the dialogue text by using a first text network to obtain an annotated text with a marking result; the marking result is used to represent the discourse function type of the turn; A sequence obtaining module, configured to obtain a marking result sequence according to the order and marking result of the turns in the annotated text; A sequence recognition module, configured to determine a sequence recognition result of the marking result sequence according to a sequence detection rule of a preset dialogue type; the sequence recognition result includes the dialogue type to which the marking result sequence belongs; the preset dialogue type is determined respectively according to different dialogue objectives in the dialogue scenario to which the dialogue audio belongs; A feedback information generation module, configured to diagnose the annotated text by using a second text network according to the sequence recognition result, and generate feedback information; the feedback information is at least used to prompt the dialogue situation to the dialogue leader in real time to optimize the dialogue.

11. An electronic device, characterized in that, Includes: A memory and a processor; The memory is used to store a computer program; The processor is used to call the computer program to implement the method according to any one of claims 1-8.

12. A readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent voice dialogue scene verbal skill intervention method and system based on customer portrait

    CN116049360A

  • Artificial intelligence character models with goal-oriented behavior

    US20230351661A1