Conversation method and device based on artificial intelligence, equipment and storage medium

By receiving voice signals and distinguishing the speakers, and generating targeted dialogue auxiliary content, the problem that existing dialogue tools cannot analyze complex dialogue scenarios in real time is solved, efficient and accurate dialogue support is achieved, and user communication effect is improved.

CN120496515APending Publication Date: 2025-08-15谭伟良
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510571537.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-04
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing dialogue tools cannot analyze conversations in real time and provide conversation assistance to users when they are talking to others, especially in complex scenarios that cannot dynamically generate targeted response suggestions.

Method used

By receiving voice signals, distinguishing the speaker, judging that the speaker belonging to the voice signals is the user or conversation object, generate dialogue auxiliary content, including reference speech, alternative speech and speech auxiliary information, and combining non-dialogue information to provide personalized dialogue support.

Benefits of technology

It improves the dialogue effect of users in scenarios such as business negotiations and social communication, solves communication barriers such as cold situations and misunderstandings, provides efficient, accurate and intelligent dialogue assistance, and improves the fluency and accuracy of dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496515A_ABST
    Figure CN120496515A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent human-computer interaction, and provides a dialogue method, device and equipment based on artificial intelligence and a storage medium. The problem that an existing dialogue tool is difficult to analyze in real time in a complex dialogue scene and assist a user is solved. The dialogue method based on artificial intelligence provided by the invention comprises the following steps: receiving a voice signal; distinguishing all or part of speakers according to the voice signals; an effective speech is obtained by judging whether a speaker is a user or a conversation object or not; associating the effective speech with the identification information; triggering and generating corresponding contents according to specific conditions; and generating dialogue auxiliary content according to the input information. Through real-time analysis of the dialogue content, speaker roles are distinguished, non-dialogue information is integrated, and targeted dialogue auxiliary support can be provided for the user. Meanwhile, complex dialogue scenes such as cold fields and interruption can be dynamically processed, dialogue interruption is avoided, and dialogue fluency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent human-computer interaction technology, and in particular to an artificial intelligence-based dialogue method, device, equipment and storage medium. Background Art

[0002] In daily life, people often face various conversation scenarios, such as business negotiations, social exchanges, and emotional communication. However, during these conversations, users may experience poor results due to nervousness, lack of experience, or insufficient language skills. For example, there may be awkward silences, misunderstandings, a lack of communication, and an inability to effectively express one's intentions.

[0003] With the development of artificial intelligence technology, chatbot programs, represented by ChatGPT, have brought significant breakthroughs in the field of intelligent human-computer interaction. These chatbot programs can be trained based on large amounts of text data, learning language patterns and structures, and thus generating natural and fluent responses. Users can interact with chatbots through text or voice to obtain information, suggestions, or answers related to their input. To a certain extent, these tools can help users answer questions and provide suggestions. However, in the process of realizing the present invention, the inventors discovered that at least these problems exist in the related art: existing conversation tools (such as chat assistants and chatbots) mainly adopt a "one-to-one interaction between user and tool" model. Their core functions are limited to answering user questions or executing commands, and the responses they generate are only specific to the user's input (such as questions or commands). Although these tools can provide personalized suggestions, they are unable to analyze third-party speech in real time in complex scenarios where users are conversing with others, and are unable to assist users in two-way conversations with third parties in real time, such as dynamically generating response suggestions in real-time exchanges with friends, colleagues, or customers. Summary of the Invention

[0004] The present invention provides an artificial intelligence-based conversation method, apparatus, device, and storage medium, which solves the problem in related technologies that conversation tools cannot analyze conversations in real time and provide conversation assistance to users when users are talking to other people.

[0005] In one aspect, the present invention provides an artificial intelligence-based dialogue method, comprising:

[0006] receiving a voice signal;

[0007] Distinguish some or all speakers to whom the received speech signal belongs;

[0008] Determine whether the speaker in the voice signal is the user or the conversation partner, and identify the speech and / or speech-derived information of the speaker as the user or the conversation partner based on the determination result, thereby obtaining a valid speech. The user speech, the conversation partner's speech, the user's speech-derived information, and the conversation partner's speech-derived information are collectively referred to as a valid speech, and the speech-derived information is derived from and related to the speech.

[0009] Associating the valid speech with the speaker through identification information;

[0010] Determine whether a specific condition is met, and if so, generate the dialogue assistance content; otherwise, do not generate the dialogue assistance content;

[0011] Conversation assistance content is generated based on input information, the conversation assistance content includes at least one of a reference speech, an alternative speech, and speech assistance information, the alternative speech is played as voice, the input information includes at least one of the valid speech, the identification information, previously generated conversation assistance content, and non-conversation information, and the non-conversation information is independent of the speaker's speech.

[0012] In a second aspect, the present invention provides an artificial intelligence-based dialogue device, comprising:

[0013] Voice acquisition module: used to receive voice signals;

[0014] Speaker differentiation module: used to distinguish some or all speakers of the received speech signal;

[0015] Speech screening module: determines whether the speaker in the voice signal is the user or the conversation partner, and identifies the speech and / or speech-derived information of the user or the conversation partner based on the judgment result, thereby obtaining valid speeches. The user speech, the conversation partner's speech, the user's speech-derived information, and the conversation partner's speech-derived information are collectively referred to as valid speeches. The speech-derived information is derived from and related to the speech.

[0016] Identification association module: associates the valid speech with the speaker through identification information;

[0017] Trigger generation detection module: determines whether specific conditions are met. If so, the dialogue assistance module is triggered to generate corresponding content. Otherwise, the dialogue assistance module is not triggered at this time.

[0018] Dialogue assistance module: used to generate dialogue assistance content based on input information, the dialogue assistance content includes at least one of a reference speech, an alternative speech, and speech assistance information, the alternative speech is played as voice, the input information includes at least one of the valid speech, the identification information, previously generated dialogue assistance content, and non-dialogue information, the non-dialogue information is independent of the speaker's speech.

[0019] Furthermore, the dialogue device also includes: an error correction module, which is used to detect whether the user or the conversation partner has spoken when the dialogue assistance module generates a reference speech. If, when generating a reference speech, it is detected that the user has not spoken and the conversation partner has spoken continuously for a period of time exceeding a certain threshold since the trigger generation detection module determines that the condition is met, the dialogue assistance module is controlled to pause outputting content, and a new reference speech is triggered through the trigger generation detection module.

[0020] Furthermore, the dialogue device also includes: a cold scene processing module, which is used to determine whether the conversation partner is silent or has continuously output speeches without substantive content for more than a preset number of times. If it is determined to be so, a reference speech and / or speech auxiliary information is generated to guide the continuation of the dialogue to avoid interruption of the dialogue.

[0021] Furthermore, the dialogue device also includes: a non-dialogue information acquisition module, which is used to obtain non-dialogue information in a specific manner at a specific time and input it into the dialogue assistance module. The specific manner includes manual addition, online search, and recommendation by the dialogue device. The non-dialogue information is independent of the speaker's speech, and the specific time is determined by the user or the dialogue device.

[0022] Furthermore, the dialogue device also includes: an interruption processing module. During the period when the dialogue assistance module outputs content, if the user's speech is interrupted by the conversation partner, and the duration of the conversation partner's continuous speech from the moment of interruption exceeds a certain threshold, the dialogue assistance module is controlled to pause the output of content, and a new reference speech is triggered by the trigger generation detection module.

[0023] Furthermore, the dialogue device also includes: an irrelevant dialogue filtering module, which is used to determine whether the valid speech is a dialogue between an irrelevant person and a user or a conversation partner. If it is determined to be so, the valid speech is filtered out before the input information enters the dialogue assistance module. The irrelevant person refers to other persons who are not users and conversation partners.

[0024] Furthermore, the dialogue device further includes: a dialogue parameter setting module, which is used to set dialogue parameters, and the dialogue parameters are configuration items used to guide and optimize the dialogue assistance content.

[0025] In a third aspect, the present invention provides a device comprising: a memory and one or more processors; wherein the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the device executes the artificial intelligence-based dialogue method as described in claim 1.

[0026] In a fourth aspect, the present invention provides a storage medium storing a computer program, which, when executed on a computer, enables the computer to execute the artificial intelligence-based dialogue method as claimed in claim 1.

[0027] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0028] It provides users with targeted reference speeches in real time based on the current conversation, solving the problem in related technologies that conversation tools cannot analyze conversations in real time and provide conversation assistance to users when users are talking to others. It effectively solves communication barriers such as awkward silences, misunderstandings, and tension, and significantly improves the user's conversation effect in various scenarios such as business negotiations, social exchanges, and emotional communication. It helps users communicate more confidently in different scenarios and provides users with an efficient, accurate, intelligent, and personalized conversation assistance tool.

[0029] By distinguishing the speaker, the system can accurately identify whether the current speaker is the user, the conversation partner, or an unrelated person, providing a clear role division for subsequent conversation processing. This distinction helps avoid confusion about the speaker's identity and improves the accuracy of conversation processing.

[0030] By distinguishing the speakers, the system can filter out the speeches of irrelevant people and form effective speeches, ensuring the effective classification and archiving of the conversation content, avoiding the speeches of irrelevant people interfering with the conversation, and reducing the burden of storage and processing.

[0031] By judging the timing of generating reference speeches, the system ensures that reference speeches are generated at the appropriate time, avoiding generating disruptive speeches while the other party is still speaking, and improving the fluency and naturalness of the conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0033] Figure 1 This is a flow chart of an artificial intelligence-based dialogue method in an embodiment of the present invention;

[0034] Figure 2 Schematic diagram of internal functional modules of an artificial intelligence-based dialogue device in an embodiment of the present invention;

[0035] Figure 3 Schematic diagram of the structure of the device in an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following description sets forth many specific details to facilitate a full understanding of the present invention. The embodiments described are only a portion of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] In the description of the present invention, it should be noted that the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0038] The term "and / or" appearing in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0039] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0040] Example 1:

[0041] In this embodiment, a dialogue method based on artificial intelligence is provided, such as Figure 1 As shown, the process of this method includes the following steps:

[0042] Step S102: receiving a voice signal;

[0043] In this embodiment, the voice signal is acquired by a sound pickup device configured in an electronic device such as a mobile phone, a computer, a tablet computer, or a wearable device.

[0044] Step S104: distinguishing some or all speakers to which the received voice signal belongs;

[0045] In this embodiment, this step uses voiceprint feature extraction technology to distinguish different speakers. When using this technology, the received audio needs to be segmented and voiceprint features extracted for each audio segment (each audio segment is usually composed of several speech frames). Therefore, each audio segment has a corresponding voiceprint feature. The similarity between the voiceprint features of different speakers (such as cosine similarity, hereinafter referred to as similarity) is relatively small. Different speakers can be distinguished by similarity. If the similarity is greater than a certain threshold, it is the same speaker; otherwise, they are different speakers.

[0046] It's worth noting that not all speakers need to be distinguished. For example, in a two-person conversation, with Zhang San and Li Si, the system will focus on the audio of the designated speaker and will not distinguish the speaker identities of unrelated speakers. Suppose the system receives audio from unrelated speakers Wang Wu and Chen Liu during the conversation. As described above, the audio signals are processed in segments. For example, at a certain moment, the system receives four audio segments, s1, s2, s3, and s4. The voiceprint feature vectors extracted from these four audio segments are v1, v2, v3, and v4, respectively. Suppose the voiceprint database already contains the voiceprint feature vectors of the current speakers, Zhang San and Li Si, u1 and u2, respectively. Let F(u,v) denote the similarity between two feature vectors u and v, with a similarity threshold of F0 = 0.7. The calculated similarities between v1 and u1 and u2 are: F(v1,u1) = 0.87 > F0, and F(v1,u2) = 0.65 < F0, respectively. The similarities between v2 and u1 and u2 are: F(v2,u1)=0.58<F0, F(v2,u2)=0.85>F0, so we can determine that the speakers of s1 and s2 are Zhang San and Li Si respectively. The similarities between v3 and u1 and u2 are: F(v3,u1)=0.51<F0, F(v3,u2)=0.42<F0, and the similarities between v4 and u1 and u2 are: F(v4,u1)=0.48<F0, F(v4,u2)=0.32<F0, so we can know that the speakers of s3 and s4 are neither Zhang San nor Li Si. Both are audio clips unrelated to the current conversation. As for whether they are from the same speaker, we don't need to distinguish them. We don't need to know who the specific speakers of these two audio clips are. Whether the speaker is Zhang San or Li Si does not affect the operation of the system.

[0047] The technical design of this step is motivated by the fact that distinguishing the speaker's identity in conversations between users is crucial for conversation management. Without the ability to distinguish between the user and the person being spoken to, the system would be unable to accurately record the conversation and generate targeted reference speech.

[0048] This step has the beneficial effect of distinguishing the speaker in the speech signal, providing a clear role division for subsequent dialogue processing. This distinction helps avoid confusion about the speaker's identity and improves the relevance and accuracy of dialogue processing.

[0049] Step S108: Determine whether the speaker in the voice signal is the user or the conversation partner, and identify the speech and / or speech-derived information of the speaker as the user or the conversation partner based on the judgment result, thereby obtaining a valid speech. The user speech, the conversation partner's speech, the user's speech-derived information, and the conversation partner's speech-derived information are collectively referred to as a valid speech, and the speech-derived information is derived from the speech and is related to the speech.

[0050] In this step, the speech signal is received continuously and may contain audio from multiple speakers. In order to facilitate the classification and archiving of the conversation content, the audio from a single speaker needs to be separated. To this end, it is necessary to determine when the speaker changes.

[0051] First, the voiceprint feature of each small audio segment is obtained through voiceprint feature extraction technology (each small audio segment is usually composed of several speech frames), and then the similarity of the voiceprint feature vectors of two adjacent signal segments is calculated and compared with a certain threshold to determine whether they come from the same speaker. Assuming that the similarity of two adjacent audio segments at time T1 and T2 is less than the threshold, it means that the speaker has changed at time t = (T1 + T2) / 2. Through this method, all the times t when the speaker changes can be determined n , n=1,2,3,.... Then, when the speaker changes, the received continuous audio signal x(t) can be cut to obtain the audio s of a single speaker n (t), where s n (t) = x(t) | t n≤ t ≤ t n+1 Then the continuous audio signal x(t) is cut to get the audio of a single speaker s n (t) is the speech mentioned above, which can be in the form of voice or text. n (t) The text information obtained after language recognition can also be called a speech. As can be seen, a speech has only one speaker.

[0052] Speech derivative information is related information generated based on the speaker's speech content through editing, refining, summarizing, converting and other processing methods. Its content is relevant to the original speech. For example, in order to make the conversation content more concise and clear and reduce redundant content, typos, slips of the tongue, modal particles and other content in the speech can be deleted to form speech derivative information. Suppose the speech is: "Um, I, um, I plan to travel to Beijing tomorrow. Um, can you recommend some fun places to me?", which can be converted into speech derivative information: "I plan to travel to Beijing tomorrow. Can you recommend some fun places to me?". If c n (t) is the value of s n (t) The audio obtained after denoising, enhancement, etc., then c n (t) is also speech-derived information.

[0053] In another embodiment, the effective speech is a summary extracted from the original speech to reduce the number of characters of the context dialogue input to the artificial intelligence model.

[0054] In this embodiment, the calculation of each audio segment s n The similarity between the voiceprint features of the audio signal (t) and all feature values in the first template library is calculated. The first template library is used to store the voiceprint features of all speakers in the current conversation. Before the conversation, only the user's voiceprint features are stored, and the original state is restored after the conversation ends. If all of the above similarities are less than a certain threshold, the voiceprint features of the audio signal are added to the first template library, and a new speaker has appeared. The threshold can be pre-set or dynamically adjusted based on the ambient noise level. For example, the threshold is 0.85 in a quiet environment and 0.75 in a noisy environment.

[0055] In this embodiment, each time a new speaker appears, the system prompts the user to set a setting or the system independently determines whether the speaker is a conversation partner. If so, the new speaker's voiceprint features are added to the second template library. This second template library is used to store the voiceprint features of the current user and conversation partner, as well as the corresponding identification information. Before the conversation, the second template library only stores the user's voiceprint features and identification information, and is restored to its original state after the conversation ends.

[0056] In this step, from a certain audio s n(t) Extract the voiceprint feature and calculate the similarity between the feature and all the voiceprint features in the second template library. If all the similarities do not exceed a certain pre-set threshold, it means that the speaker of this audio segment is an irrelevant person, neither the user nor the conversation partner, and his speech is irrelevant to the current conversation. If the calculated maximum similarity is greater than the above threshold, it means that the speaker is the user or the conversation partner. Therefore, we can distinguish which speeches and / or speech-derived information obtained from the voice signal belong to irrelevant persons and which belong to users or conversation partners. In this way, speeches and / or speech-derived information belonging to users or conversation partners can be screened out to form valid speeches. Valid speeches will be input into the artificial intelligence model in subsequent steps, while invalid speeches will not be input into the artificial intelligence model.

[0057] For example: Assume that there are two people talking to each other, Zhang San and Li Si, and the threshold is set to 0.7. Then the second template library stores the voiceprint features of the user, Zhang San, and Li Si. At a certain moment, the system obtains three audio segments of a single speaker after sampling, which are The sampling points of the three audio segments are L1, L2, L3, x n Represents the nth sampling point in the audio signal. After denoising and enhancing these audio segments, we can get the audio segments The three audio clips contain the following utterances: "Where did you go for the May Day holiday?", "I went to Zhangjiajie," and "Was it fun there?" The maximum similarities between the voiceprint features of these three audio clips, c1, c2, and c3, and all the voiceprint features in the second template library are calculated to be 0.8, 0.85, and 0.4, respectively. This indicates that the speaker in c1 and c2 is the user or conversation partner, while the speaker in c3 is an unrelated person. Therefore, the valid utterances selected are c1, c2, "Where did you go for the May Day holiday?", and "I went to Zhangjiajie."

[0058] The technical design of this step is motivated by the fact that if all speakers' speeches are indiscriminately fed into the AI model to generate reference speech, irrelevant speakers present in the conversation will cause unnecessary interference. Therefore, the system needs to determine whether the speaker is the user or the conversation partner in order to filter out relevant speech.

[0059] The beneficial effect of this step is that the identities of the user, the conversation partner, and irrelevant persons can be accurately distinguished through the speaker's voice signal, thereby filtering out the speeches of irrelevant persons, forming effective speeches, ensuring the effective classification and archiving of the conversation content, and avoiding the speeches of irrelevant persons interfering with the conversation.

[0060] Step S110: Associating the valid speech with the speaker through identification information.

[0061] In this embodiment, after the above steps, we extract the voiceprint features of each audio segment and determine whether the speaker is the user or the conversation partner, and filter out valid speeches. In this step, for each audio segment corresponding to a valid speech, a speaker recognition operation will be performed.

[0062] For each audio segment corresponding to a valid speech, the similarity between its voiceprint feature and all features in the second template library is calculated. If the similarity corresponding to a feature is the largest and exceeds a certain threshold, the identification information of the valid speech is the identification information corresponding to that feature.

[0063] Based on step S108, each time a new speaker appears, the system prompts the user to set or the system independently determines whether the speaker is a conversation partner. If the speaker is a conversation partner, the user is prompted again to set the speaker's identification information or the system automatically assigns identification information. If the speaker is a conversation partner, the new speaker's voiceprint characteristics and identification information are added to the second template library. The new speaker's valid speech will be the currently set or assigned identification information.

[0064] The identification information is one or more sequences mainly composed of words, letters, numbers, and punctuation marks. For example, if three valid speeches have been screened out, their corresponding identification information is customer 1, customer 2, and customer 3.

[0065] After this step, we identify the speaker to whom each valid speech belongs. The identification information plays the role of determining the speaker to whom each valid speech belongs.

[0066] In another embodiment, the order of the above steps S104, S108, and S110 can be changed. First, the voiceprint features of each audio segment are extracted, and then the speaker of each audio segment is identified. Then, the speech and / or speech-derived information of the speaker being the user or the conversation partner are screened out to form a valid speech.

[0067] The beneficial effect of this step is that the system can trace the ownership of valid speeches, ensuring the traceability of conversation records. This mechanism makes it easier for the system to quickly locate the specific speaker when searching the conversation history.

[0068] Step S114: Determine whether a specific condition is met. If so, execute step S116; otherwise, do not execute the dialogue assistance step.

[0069] In this embodiment, the specific conditions include the termination of the conversation partner's speech, the receipt of a specific command signal from the user, the semantic analysis result of the valid speech, and the detection of at least one of specific parameter changes by the system. In order to detect the termination of speech, the system uses acoustic features and / or semantic logic to accurately determine whether the other party has finished speaking to avoid interruptions or omissions. In terms of acoustic features, a voice activity detection (VAD) algorithm is used to detect voice activity and distinguish between speech and noise. When silence is detected for a continuous period of time, it can be considered that the speech has ended. An appropriate silence detection threshold (such as 500 to 700 milliseconds) can be set to avoid misjudgments caused by short pauses such as breathing. In terms of semantic logic, a list of sentence-end keywords is defined, such as "um", "okay", "thank you", etc., and these keywords are matched to determine whether the speech has ended.

[0070] If there is only one person in the conversation, a reference speech can be generated after that person stops speaking. If there is more than one person in the conversation, the user can choose when to send a specific command signal to trigger the system to generate a reference speech. The specific command signal is a signal that can convey the user's desire for the system to generate a reference speech, such as pressing a button on the system interface or the sound of tapping a mobile phone screen.

[0071] In another specific embodiment, the system may be configured to only detect whether the user issues a specific instruction signal, and the system is triggered to generate a reference speech after the user issues the specific instruction signal.

[0072] The motivation for this step is that the system needs to determine the appropriate time to generate reference utterances during a conversation. Generating reference utterances while the conversation partner is still speaking can lead to errors or incoherent conversations.

[0073] The beneficial effect of this step is that by judging the timing of generating reference speeches, it is possible to avoid generating reference speeches at inappropriate times, thereby improving the response efficiency of the system and user experience.

[0074] Step S116: Generate dialogue assistance content based on the input information, the dialogue assistance content includes at least one of a reference speech, an alternative speech, and speech assistance information, the alternative speech is played as voice, the input information includes at least one of the valid speech, the identification information, previously generated dialogue assistance content, and non-dialogue information, and the non-dialogue information is independent of the speaker's speech.

[0075] In this embodiment, a large language model or a multimodal large model is used as an artificial intelligence model to generate content. These models can generate conversation-assisted content that conforms to the context based on the input current speech, the context of the effective speech, non-conversation information, and other additional information. The conversation-assisted content includes at least one of a reference speech and speech-assisted information. The conversation-assisted content can be in text or voice form. The reference speech serves as a reference for the user's speech, and the user can make selective speeches based on it. In non-face-to-face communication scenarios (such as calling to book a hotel), alternative speeches can be generated according to the user's actual needs, and the alternative speeches are played as voice to replace the user's speech.

[0076] This speech-support information refers to additional, relevant information, suggestions, or tips, as well as conversation analysis conclusions, provided during a conversation to help users better express themselves, understand, or advance the conversation. For example, in formal situations like business negotiations or interviews, the system can offer tips on presentation techniques, such as "use polite language, speak at a moderate speed, and avoid overly absolute terms" and "use examples to enhance the credibility of your ideas," to help users improve their professionalism and approachability.

[0077] In another specific embodiment, the system can analyze the conversation partner's voice, intonation, and word choice to determine their emotional state and provide this information as auxiliary prompts to the user. If the conversation partner expresses impatience, the system can remind the user to adjust their tone and expression appropriately to ease the other party's emotions and avoid escalating the conflict.

[0078] Non-dialogue information refers to information independent of the speaker's speech. This information does not directly originate from the conversation process, but rather is obtained from other channels or pre-set content. It is used to assist the conversation, provide context, or enrich the conversation. For example, when a user speaks and the conversation partner is silent, the system generates non-dialogue information without user intervention to prevent the conversation from becoming awkward. This information is then fed into the large language model or multimodal large model to generate reference speech to guide the conversation forward, ensuring an uninterrupted conversation. For example, the non-dialogue information in this case might read, "The conversation partner has been detected to be silent. Please provide a new chat topic to guide the conversation forward."

[0079] The technical design of this step is motivated by the fact that in complex or stressful conversations, users may struggle to quickly organize their words or accurately express their intentions. The system analyzes the conversation from the user's perspective and generates reference statements that are closely aligned with the current conversation. This helps users quickly find appropriate expressions, reduces their psychological burden, and improves the fluency and efficiency of conversations.

[0080] This step has the beneficial effect of generating reference speeches, providing users with a reference for their speech, helping them participate in conversations more efficiently. The generated reference speech incorporates the current speech, the context of valid speeches, and non-conversational information, making it more relevant to the current conversation, improving user experience and conversation efficiency. For example, in business scenarios, the reference speech should be formal and professional; in emotional communication scenarios, the reference speech should focus on emotional support and resonance.

[0081] To sum up, the beneficial effects of the above technical solution are: the system provides users with reference speeches that are in line with the current dialogue scenario based on non-dialogue information, the user's actual speech and the speech of the conversation partner, solving the problem that existing dialogue tools are unable to analyze dialogues in real time and provide dialogue assistance to users when users are talking with others. It effectively avoids communication barriers caused by factors such as awkward silences, misunderstandings, and tension, significantly improves the dialogue effects of users in various scenarios such as business negotiations, social exchanges, and emotional communication, helps users communicate more confidently in different scenarios, and provides users with an efficient, accurate, intelligent, and personalized dialogue assistance tool.

[0082] Example 2:

[0083] In the embodiment of the present invention, a dialogue device based on artificial intelligence is provided, such as Figure 2 ,include:

[0084] Voice collection module 202: used to receive voice signals. For technical details of this module, please refer to step S102 of Example 1.

[0085] Speaker identification module 204: used to identify some or all speakers of the received speech signal. For technical details of this module, please refer to step S104 of embodiment 1.

[0086] Speech screening module 206 determines whether the speaker in the speech signal is the user or the conversation partner. Based on the determination result, it identifies the speech and / or speech-derived information of the user or the conversation partner, thereby obtaining valid utterances. User utterances, conversation partner utterances, user utterance-derived information, and conversation partner utterance-derived information are collectively referred to as valid utterances. The speech-derived information originates from and is related to the utterance. For technical details of this module, see step S108 of Example 1.

[0087] Identification association module 208: associates the valid speech with the speaker through identification information. For technical details of this module, see step S110 of embodiment 1.

[0088] The trigger generation detection module 212 determines whether a specific condition is met. If so, the dialogue assistance module is triggered to generate the corresponding content. Otherwise, the dialogue assistance module is not triggered. For technical details of this module, see step S114 of embodiment 1.

[0089] Conversation assistance module 214 is configured to generate conversation assistance content based on input information. The conversation assistance content includes at least one of a reference speech, an alternative speech, and speech assistance information. The alternative speech is played as audio. The input information includes at least one of the valid speech, the identification information, previously generated conversation assistance content, and non-conversation information. The non-conversation information is independent of the speaker's speech. For technical details of this module, see step S116 of Example 1.

[0090] Example 3:

[0091] Based on Example 2, the artificial intelligence-based dialogue device further includes:

[0092] The misjudgment correction module 216 is used to detect whether the user or the conversation partner has spoken when the dialogue assistance module generates a reference speech. If, when generating a reference speech, it is detected that the user has not spoken and the conversation partner has spoken continuously for a period of time exceeding a certain threshold since the trigger generation detection module determines that the condition is met, the dialogue assistance module is controlled to pause outputting content and a new reference speech is triggered by the trigger generation detection module.

[0093] In this embodiment, while generating reference utterances, the system uses a voice activity detection (VAD) algorithm to analyze features such as the energy and zero-crossing rate of the speech signal to determine whether there is currently any speech activity. When the energy of the speech signal exceeds a threshold, it is determined to be a speech; otherwise, it is determined to be a speechless speech. Speaker recognition technology is then used to determine whether the current speech is from the user or the conversation partner. If the user is not speaking and the conversation partner has been speaking continuously for a period exceeding a certain threshold since the trigger generation detection module determined that the speech was valid, indicating that the timing for generating a reference utterance has been misjudged, an interrupt signal is sent to the conversation assistance module, immediately halting the current reference utterance generation process and triggering a new reference utterance through the trigger generation detection module.

[0094] The motivation for designing this module is that in actual conversations, there may be situations where a speech is mistakenly judged to be terminated (for example, the conversation partner is still speaking, but the system or user mistakenly believes that the speech has been terminated). When the conversation is mistakenly judged to be terminated, the currently generated reference speech is invalid. If reference speeches are continued to be generated, it will cause unnecessary waste of system resources.

[0095] The beneficial effect of this module is: the misjudgment correction module detects the speaking status of the user or conversation partner in real time, ensuring that the system can promptly terminate and re-receive the correct speech when a misjudgment is made, avoiding the generation of erroneous reference speeches in this case and improving the accuracy of the reference speeches.

[0096] Example 4:

[0097] Based on Example 2, the artificial intelligence-based dialogue device also includes: a cold spot processing module 218, which is used to determine whether the conversation partner is silent or has continuously output speeches without substantive content for more than a preset number of times. If it is determined to be so, a reference speech and / or speech auxiliary information is generated to guide the continuation of the conversation to avoid interruption of the conversation.

[0098] In this embodiment, if the system detects that the conversation partner has been silent for longer than a threshold, or has continuously spoken without substantive content for more than a preset number of times, for example, saying "hmm," "oh," or "haha" three times in a row, it means that the current conversation has fallen into a lull. At this time, a guiding reference speech should be generated, such as "By the way, a new milk tea shop has opened near the company recently, and it's quite delicious."

[0099] The motivation for designing this module is: if this module is missing, when a user speaks and the other party is silent, the system will not be able to detect the other party's speech and will not actively generate reference speeches, which may cause the conversation to be interrupted.

[0100] The beneficial effects of this module are: improving the fluency of conversations, avoiding embarrassment or misunderstandings caused by awkward silences, and helping users maintain the initiative, especially in scenarios such as negotiations, sales, or emotional exchanges.

[0101] Example 5:

[0102] Based on Example 2, the artificial intelligence-based dialogue device also includes: a non-dialogue information acquisition module 220, which is used to obtain non-dialogue information in a specific manner at a specific time and input it into the dialogue assistance module. The specific manner includes manual addition, online search, and recommendation by the dialogue device. The non-dialogue information is independent of the speaker's speech, and the specific timing is determined by the user or the dialogue device.

[0103] In this embodiment, a script library can be established in the system to categorize and store script templates for different scenarios. For example, scenarios such as sales, customer service, and negotiation can be categorized. The system supports dynamic updating of script templates to adapt to the needs of different users and changing scenarios. Users can add new script templates or modify existing templates based on actual conversations.

[0104] In this embodiment, a knowledge base can also be established within the system to categorize and store knowledge in different fields, such as medicine, law, and computers. The system supports dynamic updates of knowledge content to adapt to the needs of different users and changing scenarios. Users can add new knowledge or modify existing knowledge based on actual conversations.

[0105] In this embodiment, it is also possible to obtain time-sensitive information through online search, for example:

[0106] Real-time news: Get the latest news from authoritative news websites or API interfaces to ensure that the information cited in the reference speech is timely and accurate.

[0107] Policy updates: Obtain the latest policies and regulations from official government websites or industry regulatory agencies to help users cite authoritative information in conversations.

[0108] Market data: Obtain the latest market trends, price fluctuations, and other information from financial data platforms or industry reports to support user decision-making in business negotiations or consulting scenarios.

[0109] Social media hot topics: Use crawler technology or API interfaces to obtain current hot topics from social media platforms, helping users integrate hot topics into conversations and enhance interactivity.

[0110] The motivation for designing this module is that although large language models or large multimodal models are trained on large-scale data, the responses they generate are often general, the generated content is not timely, and lacks optimization for specific scenarios. Such general responses may not meet the specific needs of users in complex or professional scenarios. Different speech techniques and strategies are required in specific scenarios (such as sales, negotiation, customer service, etc.), and knowledge support in specific fields is required in complex scenarios (such as medical, legal, technical, etc.). When it comes to time-sensitive topics, it is necessary to search the Internet for real-time information. For example, in sales scenarios, guiding speech techniques are needed; in customer service scenarios, problem-solving speech techniques are needed.

[0111] The beneficial effects of this module are: by obtaining non-dialogue information, such as verified speech templates, professional knowledge in various fields, and online search information, it can better meet the communication needs of users in specific scenarios, ensure that the system can provide accurate and professional reference speeches based on the scenarios, and improve the quality of the reference speech content.

[0112] Example 6:

[0113] Based on Example 2, the artificial intelligence-based dialogue device also includes: an interruption processing module 222. During the period when the dialogue assistance module outputs content, if the user's speech is interrupted by the conversation partner, and the duration of the conversation partner's continuous speech from the moment of interruption exceeds a certain threshold, the dialogue assistance module is controlled to pause the output of content, and a new reference speech is triggered by the trigger generation detection module.

[0114] In this embodiment, while the conversation assistance module is outputting content, the system uses speaker recognition technology to determine the speaker's identity. If the system detects that the speaker is not only the user but also a conversation partner, and if the user does not speak while the conversation partner is speaking, it indicates that the user's speech has been interrupted. The duration of the conversation partner's speech is then measured from the moment of the interruption. If the duration exceeds a threshold, an interrupt signal is sent to the conversation assistance module, immediately halting the current reference speech generation process. Simultaneously, a new reference speech is triggered by the trigger generation detection module.

[0115] The motivation for designing this module is that in real-world conversations, users may be interrupted by their counterparts, especially in situations of heated emotion or heated discussion. If the conversational device cannot handle such interruptions, it may result in incomplete conversation context information, affecting the coherence and accuracy of the conversation.

[0116] The beneficial effects of this module are: the interruption processing module can detect interruptions in real time and record the speeches before and after the interruption in the context of the conversation record. The module can ensure that the reference speech is based on the complete conversation history, ensuring the coherence and integrity of the conversation.

[0117] Example 7:

[0118] On the basis of Example 2, the artificial intelligence-based dialogue device also includes: an irrelevant dialogue filtering module 224, and an irrelevant dialogue filtering module for determining whether the valid speech is a dialogue between an irrelevant person and a user or a conversation partner. If so, the valid speech is filtered out before the input information enters the dialogue assistance module. The irrelevant person refers to other persons who are not users and conversation partners.

[0119] In this embodiment, the system uses semantic analysis, keyword detection, and contextual coherence analysis to determine whether the user or conversation partner is conversing with an unrelated person. For example, if the user or conversation partner detects "Hello" or "I'm taking a call" in their conversation, the system can determine that the user is currently on the phone and is speaking to an unrelated person. The conversation partner's speech will then be intercepted by the conversation assistance module. The system allows users to selectively set the current unrelated person as a new conversation partner as needed. If so, no interception measures will be taken.

[0120] The motivation for designing this module is that during a conversation, the user or the conversation partner may talk to an unrelated person at any time. The conversation at this time is irrelevant to the current conversation scenario and will be input into the conversation assistance module as input information. Therefore, it is necessary to identify and filter between the conversation assistance modules to ensure that the input information entering the conversation assistance module only contains the conversation content between the user and the conversation partner.

[0121] The beneficial effects of this module are: through precise identification and filtering, it reduces the system's misjudgment of irrelevant conversations, improves the accuracy of reference speeches, and avoids the generation of irrelevant reference speeches due to irrelevant content.

[0122] Example 8:

[0123] Based on any of Examples 2 to 7, the AI-based dialogue device further includes a dialogue parameter setting module 226 for setting dialogue parameters, which are configuration items used to guide and optimize the dialogue assistance content. Common dialogue parameters include: conversation partner, conversation scenario, conversation intent, conversation topic, language style, conversation partner's personality, conversation taboos, and the relationship between the user and the conversation partner.

[0124] In this embodiment, before the conversation, the user can input or select conversation parameters through the system interface.

[0125] Regarding the setting of conversation scenes, the system provides multiple preset scene options, such as business meetings, chatting with friends, customer service, etc. Users can select or create new conversation scenes.

[0126] Regarding the settings of the conversation partner, the system provides a drop-down menu or search box, allowing the user to select or enter relevant information about the conversation partner.

[0127] Regarding the setting of conversation intention, the system provides a text box where users can describe in detail the purpose of the conversation and the expected results.

[0128] Regarding the setting of the conversation topic, the system provides a text box where users can enter the topic of the conversation to guide the direction of the conversation.

[0129] Regarding the language style setting, the system provides radio buttons or drop-down menus for users to select different language styles such as formal, casual, humorous, etc.

[0130] For the setting of the conversation partner's personality, the system provides multiple personality trait options, such as extroversion, introversion, impatience, etc. Users can choose based on their understanding of the conversation partner.

[0131] For setting conversation taboos, the system provides a text box where users can enter topics or content that need to be avoided.

[0132] Regarding the relationship settings between users and conversation partners, the system provides a drop-down menu for users to select different relationship types such as friends, colleagues, superiors, etc.

[0133] When a user enters parameters, the system automatically verifies their completeness and rationality (e.g., whether required fields are filled in, whether the language style matches the scenario, etc.). If the parameters are incomplete or unreasonable, the module prompts the user to correct them. The system supports real-time parameter adjustment during a conversation (e.g., changing the language style or conversation intent). The updated parameters take effect immediately, affecting the generation of subsequent reference speeches. The system supports saving commonly used parameter combinations as templates (e.g., "Business Negotiation Template" or "Friend Chat Template"). Users can quickly load templates to reduce repeated settings. When a user sets a conversation partner, if they have previously had a conversation through the system, the set parameters will be automatically matched to avoid repeated settings.

[0134] This module was designed because responses generated by large language models or multimodal models are generally general and lack specificity. By setting conversation parameters, we can provide specific guidance to the conversation assistance module, making the generated reference utterances more tailored to the current conversation scenario and user needs.

[0135] The beneficial effect of this module is that different users have different requirements for conversations in different scenarios, such as language style and conversational intent. The conversation parameter setting module can meet these personalized needs, better meet the current scenario and user needs, and improve the accuracy and relevance of the conversation.

[0136] Example 9:

[0137] The present invention provides a device, such as Figure 3 As shown, it includes at least one processor 301; and a memory 302 that is communicatively connected to the at least one processor 301; wherein the memory 302 stores instructions that can be executed by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute the above-mentioned artificial intelligence-based dialogue method.

[0138] The memory 302 and processor 301 are connected using a bus. The bus can include any number of interconnected buses and bridges, connecting various circuits of one or more processors 301 and memory 302. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and are therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor 301 is transmitted over a wireless medium via an antenna. Furthermore, the antenna receives data and transmits it to the processor 301.

[0139] The processor 301 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 302 can be used to store data used by the processor 301 when performing operations.

[0140] Example 10:

[0141] The present invention provides a storage medium storing a computer program, which implements the above method embodiment when running on a computer.

[0142] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., various media that can store program code.

[0143] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above description is only the specific implementation method of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of the present invention.

Claims

1. A conversation method based on artificial intelligence, characterized in that: include: receiving a voice signal; Distinguish some or all speakers to whom the received speech signal belongs; Determine whether the speaker in the voice signal is the user or the conversation partner, and identify the speech and / or speech-derived information of the speaker as the user or the conversation partner based on the determination result, thereby obtaining a valid speech. The user speech, the conversation partner's speech, the user's speech-derived information, and the conversation partner's speech-derived information are collectively referred to as a valid speech, and the speech-derived information is derived from and related to the speech. Associating the valid speech with the speaker through identification information; Determine whether a specific condition is met. If so, execute the dialogue assistance step; otherwise, do not execute the dialogue assistance step. Dialogue assistance step: generating dialogue assistance content based on input information, the dialogue assistance content including at least one of a reference speech, an alternative speech, and speech assistance information, the alternative speech being played as voice, the input information including at least one of the valid speech, the identification information, previously generated dialogue assistance content, and non-dialogue information, the non-dialogue information being independent of the speaker's speech.

2. A conversation device based on artificial intelligence, characterized in that: include: Voice acquisition module: used to receive voice signals; Speaker differentiation module: used to distinguish some or all speakers of the received speech signal; Speech screening module: determines whether the speaker in the voice signal is the user or the conversation partner, and identifies the speech and / or speech-derived information of the user or the conversation partner based on the judgment result, thereby obtaining valid speeches. The user speech, the conversation partner's speech, the user's speech-derived information, and the conversation partner's speech-derived information are collectively referred to as valid speeches. The speech-derived information is derived from and related to the speech. Identification association module: associates the valid speech with the speaker through identification information; Trigger generation detection module: determines whether specific conditions are met. If so, the dialogue assistance module is triggered to generate corresponding content. Otherwise, the dialogue assistance module is not triggered at this time. Dialogue assistance module: used to generate dialogue assistance content based on input information, the dialogue assistance content includes at least one of a reference speech, an alternative speech, and speech assistance information, the alternative speech is played as voice, the input information includes at least one of the valid speech, the identification information, previously generated dialogue assistance content, and non-dialogue information, the non-dialogue information is independent of the speaker's speech.

3. The artificial intelligence-based dialogue device according to claim 2, characterized in that: Also includes: The misjudgment correction module is used to detect whether the user or the conversation partner has spoken when the dialogue assistance module generates a reference speech. If, when generating the reference speech, it is detected that the user has not spoken and the conversation partner has spoken continuously for a period of time exceeding a certain threshold since the trigger generation detection module determines that the condition is met, the dialogue assistance module is controlled to pause outputting content and a new reference speech is triggered by the trigger generation detection module.

4. The artificial intelligence-based dialogue device according to claim 2, characterized in that: Also includes: The cold silence processing module is used to determine whether the conversation partner is silent or has continuously output non-substantive speeches for more than a preset number of times. If so, it generates reference speeches and / or speech auxiliary information to guide the continuation of the conversation to avoid conversation interruption.

5. The artificial intelligence-based dialogue device according to claim 2, characterized in that: Also includes: A non-dialogue information acquisition module is used to acquire non-dialogue information at a specific time and in a specific manner and input it into the dialogue assistance module. The specific manner includes at least one of manual addition, online search, and recommendation by the dialogue device. The non-dialogue information is independent of the speaker's speech, and the specific timing is determined by the user or the dialogue device.

6. The artificial intelligence-based dialogue device according to claim 2, characterized in that: It also includes an interruption processing module. During the period when the dialogue assistance module is outputting content, if the user's speech is interrupted by the conversation partner, and the duration of the conversation partner's continuous speech from the moment of interruption exceeds a certain threshold, the dialogue assistance module is controlled to pause outputting content, and a new reference speech is triggered by the trigger generation detection module.

7. The artificial intelligence-based dialogue device according to claim 2, characterized in that: It also includes an irrelevant conversation filtering module for determining whether the valid speech is a conversation between an irrelevant person and the user or conversation partner. If so, the valid speech is filtered out before the input information enters the conversation assistance module. The irrelevant person refers to other persons who are not users or conversation partners.

8. The artificial intelligence-based dialogue device according to any one of claims 2 to 7, characterized in that: It also includes a dialogue parameter setting module for setting dialogue parameters, where the dialogue parameters are configuration items for guiding and optimizing the dialogue assistance content.

9. A device, characterized in that The device includes a memory and one or more processors; wherein the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the device executes the artificial intelligence-based dialogue method as described in claim 1.

10. A storage medium storing a computer program, characterized in that: When the computer program is run on a computer, the computer is enabled to execute the artificial intelligence-based dialogue method according to claim 1 .

Citation Information

Cited By

  • Multi-role configuration and effective judgment method based on semantic arbitration

    CN120809299A