Continuous dialogue method and device, electronic device, readable storage medium

By detecting keywords and extracting voiceprint features, and combining user voice data to calculate the probability of continuous dialogue triggering, intelligent mode switching is achieved. This solves the problems of false triggering and low interaction fluency in the continuous dialogue function of existing smart devices, and improves user experience and interaction security.

CN121808023BActive Publication Date: 2026-05-29BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

The continuous dialogue function of existing smart devices suffers from problems such as users having to repeatedly trigger the wake word, susceptibility to environmental noise or interference from other people's voices leading to false triggers or erratic responses, low interaction fluency, and difficulty in meeting users' needs for natural, efficient, and secure intelligent interaction.

Method used

After detecting the first keyword, the system acquires the user's voice data, extracts voiceprint features, and matches them with default voiceprint features to enter the normal dialogue mode. In the normal dialogue mode, the system calculates the probability of triggering continuous dialogue based on the user's voice data, adjusts the probability to enter the continuous dialogue mode, realizes intelligent mode switching, and designs differences in the method of generating reply text and the duration of audio reception in different modes.

Benefits of technology

It enables users to have natural, smooth and continuous dialogue with AI, avoids accidental triggering by environmental noise or other people's voices, improves the personalization and security of the interaction, adapts to the needs of different interaction scenarios, improves the smoothness and naturalness of the interaction, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808023B_ABST
    Figure CN121808023B_ABST
Patent Text Reader

Abstract

The application provides a continuous conversation method and device, an electronic device and a readable storage medium, and belongs to the technical field of continuous conversation. The method comprises the following steps: in response to detecting a first keyword, acquiring first user voice data; extracting a first voiceprint feature from voice data corresponding to the first keyword, calculating a first voiceprint matching degree of the first voiceprint feature and a default voiceprint feature, and entering a normal conversation mode if the first voiceprint matching degree is not less than a first matching degree threshold; after entering the normal conversation mode, calculating a continuous conversation trigger probability based on the first user voice data; entering a continuous conversation mode if the continuous conversation trigger probability is not less than a first probability threshold; after entering the continuous conversation mode, determining a first conversation topic and a first user intention based on the first user voice data, and generating and playing a first reply text based on the first user intention and the first conversation topic. The application can realize natural and smooth continuous conversation between a user and an AI, and meet the user demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of continuous dialogue technology, and more specifically, relates to continuous dialogue methods and apparatus, electronic devices, and readable storage media. Background Technology

[0002] With the development of artificial intelligence technology, intelligent interactive devices have been widely used in daily life, and continuous dialogue functionality is a core technology for enhancing the interactive experience. However, existing intelligent devices have significant limitations in dialogue interaction: on the one hand, users need to repeatedly trigger wake words to initiate each round of dialogue, making the interaction process cumbersome and difficult to meet the needs of multi-round continuous communication; on the other hand, while some devices support simple continuous dialogue, they are easily affected by environmental noise or interference from other people's voices, leading to false triggers or erratic responses, resulting in low interaction fluency. These problems mean that the practicality and user experience of existing continuous dialogue functions need improvement, making it difficult to meet users' needs for natural, efficient, and secure intelligent interaction. Summary of the Invention

[0003] The purpose of this application is to provide a continuous dialogue method and apparatus, electronic device, and readable storage medium to enable users to have natural and fluent continuous dialogue with AI and meet user needs.

[0004] A first aspect of this application provides a continuous dialogue method, including:

[0005] In response to the detection of the first keyword, the first user's voice data is acquired;

[0006] Extract the first voiceprint feature from the voice data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode.

[0007] After entering the normal dialogue mode, the topic type, first probability adjustment coefficient, and second probability adjustment coefficient are determined based on the first user's voice data; the continuous dialogue probability level corresponding to the topic type is determined from the mapping table between topic types and continuous dialogue probability levels; the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level are determined; the average continuous dialogue probability is adjusted based on the first probability adjustment coefficient, the second probability adjustment coefficient, and the probability range to obtain the continuous dialogue trigger probability; if the continuous dialogue trigger probability is not less than the first probability threshold, the continuous dialogue mode is entered; the response text generation method and the recording duration of a single dialogue are different between the continuous dialogue mode and the normal dialogue mode.

[0008] After entering continuous dialogue mode, the first dialogue topic and the first user intent are determined based on the first user's voice data, and the first reply text is generated and played based on the first user intent and the first dialogue topic.

[0009] A second aspect of this application provides a continuous dialogue device, comprising:

[0010] The voice acquisition module is used to acquire the first user's voice data in response to the detection of the first keyword;

[0011] The voiceprint recognition module is used to extract the first voiceprint feature from the voice data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode.

[0012] The continuous dialogue judgment module is used to determine the topic type, first probability adjustment coefficient, and second probability adjustment coefficient based on the first user's voice data after entering the normal dialogue mode; determine the continuous dialogue probability level corresponding to the topic type from the mapping table between topic types and continuous dialogue probability levels; determine the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level; adjust the average continuous dialogue probability based on the first probability adjustment coefficient, the second probability adjustment coefficient, and the probability range to obtain the continuous dialogue trigger probability; if the continuous dialogue trigger probability is not less than the first probability threshold, then enter the continuous dialogue mode; the response text generation method and the audio recording duration of a single dialogue are different between the continuous dialogue mode and the normal dialogue mode.

[0013] The first continuous dialogue module is used to determine the first dialogue topic and the first user intent based on the first user's voice data after entering the continuous dialogue mode, and to generate and play the first reply text based on the first user intent and the first dialogue topic.

[0014] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described continuous dialogue method.

[0015] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described continuous dialogue method.

[0016] The beneficial effects of the continuous dialogue method and apparatus, electronic device, and readable storage medium provided in the embodiments of this application are as follows:

[0017] In this embodiment, after detecting the first keyword, the system enters a normal dialogue mode through voiceprint feature matching. Only users whose voiceprints match the criteria are allowed to initiate the interaction, thus preventing accidental triggering due to environmental noise or other people's voices and ensuring the personalization and security of the interaction. In normal dialogue mode, the probability of continuous dialogue triggering is calculated based on the first user's voice data, realizing intelligent mode switching without manual operation by the user. This avoids the tediousness of repeated wake-up and prevents invalid sound reception when continuous dialogue is not needed, making the interaction more in line with the user's actual needs.

[0018] This application clearly defines the differences in response text generation methods and single-session audio recording duration between the normal dialogue mode and the continuous dialogue mode. It provides adaptability services for different interaction scenarios. For example, the continuous dialogue mode extends the audio recording duration and generates more coherent responses, improving the fluency and naturalness of the interaction and meeting users' needs for efficient and intelligent interaction. Overall, the method of this application is simple, practical, and effectively balances the convenience and security of interaction, significantly optimizing the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a continuous dialogue method provided in an embodiment of this application;

[0021] Figure 2 A structural block diagram of a continuous dialogue device provided in an embodiment of this application;

[0022] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0024] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0025] It should be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0026] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a continuous dialogue method provided in an embodiment of this application. The method can be executed by an electronic device, and specifically, the method may include S101 to S104.

[0027] S101: In response to the detection of the first keyword, acquire the first user voice data.

[0028] In this embodiment, the first keyword is a preset word used to wake up the smart device; the smart device can be smart glasses, a smart voice assistant, or a smart speaker, etc. The first user voice data is the voice information subsequently input by the user after the first keyword is detected. This embodiment first wakes up the device using the first keyword, then ensures interaction security through voiceprint matching, and subsequently determines the need for continuous dialogue based on the user's voice. By designing differentiated modes, it balances the convenience and rationality of the interaction, thereby improving the user experience.

[0029] For example, the smart glasses continuously monitor ambient voice during operation. When the microphone array of the smart glasses detects the first keyword, it begins to acquire and save the first user's voice data. Then, it extracts the voiceprint features corresponding to the first keyword and compares them with locally stored default voiceprint features. If the matching degree is not less than a first threshold, the smart glasses enter a normal dialogue mode. In this mode, the smart glasses can parse the first user's voice data to calculate the probability of triggering continuous dialogue. If the probability reaches the threshold, it switches to continuous dialogue mode. After switching, the smart glasses can analyze the voice data to determine the dialogue topic and user intent. In continuous dialogue mode, the recording duration is extended according to adaptation rules, generating coherent reply text and playing it through the speaker, achieving multi-turn seamless interaction.

[0030] S102: Extract the first voiceprint feature from the speech data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode.

[0031] In this embodiment, the first voiceprint feature is a unique biometric feature extracted from voice data containing the first keyword, representing the identity of the voice speaker; the default voiceprint feature is a target user voiceprint baseline feature pre-stored in the smart device for identity verification; the first voiceprint matching degree is a quantitative index of similarity between the first voiceprint feature and the default voiceprint feature; the first matching degree threshold is a pre-set threshold used to determine whether the voiceprint matching is successful; the normal dialogue mode is the basic interaction mode that the device enters after the voiceprint matching is successful, supporting the interaction process corresponding to a single wake-up.

[0032] The core consideration of this embodiment is to ensure the security and personalization of the interaction. This embodiment realizes identity verification through the uniqueness of voiceprint biometrics, and only allows users whose voiceprints match the criteria to start the normal dialogue mode. This blocks the accidental triggering by non-target users or environmental noise from the source of the interaction, while laying a secure foundation for subsequent mode switching, thus balancing the convenience of interaction and the security of use.

[0033] For example, the intelligent in-vehicle assistant is in standby listening mode, with its built-in microphone continuously collecting voice signals from the in-vehicle environment. When the intelligent in-vehicle assistant detects a first keyword, it extracts a voice segment containing that keyword as the corresponding voice data. The intelligent in-vehicle assistant preprocesses this voice data using a voiceprint extraction module, removing background noise and extracting the first voiceprint feature. The intelligent in-vehicle assistant calls the locally stored default voiceprint features and calculates the first voiceprint matching degree between the two using a voiceprint comparison algorithm. The intelligent in-vehicle assistant retrieves a preset first matching degree threshold and compares the calculated first voiceprint matching degree with this threshold. If the first voiceprint matching degree is not less than the first matching degree threshold, the identity verification is deemed successful, and the in-vehicle assistant automatically enters a normal dialogue mode. At this time, the user can initiate specific requests, and the intelligent in-vehicle assistant will respond based on the interaction rules of this mode. The entire process is completed quickly locally, ensuring the real-time nature of identity verification and the continuity of interaction.

[0034] S103: After entering the normal dialogue mode, determine the topic type, the first probability adjustment coefficient, and the second probability adjustment coefficient based on the first user's voice data; determine the continuous dialogue probability level corresponding to the topic type from the mapping table between topic types and continuous dialogue probability levels; determine the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level; adjust the average continuous dialogue probability based on the first probability adjustment coefficient, the second probability adjustment coefficient, and the probability range to obtain the continuous dialogue trigger probability; if the continuous dialogue trigger probability is not less than the first probability threshold, then enter the continuous dialogue mode; the response text generation method and the recording duration of a single dialogue are different between the continuous dialogue mode and the normal dialogue mode.

[0035] In this embodiment, the probability of triggering continuous dialogue is a quantitative indicator of the likelihood that the user has a need for continuous dialogue, based on the first user's voice data; the first probability threshold is a pre-set critical value used to determine whether to start the continuous dialogue mode; the continuous dialogue mode is an interaction mode switched after the trigger probability reaches the threshold, supporting multiple rounds of continuous interaction; the reply text generation method is the specific logic and rules by which the smart device generates response text according to the user's needs; the recording duration of a single dialogue is the length of time that the smart device keeps the microphone in recording state during a single interaction.

[0036] The core consideration of this embodiment is to adapt to diverse user interaction needs and balance convenience with resource efficiency. The normal dialogue mode is the basic interaction state, avoiding resource waste when continuous dialogue is not required. This embodiment calculates the trigger probability using the first user's voice data, accurately capturing potential continuous user needs without requiring manual mode switching, thus improving interaction smoothness. Simultaneously, this embodiment clarifies the differences in response generation and audio recording duration between the two modes. This is because continuous dialogue requires a longer audio recording duration to accommodate pauses in thought, and responses need to maintain continuity, while normal dialogue emphasizes efficient response. This targeted design optimizes the user experience in different scenarios.

[0037] For example, after entering normal dialogue mode, the smart glasses can immediately retrieve the acquired first user voice data and analyze the voice content, semantic relationships, and other information through a preset judgment mechanism to calculate the probability of triggering continuous dialogue. The smart glasses compare the calculated trigger probability with a built-in first probability threshold. If the trigger probability is not less than the first probability threshold, the smart glasses automatically switch from normal dialogue mode to continuous dialogue mode. After the switch is completed, the continuous dialogue mode adjusts parameters according to preset rules: the recording duration is extended from 10 seconds in normal mode to 30 seconds to adapt to the pause intervals of multiple rounds of user input; when generating reply text, it combines the dialogue topic and intent in the first user voice data to generate coherent response content, avoiding fragmented replies. For example, if a user says "check the weather" and then asks about clothing, normal mode only replies with the weather, while continuous dialogue mode adds clothing suggestions after replying with the weather, and maintains 30 seconds of recording to wait for the user's subsequent questions, achieving seamless multi-round interaction.

[0038] In this embodiment, determining the topic type, the first probability adjustment coefficient, and the second probability adjustment coefficient based on the first user's voice data specifically includes:

[0039] The first user's voice data is converted into a text string, and the text string is segmented and encoded to obtain a word vector sequence;

[0040] Extract global semantic features from word vector sequences, and determine topic types based on global semantic features;

[0041] The first probability adjustment coefficient and the second probability adjustment coefficient are determined based on the text string.

[0042] In this embodiment, determining the first probability adjustment coefficient and the second probability adjustment coefficient based on the text string specifically includes: splitting the text string into short sentences to obtain multiple short sentences, calculating the semantic correlation between every two adjacent short sentences, and determining the first probability adjustment coefficient based on the multiple semantic correlations; calculating the intent complexity of the text string, and determining the second probability adjustment coefficient based on the intent complexity.

[0043] In this embodiment, the text string is a character sequence obtained after the first user's voice data is processed by speech-to-text; word segmentation is the operation of splitting the text string into independent words; encoding is the process of converting words into a vector form that can be recognized by a computer; the word vector sequence is a continuous set of vectors formed by the encoded words after word segmentation; global semantic features are features extracted from the word vector sequence that reflect the overall meaning of the text; the continuous dialogue probability level is a preset hierarchical classification corresponding to the topic type and representing the probability of continuous dialogue; the mapping table is a preset data table that stores the correspondence between topic types and continuous dialogue probability levels; the average continuous dialogue probability is a preset baseline probability value corresponding to each continuous dialogue probability level; the probability range is the allowed probability fluctuation range for each continuous dialogue probability level; the first probability adjustment coefficient is a parameter determined based on the semantic correlation between adjacent short sentences and used to adjust the average continuous dialogue probability; the second probability adjustment coefficient is a parameter determined based on the intent complexity of the text string and used to adjust the average continuous dialogue probability; the semantic correlation is a quantitative indicator of the degree of semantic correlation between adjacent short sentences; the intent complexity is a quantitative feature of the number of independent intents contained in the text string and the degree of correlation.

[0044] The core consideration in this embodiment is to improve the accuracy of calculating the probability of continuous dialogue triggering through precise adjustments across multiple dimensions, thus better aligning with users' actual needs. Topic type reflects the basic probability of continuous dialogue, with a mapping table providing a standardized benchmark. Semantic relevance reflects the coherence of user expression; higher relevance indicates a greater likelihood of continuous dialogue, corresponding to the first probability adjustment coefficient. Intent complexity reflects the complexity of user needs; complex needs often require multiple rounds of dialogue, corresponding to the second probability adjustment coefficient. This embodiment combines the benchmark probability with adjustment coefficients from both dimensions, ensuring standardized calculations while also considering the personalized characteristics of the text, avoiding biases caused by single-dimensional judgments, and making the triggering probability more closely reflect actual interaction scenarios.

[0045] For example, after acquiring the first user's voice data, the intelligent voice assistant can convert it into a text string using speech-to-text technology. For instance, the text string corresponding to the user's voice data might be: "Check the weather in location A tomorrow, check the probability of precipitation the day after tomorrow, and recommend suitable travel methods." This text string is then segmented into individual words and encoded to generate a word vector sequence. Global semantic features are extracted from the word vector sequence, and combined with preset topic classification rules, the topic type of the text string is determined to be a tool-type weather and travel-related topic. The intelligent voice assistant can then call a pre-stored mapping table of topic types and continuous dialogue probability levels, and find that the continuous dialogue probability level for tool-type weather and travel-related topics is medium. Based on this probability level, a preset average continuous dialogue probability of 0.5 and a probability range of 0.3 to 0.7 are retrieved.

[0046] The intelligent voice assistant can break down a text string into short sentences, resulting in three phrases: "Check the weather in location A tomorrow, check the probability of precipitation the day after tomorrow, and recommend suitable travel methods." The assistant calculates the semantic relevance between any two adjacent phrases. The first and second phrases both revolve around weather queries, resulting in a semantic relevance of 0.8; the second and third phrases involve both weather and travel suggestions, resulting in a semantic relevance of 0.7. The average of 0.75 is taken as the final semantic relevance. Based on preset rules, a semantic relevance of 0.7 to 0.8 corresponds to a first probability adjustment coefficient of 0.15. Simultaneously, the assistant calculates the intent complexity of the text string. This text contains three independent intents: "check the weather," "check the probability of precipitation," and "get travel recommendations." According to preset standards, the higher the intent complexity level among these three intents, the higher the second probability adjustment coefficient is determined to be, resulting in a second probability adjustment coefficient of 0.2.

[0047] The intelligent voice assistant can adjust the average continuous dialogue probability of 0.5 based on a first probability adjustment coefficient of 0.15, a second probability adjustment coefficient of 0.2, and a probability range of 0.3 to 0.7. According to the preset adjustment rules, the average continuous dialogue probability is summed with the two adjustment coefficients, resulting in 0.5 + 0.15 + 0.2 = 0.85. This value falls within the probability range of 0.3 to 0.7; therefore, 0.85 is the final continuous dialogue trigger probability. If the adjusted value exceeds the probability range, the upper or lower limit of the probability range is taken as the continuous dialogue trigger probability to ensure the reasonableness and validity of the result.

[0048] This embodiment provides standardized baseline probabilities for different topics through a mapping table between topic types and continuous dialogue probability levels, ensuring the standardization and consistency of calculations. It determines a first probability adjustment coefficient based on the semantic correlation between adjacent short sentences and a second probability adjustment coefficient based on intent complexity, accurately capturing the coherence of the text and the complexity of the demand, thus achieving personalized probability calibration. By combining baseline probabilities with dual-dimensional adjustment coefficients, this embodiment balances standardization and personalization, improving the accuracy of continuous dialogue trigger probability calculations, effectively avoiding false triggers or missed triggers caused by single-dimensional judgments, making continuous dialogue mode switching more closely match the user's true intent, significantly optimizing interaction smoothness and user experience, and enhancing the interactive adaptability of smart devices.

[0049] S104: After entering the continuous dialogue mode, determine the first dialogue topic and the first user intent based on the first user's voice data, and generate and play the first reply text based on the first user intent and the first dialogue topic.

[0050] In this embodiment, the first dialogue topic is the core discussion direction extracted from the first user's voice data, and it is the core content that runs through the continuous dialogue; the first user intent is the specific need or goal that the user wants to achieve through the first user's voice data; the first reply text is the targeted text content generated by the smart device based on the first user intent and the first dialogue topic in response to the user's needs; generation is the process by which the smart device constructs the reply text according to preset logic and rules; playback is the operation by which the smart device converts the reply text into a voice signal and transmits it to the user through the audio output module.

[0051] The core consideration of this embodiment is to ensure the accuracy and consistency of responses in continuous dialogue mode, aligning with users' actual needs. The core value of continuous dialogue mode lies in multi-round coherent interaction, and accurately identifying the first dialogue topic ensures that subsequent interactions do not deviate from the core, while clearly defining the first user intent ensures that the response directly addresses the user's needs. This embodiment generates response text based on these two factors, avoiding a disconnect between the response and user needs, while maintaining content coherence, which conforms to the characteristics of continuous dialogue interaction scenarios. By accurately focusing on the topic and intent, this embodiment further enhances the user experience in multi-round interactions, solving the problems of fragmented and insufficiently targeted responses in existing technologies.

[0052] For example, after the intelligent in-vehicle assistant enters continuous dialogue mode, it can retrieve the stored first user voice data, parse the voice data through the semantic analysis module, extract core keywords and semantic association information, and determine the first dialogue topic. The intelligent in-vehicle assistant can then use the intent recognition module, combined with keywords, context, and a common needs database, to clarify the first user's intent. For instance, if the user's voice data is "checking tomorrow's commute route and reminding me to bring documents," parsing determines the first dialogue topic to be commuting-related matters, and the first user's intent to check commuting routes and set reminders for items to bring. Based on this, the intelligent in-vehicle assistant can generate a coherent first response text containing route information and reminder content, according to the rules of continuous dialogue mode. Finally, the first response text is converted into clear speech and played to the user, while the system remains on record, waiting for subsequent user input to ensure a smooth transition in the continuous dialogue.

[0053] As can be seen from the above, this embodiment, after detecting the first keyword, enters the normal dialogue mode through voiceprint feature matching. Only users whose voiceprints match the criteria are allowed to initiate the interaction, thus avoiding false triggers caused by environmental noise or other people's voices from the source, ensuring the personalization and security of the interaction. In the normal dialogue mode, the probability of continuous dialogue triggering is calculated based on the first user's voice data, realizing intelligent mode switching without manual operation by the user. This avoids the tediousness of repeated wake-up and prevents invalid sound reception when continuous dialogue is not needed, making the interaction more in line with the user's actual needs.

[0054] This embodiment clearly defines the differences between the normal dialogue mode and the continuous dialogue mode in terms of reply text generation methods and the duration of audio reception per dialogue. It can provide adaptive services for different interaction scenarios. For example, in the continuous dialogue mode, the audio reception duration can be extended, generating more coherent replies, improving the fluency and naturalness of the interaction, and meeting users' needs for efficient and intelligent interaction. Overall, the method and process of this embodiment are simple, highly practical, effectively balancing the convenience and security of interaction, and significantly optimizing the user experience.

[0055] In one embodiment of this application, after generating and playing the first reply text based on the first user intent and the first dialogue topic, the method further includes:

[0056] Recording is performed based on the first recording duration;

[0057] If the second user's voice data is received within the first reception duration, then the second voiceprint feature is extracted from the second user's voice data.

[0058] Calculate the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature. If the second voiceprint matching degree is not less than the first matching degree threshold, then determine the second dialogue topic and the second user intent based on the second user voice data.

[0059] The third dialogue topic is determined based on the second and first dialogue topics;

[0060] Generate and play a second response text based on the second user intent and the third dialogue topic;

[0061] If no voice data is received from the user within the first reception duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start a new conversation.

[0062] In this embodiment, determining the third dialogue topic based on the second dialogue topic and the first dialogue topic includes:

[0063] Calculate the topic matching degree between the second dialogue topic and the first dialogue topic;

[0064] If the topic matching degree is not less than the first matching degree threshold, the first dialogue topic is supplemented based on the second dialogue topic to obtain the third dialogue topic;

[0065] If the topic matching degree is less than the first matching degree threshold, then the second dialogue topic will be used as the third dialogue topic.

[0066] In this embodiment, after calculating the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature, the method further includes:

[0067] If the matching degree of the second voiceprint is less than the matching degree threshold of the first voiceprint, then the first template text is generated and played; the first template text is used to instruct the user to determine whether to add the second voiceprint feature as a temporary voiceprint feature.

[0068] Receive the user's judgment result. If the judgment result is yes, add the second voiceprint feature as a temporary voiceprint feature and enter the multi-person continuous dialogue mode. In the multi-person continuous dialogue mode, the temporary voiceprint feature is used as the default voiceprint feature.

[0069] In this embodiment, the first reception duration is a preset time length during which the smart device maintains reception after the first reply text is generated and played in continuous dialogue mode; the second user voice data is the subsequent voice information input by the user within the first reception duration; the second voiceprint feature is the speaker identity feature extracted from the second user voice data; the second voiceprint matching degree is a quantitative index of the similarity between the second voiceprint feature and the default voiceprint feature; the second dialogue topic is the core discussion direction extracted based on the second user voice data; the second user intent is the specific need that the user wants to achieve through the second user voice data; and the third dialogue topic is the core discussion direction determined by combining the second dialogue topic and the first dialogue topic, used for subsequent responses.

[0070] Topic matching degree is a quantitative indicator of the relevance between the second dialogue topic and the first dialogue topic; the first template text is a preset fixed text used to ask the user whether to add temporary voiceprint features; temporary voiceprint features are non-initial default voiceprint features added after user confirmation and used for multi-person interaction; multi-person continuous dialogue mode is an interaction mode that supports users corresponding to temporary voiceprint features to participate in continuous dialogue; the judgment result is the user's confirmation or negative response to whether to add temporary voiceprint features.

[0071] The core consideration of this embodiment is to expand the adaptable scenarios of continuous dialogue, balancing the continuity of single-person multi-turn interaction with the flexibility of multi-person interaction, while ensuring interaction security. This embodiment ensures the orderliness of single-person continuous dialogue through a first sound reception duration setting; the dialogue ends when no input is received to avoid resource waste. Secondary voiceprint verification prevents interference from non-target users, ensuring personalized interaction. Topic matching determination enables the natural continuation or reasonable switching of dialogue topics, maintaining the continuity of multi-turn interaction. A temporary voiceprint feature mechanism allows authorization for others to participate in the dialogue, meeting the needs of multi-person collaborative scenarios. This embodiment solves the limitations of single-person interaction while balancing flexibility and security through an identity verification mechanism.

[0072] For example, after the intelligent voice assistant enters continuous dialogue mode and plays the first reply text, it immediately starts a first recording duration timer, maintaining continuous microphone recording according to preset rules, assuming the first recording duration is set to 15 seconds. If, during the timer, a user's family member says they want to know about restaurants near the route, the assistant collects this voice information through the microphone as the second user's voice data, then extracts the second voiceprint features and calculates its matching degree with the default voiceprint features. If the matching degree is not less than the first matching degree threshold, the assistant parses the second user's voice data, determines that the second dialogue topic is querying restaurants near the route, and the second user's intention is to obtain restaurant recommendations.

[0073] This embodiment calculates the topic matching degree between the topic and the first dialogue topic, commuting route. Since the two are highly correlated and the degree is not less than the first matching degree threshold, the first dialogue topic is supplemented based on the second dialogue topic to obtain the third dialogue topic: commuting route and nearby restaurant recommendations. Then, combined with the second user intent, a second reply text containing information on high-quality restaurants near the route is generated and played. If the second voiceprint matching degree is less than the first matching degree threshold, the assistant generates and plays the first template text asking whether to add the voiceprint as a temporary voiceprint feature. If the user responds with "yes" via voice, the assistant adds the second voiceprint feature as a temporary voiceprint feature and automatically switches to a multi-person continuous dialogue mode. Subsequent voice input from this family member will be matched and verified using the temporary voiceprint feature as the default voiceprint feature. If no user voice data is detected after the first 15-second recording period, the assistant determines that the dialogue has ended, and the user must re-trigger the first keyword to start a new dialogue.

[0074] In this embodiment, the setting of the first audio reception duration and the rule of ending the dialogue if no audio is received avoid meaningless continuous resource consumption during continuous dialogue, ensuring the orderliness of the interaction; the second voiceprint matching verification mechanism further blocks interference from non-target users, enhancing the security of the interaction; the topic matching determination enables flexible continuation or switching of the dialogue topic, ensuring that multi-round interactions do not deviate from user needs and improving the continuity of response; the design of temporary voiceprint features and multi-person continuous dialogue mode solves the limitations of single-person continuous dialogue, adapts to multi-person interaction scenarios such as home and office, and improves the practicality of the function; the overall process takes into account the security, continuity and flexibility of the interaction, comprehensively optimizes the user experience, and meets the needs of continuous dialogue in different scenarios.

[0075] In one embodiment of this application, after calculating the probability of continuous dialogue triggering based on the first user's voice data, the method further includes:

[0076] If the probability of triggering a continuous dialogue is less than the first probability threshold, then:

[0077] The first user's intent is determined based on the first user's voice data, and a second response text is generated and played based on the first user's intent.

[0078] The sound is received based on the second reception duration; the second reception duration is shorter than the first reception duration.

[0079] If a third user voice data is received within the second reception duration, the third user's intent is determined based on the third user voice data, and a third reply text is generated and played based on the third user's intent.

[0080] If no third user voice data is received within the second recording duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start a new conversation.

[0081] In this embodiment, the second response text refers to the targeted response text generated based on the first user intent when the probability of continuous dialogue triggering does not meet the threshold. The second reception duration refers to the preset duration for which the device maintains reception when the probability of continuous dialogue triggering is less than the first probability threshold. The third response text refers to the response text generated based on the third user intent in the second user voice data. The third user voice data refers to the subsequent voice information input by the user within the second reception duration. The third user intent refers to the specific requirement that the user wants to achieve through the third user voice data.

[0082] The core consideration of this embodiment is to differentiate and adapt to user interaction needs, balancing resource utilization efficiency with the convenience of basic interaction. A low probability of triggering continuous dialogue indicates that the user is unlikely to have multiple rounds of continuous requests. In this case, generating a second response text accurately addresses the core need. Simultaneously, setting a second reception duration shorter than the first reception duration preserves the possibility of the user supplementing their needs in real time while avoiding resource waste caused by prolonged reception. If the second user voice data is received, a timely response is initiated; otherwise, the dialogue ends and requires re-activation. This ensures the integrity of each interaction while preventing unnecessary occupation of device resources, achieving flexible adaptation to different demand scenarios.

[0083] For example, after the smart glasses enter normal dialogue mode, they can calculate the probability of continuous dialogue triggering based on the first user's voice data. If this probability is less than a first probability threshold, the smart glasses immediately parse the first user's voice data to determine the first user's intent. For instance, if the user's voice data is to inquire about the day's weather, the first user's intent is determined to be to obtain the day's weather information. Based on this intent, a second reply text containing information such as temperature and precipitation probability is generated and played through the speaker. After playback, the smart glasses start a second reception duration timer. Assuming the second reception duration is set to 3 seconds and the first reception duration is 15 seconds, this meets the setting that the second reception duration is less than the first reception duration. If the user adds a question about wind conditions within 3 seconds, the smart glasses collect this voice as the third user's voice data, parse it to determine that the third user's intent is to inquire about the day's wind conditions, generate the corresponding third reply text, and play it. If no user voice data is received after the 3-second timer expires, the smart glasses determine that the dialogue has ended, and the user needs to re-trigger the first keyword to start a new dialogue.

[0084] Corresponding to the continuous dialogue method in the above embodiments, Figure 2 This is a structural block diagram of a continuous dialogue device provided according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. Reference Figure 2 The continuous dialogue device 20 includes: a voice acquisition module 21, a voiceprint recognition module 22, a continuous dialogue judgment module 23, and a first continuous dialogue module 24.

[0085] Among them, the voice acquisition module 21 is used to acquire the first user's voice data in response to the detection of the first keyword;

[0086] The voiceprint recognition module 22 is used to extract the first voiceprint feature from the voice data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode.

[0087] The continuous dialogue judgment module 23 is used to determine the topic type, first probability adjustment coefficient, and second probability adjustment coefficient based on the first user voice data after entering the normal dialogue mode; determine the continuous dialogue probability level corresponding to the topic type from the mapping table between topic type and continuous dialogue probability level; determine the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level; adjust the average continuous dialogue probability based on the first probability adjustment coefficient, second probability adjustment coefficient, and probability range to obtain the continuous dialogue trigger probability; if the continuous dialogue trigger probability is not less than the first probability threshold, then enter the continuous dialogue mode; the response text generation method and the audio recording duration of a single dialogue are different between the continuous dialogue mode and the normal dialogue mode.

[0088] The first continuous dialogue module 24 is used to determine the first dialogue topic and the first user intent based on the first user's voice data after entering the continuous dialogue mode, and to generate and play the first reply text based on the first user intent and the first dialogue topic.

[0089] In one embodiment of this application, after generating and playing the first reply text based on the first user intent and the first dialogue topic, the continuous dialogue device 20 further includes: a second continuous dialogue module, used for:

[0090] Recording is performed based on the first recording duration;

[0091] If the second user's voice data is received within the first reception duration, then the second voiceprint feature is extracted from the second user's voice data.

[0092] Calculate the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature. If the second voiceprint matching degree is not less than the first matching degree threshold, then determine the second dialogue topic and the second user intent based on the second user voice data.

[0093] The third dialogue topic is determined based on the second and first dialogue topics;

[0094] Generate and play a second response text based on the second user intent and the third dialogue topic;

[0095] If no voice data is received from the user within the first reception duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start a new conversation.

[0096] In one embodiment of this application, when determining a third dialogue topic based on a second dialogue topic and a first dialogue topic, the second continuous dialogue module is specifically used for:

[0097] Calculate the topic matching degree between the second dialogue topic and the first dialogue topic;

[0098] If the topic matching degree is not less than the first matching degree threshold, the first dialogue topic is supplemented based on the second dialogue topic to obtain the third dialogue topic;

[0099] If the topic matching degree is less than the first matching degree threshold, then the second dialogue topic will be used as the third dialogue topic.

[0100] In one embodiment of this application, after calculating the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature, the continuous dialogue device 20 further includes: a multi-person continuous dialogue module, used for:

[0101] If the matching degree of the second voiceprint is less than the matching degree threshold of the first voiceprint, then the first template text is generated and played; the first template text is used to instruct the user to determine whether to add the second voiceprint feature as a temporary voiceprint feature.

[0102] Receive the user's judgment result. If the judgment result is yes, add the second voiceprint feature as a temporary voiceprint feature and enter the multi-person continuous dialogue mode. In the multi-person continuous dialogue mode, the temporary voiceprint feature is used as the default voiceprint feature.

[0103] In one embodiment of this application, after calculating the continuous dialogue trigger probability based on the first user voice data, the continuous dialogue device 20 further includes: a normal dialogue module, used for:

[0104] If the probability of triggering a continuous dialogue is less than the first probability threshold, then:

[0105] The first user's intent is determined based on the first user's voice data, and a second response text is generated and played based on the first user's intent.

[0106] The sound is received based on the second reception duration; the second reception duration is shorter than the first reception duration.

[0107] If a third user voice data is received within the second reception duration, the third user's intent is determined based on the third user voice data, and a third reply text is generated and played based on the third user's intent.

[0108] If no third user voice data is received within the second recording duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start a new conversation.

[0109] In one embodiment of this application, the continuous dialogue determination module 23, when determining the topic type, the first probability adjustment coefficient, and the second probability adjustment coefficient based on the first user voice data, is specifically used for:

[0110] The first user's voice data is converted into a text string, and the text string is segmented and encoded to obtain a word vector sequence;

[0111] Extract global semantic features from word vector sequences, and determine topic types based on global semantic features;

[0112] The first probability adjustment coefficient and the second probability adjustment coefficient are determined based on the text string.

[0113] In one embodiment of this application, the continuous dialogue judgment module 23, when determining the first probability adjustment coefficient and the second probability adjustment coefficient based on the text string, is specifically used for:

[0114] The text string is split into short sentences to obtain multiple short sentences. The semantic correlation between each pair of adjacent short sentences is calculated, and the first probability adjustment coefficient is determined based on the multiple semantic correlations.

[0115] Calculate the intent complexity of the text string, and determine the second probability adjustment coefficient based on the intent complexity.

[0116] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the voice acquisition module 21, voiceprint recognition module 22, continuous dialogue judgment module 23, and first continuous dialogue module 24 are shown.

[0117] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0118] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0119] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store information about default voiceprint characteristics.

[0120] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the embodiments of the continuous dialogue method provided in the embodiments of this application, or they can execute the implementation methods of the electronic device 300 described in the embodiments of this application, which will not be repeated here.

[0121] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0122] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0123] Those skilled in the art will recognize that the modules / units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules, units, or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or modules / units, or it may be an electrical, mechanical, or other form of connection.

[0126] The modules / units described as separate components may or may not be physically separate. Similarly, the components shown as modules / units may or may not be physical modules / units; they may be located in one place or distributed across multiple network modules / units. Some or all of the modules / units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0127] Furthermore, the functional modules / units in the various embodiments of this application can be integrated into one processing module / unit, or each module / unit can exist physically separately, or two or more modules / units can be integrated into one module / unit. The integrated modules / units described above can be implemented in hardware or in the form of software functional modules / units.

[0128] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A continuous dialogue method, characterized in that, include: In response to the detection of the first keyword, the first user's voice data is acquired; Extract the first voiceprint feature from the voice data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode. After entering the normal dialogue mode, the topic type, the first probability adjustment coefficient and the second probability adjustment coefficient are determined based on the first user voice data; the continuous dialogue probability level corresponding to the topic type is determined from the mapping table between topic type and continuous dialogue probability level. Determine the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level; adjust the average continuous dialogue probability based on the first probability adjustment coefficient, the second probability adjustment coefficient, and the probability range to obtain the continuous dialogue trigger probability; If the probability of triggering a continuous dialogue is not less than the first probability threshold, then the continuous dialogue mode is entered; the continuous dialogue mode differs from the normal dialogue mode in the method of generating reply text and the duration of audio reception for a single dialogue. After entering continuous dialogue mode, the first dialogue topic and the first user intent are determined based on the first user voice data, and the first reply text is generated and played based on the first user intent and the first dialogue topic.

2. The continuous dialogue method as described in claim 1, characterized in that, After generating and playing the first reply text based on the first user intent and the first dialogue topic, the method further includes: Recording is performed based on the first recording duration; If second user voice data is received within the first reception duration, then the second voiceprint feature is extracted from the second user voice data. Calculate the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature. If the second voiceprint matching degree is not less than the first matching degree threshold, then determine the second dialogue topic and the second user intent based on the second user voice data. A third dialogue topic is determined based on the second dialogue topic and the first dialogue topic; A second response text is generated and played based on the second user intent and the third dialogue topic; If no voice data from the user is received within the first reception duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start a new conversation.

3. The continuous dialogue method as described in claim 2, characterized in that, Determining the third dialogue topic based on the second dialogue topic and the first dialogue topic includes: Calculate the topic matching degree between the second dialogue topic and the first dialogue topic; If the topic matching degree is not less than the first matching degree threshold, then the first dialogue topic is supplemented based on the second dialogue topic to obtain a third dialogue topic; If the topic matching degree is less than the first matching degree threshold, then the second dialogue topic is used as the third dialogue topic.

4. The continuous dialogue method as described in claim 2, characterized in that, After calculating the second voiceprint matching degree between the second voiceprint feature and the default voiceprint feature, the method further includes: If the matching degree of the second voiceprint is less than the matching degree threshold of the first, then the first template text is generated and played; the first template text is used to instruct the user to determine whether to add the second voiceprint feature as a temporary voiceprint feature. If the user's judgment result is received, and the judgment result is yes, the second voiceprint feature is added as a temporary voiceprint feature, and the multi-person continuous dialogue mode is entered; in the multi-person continuous dialogue mode, the temporary voiceprint feature is used as the default voiceprint feature.

5. The continuous dialogue method as described in claim 2, characterized in that, After calculating the continuous dialogue trigger probability based on the first user voice data, the method further includes: If the probability of triggering continuous dialogue is less than the first probability threshold, then: Based on the first user's voice data, determine the first user's intent, and based on the first user's intent, generate and play the second reply text; The sound is received based on a second reception duration; the second reception duration is less than the first reception duration; If a third user voice data is received within the second reception duration, the third user intent is determined based on the third user voice data, and a third reply text is generated and played based on the third user intent. If no third user voice data is received within the second reception duration, the conversation is considered to have ended. After the conversation ends, the first keyword needs to be detected again to start the conversation again.

6. The continuous dialogue method as described in claim 1, characterized in that, The step of determining the topic type, the first probability adjustment coefficient, and the second probability adjustment coefficient based on the first user's voice data includes: The first user's voice data is converted into a text string, and the text string is segmented and encoded to obtain a word vector sequence; Global semantic features are extracted from the word vector sequence, and the topic type is determined based on the global semantic features; The first probability adjustment coefficient and the second probability adjustment coefficient are determined based on the text string.

7. The continuous dialogue method as described in claim 6, characterized in that, Determining the first probability adjustment coefficient and the second probability adjustment coefficient based on the text string includes: The text string is split into short sentences to obtain multiple short sentences. The semantic correlation between any two adjacent short sentences is calculated. A first probability adjustment coefficient is determined based on the multiple semantic correlations. Calculate the intent complexity of the text string, and determine a second probability adjustment coefficient based on the intent complexity.

8. A continuous dialogue device, characterized in that, include: The voice acquisition module is used to acquire the first user's voice data in response to the detection of the first keyword; The voiceprint recognition module is used to extract the first voiceprint feature from the voice data corresponding to the first keyword, calculate the first voiceprint matching degree between the first voiceprint feature and the default voiceprint feature, and if the first voiceprint matching degree is not less than the first matching degree threshold, then enter the normal dialogue mode. The continuous dialogue judgment module is used to determine the topic type, the first probability adjustment coefficient and the second probability adjustment coefficient based on the first user voice data after entering the normal dialogue mode; and to determine the continuous dialogue probability level corresponding to the topic type from the mapping table between topic type and continuous dialogue probability level. Determine the average continuous dialogue probability and probability range corresponding to the continuous dialogue probability level; adjust the average continuous dialogue probability based on the first probability adjustment coefficient, the second probability adjustment coefficient, and the probability range to obtain the continuous dialogue trigger probability; If the probability of triggering a continuous dialogue is not less than the first probability threshold, then the continuous dialogue mode is entered; the continuous dialogue mode differs from the normal dialogue mode in the method of generating reply text and the duration of audio reception for a single dialogue. The first continuous dialogue module is used to determine the first dialogue topic and the first user intent based on the first user voice data after entering the continuous dialogue mode, and to generate and play the first reply text based on the first user intent and the first dialogue topic.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice data processing method and device

    CN106782564A

  • Voice control method and device, electronic equipment and readable storage medium

    CN112581969A