Multi-mode collaborative off-line and on-line voice interaction method, system, equipment and medium

Through the multi-mode collaboration off-line voice interaction method, the problems of low accuracy, single reply content and unavailability of voice interaction systems in the prior art in a single mode are solved, and voice interaction is always available and efficient and accurate under different network conditions.

CN120071937AActive Publication Date: 2025-05-30SHENZHEN WAYTRONIC ELECTRONICS CO LTD

Patent Information

Application Number
CN202510539496.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Existing voice interaction systems have problems in single mode with low accuracy, single reply content, and unavailability of the network when it is not available.

Method used

A multi-mode collaboration off-line voice interaction method is proposed, which supports the coordinated work of offline mode, online mode and hybrid mode. By obtaining target voice data, feature extraction and replying to voice determination are used, and the accuracy and diversity of interaction is improved by using cloud resources.

Benefits of technology

When the network is in poor condition or no network connection, enable offline mode to ensure the availability of voice interaction; the online mode uses cloud resources to improve interaction accuracy and diversity; the hybrid mode combines offline and online advantages to optimize interaction effects and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071937A_ABST
    Figure CN120071937A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, and discloses a multi-mode collaborative off-line and on-line voice interaction method, system and device and a medium, and the method comprises the steps: carrying out the feature extraction according to target voice data in a target mode, obtaining a target feature, determining a reply voice according to the target feature, and carrying out the recognition of the reply voice; the target mode is one of an offline mode, an online mode and a mixed mode; if the target mode is a mixed mode, determining first data according to the target feature; and obtaining second data from the cloud according to the target feature, fusing the first data and the second data, determining a target text, and performing voice synthesis according to the target text to obtain a reply voice. Under the condition of ensuring that the voice is always available, the accuracy of voice interaction and the diversity of reply are improved by utilizing online resources as much as possible, and the mixed mode fully utilizes the advantages of cloud resources and local processing, so that the accuracy and richness of data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of speech recognition technology and artificial intelligence technology, and in particular to a multi-mode collaborative offline and online speech interaction method, system, device and medium. Background Art

[0002] Traditional speech interaction systems mostly use a single fixed mode (offline or online) for response processing. In the pure offline mode, limited by local computing power and knowledge base scale, there are the following disadvantages: the speech recognition accuracy is relatively low and the reply content is single, making it difficult to meet the requirements of complex scenarios; while the pure online mode highly depends on network stability, and there are the following disadvantages: when the network is interrupted or the latency is high, it cannot be guaranteed to be always available. In the prior art, it is impossible to overcome the disadvantages of the offline mode and the online mode existing alone when it cannot be guaranteed to be always available. Summary of the Invention

[0003] Based on this, in view of the technical problem that the single-mode implementation of speech interaction in the prior art cannot overcome the disadvantages of the offline mode and the online mode existing alone when it cannot be guaranteed to be always available, a multi-mode collaborative offline and online speech interaction method, system, device and medium are proposed.

[0004] In a first aspect, a multi-mode collaborative offline and online speech interaction method is provided. The method is applied to a target device, the target device is communicatively connected to a cloud, and the cloud is a remote server. The method includes: Obtain target speech data; In a target mode, perform feature extraction according to the target speech data to obtain target features, and determine a reply speech according to the target features, where the target mode is one of an offline mode, an online mode, and a hybrid mode.

[0005] In a second aspect, a multi-mode collaborative offline and online speech interaction system is provided. The system includes: a cloud and a target device. The target device is communicatively connected to the cloud, and the cloud is a remote server. The target device is configured to implement the steps of the above multi-mode collaborative offline and online speech interaction method.

[0006] In a third aspect, a target device is provided. The target device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above multi-mode collaborative offline and online speech interaction method are implemented.

[0007] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above multi-mode collaborative offline and online speech interaction method are implemented.

[0008] The multi - mode collaborative offline - online voice interaction method, system, device and medium of this application support the collaborative work of multiple modes, including offline mode, online mode and hybrid mode, ensuring that voice interaction is always available. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be carried out when the network is unavailable in the existing single online mode. The online mode can make full use of the powerful computing resources and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. The hybrid mode combines the advantages of offline and online, further optimizing the interaction effect. In the process of obtaining target voice data and extracting features to determine the response voice, while ensuring always available, online resources are utilized as much as possible to improve the accuracy of voice interaction and the diversity of responses, thereby enhancing the user experience. In the hybrid mode, by extracting Mel - Frequency Cepstral Coefficient (MFCC) features, voiceprint features and context semantic features from the target voice data, this multi - feature extraction method helps to accurately grasp various aspects of the voice characteristics. Furthermore, operations such as entity recognition and intent recognition based on the target features can accurately understand the user's needs, determine entity data and the first data. Obtain the second data from the cloud (based on the target features or the combination of target features and entity data), then fuse the first data and the second data to determine the target text, and finally synthesize the response voice. This process can make full use of the advantages of cloud resources and local processing, ensuring both the accuracy and richness of data, and adapting to the requirements of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thus providing high - quality voice responses for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0010] Among them: Figure 1 It is an application environment diagram of the multi - mode collaborative offline - online voice interaction method in an embodiment; Figure 2 It is a flowchart of the multi - mode collaborative offline - online voice interaction method in an embodiment; Figure 3 It is a flowchart of the multi - mode collaborative offline - online voice interaction method in an embodiment; Figure 4 It is a structural block diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0012] The multi-mode collaborative off-line and on-line voice interaction method provided by the embodiments of the present invention can be applied to an application environment such as Figure 1 . Among them, the target device 110 communicates with the cloud 120 through a network. The target device is communicatively connected to the cloud 120, and the cloud 120 is a remote server.

[0013] The target device 110 is configured to implement the multi-mode collaborative off-line and on-line voice interaction method of the present application, specifically including: obtaining target voice data; in the target mode, performing feature extraction according to the target voice data to obtain target features, and determining a reply voice according to the target features, where the target mode is one of an off-line mode, an on-line mode, and a hybrid mode.

[0014] By supporting the collaborative work of multiple modes such as the off-line mode, the on-line mode, and the hybrid mode, the present application ensures that the voice interaction is always available. When the network condition is poor or there is no network connection, the off-line mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be performed when the network is unavailable in the existing single on-line mode; while the on-line mode can make full use of the powerful computing resources and data resources of the cloud, improving the accuracy of voice interaction and the diversity of replies; the hybrid mode combines the advantages of off-line and on-line, further optimizing the interaction effect. In the process of obtaining target voice data and performing feature extraction to determine a reply voice, while ensuring always available, try to use on-line resources to improve the accuracy of voice interaction and the diversity of replies, thereby enhancing the user experience.

[0015] It can be understood that in the off-line mode, the target device 110 independently completes the voice interaction; in the on-line mode, the cloud 120 is the main and the target device 110 is the auxiliary to complete the voice interaction; in the hybrid mode, the cloud 120 and the target device 110 cooperate with each other, combining the advantages of off-line and on-line.

[0016] Among them, the target device 110 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. It can be understood that the target device 110 can be a computer device such as Figure 4 described.

[0017] The cloud 120 can be implemented by an independent server or a server cluster composed of multiple servers.

[0018] The present invention will be described in detail below through specific embodiments.

[0019] Please refer to Figure 2 as shown Figure 2 which is a schematic flowchart of a multi-mode collaborative offline and online voice interaction method provided by an embodiment of the present invention. The method is applied to a target device, and the target device is communicatively connected to the cloud. The cloud is a remote server. The method includes: S1: Obtain target voice data; Specifically, the user inputs voice to the target device through a microphone, and directly uses the voice as the target voice data, or processes the voice (for example, noise reduction), and uses the processed data as the target voice data.

[0020] S2: In the target mode, perform feature extraction according to the target voice data to obtain target features, and determine a reply voice according to the target features, where the target mode is one of an offline mode, an online mode, and a hybrid mode.

[0021] Specifically, obtain the pre-determined target mode; when the target mode is the offline mode, only the target device is used to implement the steps of performing feature extraction according to the target voice data to obtain target features, and determining a reply voice according to the target features; when the target mode is the online mode, the cloud 120 is the main and the target device 110 is the auxiliary to complete the voice interaction; in the hybrid mode, the cloud 120 and the target device 110 cooperate with each other, combining the advantages of offline and online.

[0022] The target mode is adaptively determined according to monitoring data such as network connection, local resource load, user preference, energy consumption constraint, and the current voice interaction type. That is to say, the entire process dynamically adjusts and determines the target mode according to multiple factors, thereby realizing multi-mode collaboration.

[0023] When performing feature extraction according to the target voice data, at least context semantic features are extracted, and all the extracted features are used as target features. Among them, a model trained based on ASR (Automatic Speech Recognition technology) is used to perform speech-to-text processing on the target voice data, generate a reply text according to the converted text, and a model trained based on TTS (Text To Speech technology) is used to perform text-to-speech on the reply text, and use the voice as the reply voice.

[0024] After the step of, in the target mode, extracting features from the target voice data to obtain target features, and determining a reply voice according to the target features, where the target mode is one of an offline mode, an online mode, and a hybrid mode, the method further includes: controlling a target device to play the reply voice.

[0025] Optionally, the step of generating a reply text according to the converted text specifically includes: performing entity recognition according to the converted text, performing user intent recognition according to the results of the entity recognition and the converted text, and inputting the results of the user intent recognition, the results of the entity recognition, and the converted text after splicing into a question-answering model to generate a reply text.

[0026] The question-answering model is based on a deep neural network architecture and realizes end-to-end intelligent interaction through natural language understanding and generation technologies: after inputting a user question (i.e., the converted text), the model first performs semantic parsing and context modeling, and uses an attention mechanism to capture the core intent; then it fuses the local knowledge base and real-time cloud data, combines multi-source information to generate a candidate answer set; and then through confidence evaluation and logical consistency verification, dynamically filters the optimal answer (i.e., the reply text) for output.

[0027] Optionally, the step of generating a reply text according to the converted text specifically includes: performing entity recognition according to the converted text, performing user intent recognition according to the results of the entity recognition and the converted text, performing context tracking according to the results of the user intent recognition, the results of the entity recognition, and the converted text, and generating a reply text based on the results of the context tracking, the results of the user intent recognition, the results of the entity recognition, and the converted text based on a multi-turn dialogue strategy.

[0028] The methods of entity recognition, user intent recognition, context tracking, and generating a reply text based on a multi-turn dialogue strategy can all be selected from the prior art.

[0029] Performing entity recognition according to the converted text specifically includes: first, constructing a labeled corpus containing various entities (such as person names, place names, organization names, etc.). Then, using machine learning algorithms, such as feature-based classification algorithms or neural network models. For the feature-based method, features such as part of speech, word frequency, and context are extracted from the text, and these features are matched and compared with the entity features in the labeled corpus. If a neural network model is used, the converted text is vectorized and input, and the semantic and syntactic information in the text is learned through a multi-layer network to identify the entities therein. During the recognition process, some rules and constraints can also be used to improve the accuracy, such as defining the boundaries of entities and excluding some interfering words.

[0030] User intention recognition is performed based on the results of entity recognition and the transformed text, specifically including: First, analyze the content recognized by entity recognition to clarify the key entities and their relationships. Then, in combination with the complete transformed text, examine information such as the vocabulary and grammatical structure in the text. For example, if entities such as "hotel" and "date" are recognized and the word "reservation" appears in the text, it may imply the intention to reserve a hotel. An intention classification model can be constructed and trained using a large amount of labeled corpus. Based on entities, keywords, semantic information, etc. in the text, the model classifies user intentions into different categories. Techniques such as semantic role labeling can also be adopted to further analyze the text semantics to determine whether the user is querying, requesting an operation, or expressing a certain attitude, etc., so as to accurately recognize the user intention.

[0031] Context tracking is performed based on the results of user intention recognition, entity recognition, and the transformed text, specifically including: First, determine the key elements in the text based on the results of entity recognition, and these elements are important components of the context. For example, in a dialogue scenario, entities such as the recognized characters and locations can construct a basic topic framework. Then, in combination with the results of user intention recognition, clarify the direction and purpose of the dialogue. If the intention is to query information about a certain product, the subsequent text content should focus on product-related attributes, evaluations, etc. For the transformed text, analyze the logical relationships between its sentences, such as cause and effect, progression, parallelism, etc. When it is found that the intention changes from querying product information to comparing different products, the product names recognized by entity recognition, etc. become the key basis for tracking this change. A data structure can be established to store the entities, intentions, and text fragments of each round, facilitating backtracking at any time. At the same time, use semantic understanding techniques to ensure that even if the entity expressions change, as long as the semantic associations still point to the same topic, the context can be accurately tracked. For example, different expressions of the same location name can be associated with the correct context related to the geographical location. By continuously updating this storage structure and based on the comprehensive analysis of the above results, effective context tracking is achieved.

[0032] Based on the multi-round dialogue strategy, reply text generation is performed according to the results of context tracking, user intention recognition, entity recognition, and the converted text. Specifically, first, the results of context tracking can clarify the flow and direction of the dialogue and understand the key information of the previous communication. The results of entity recognition provide the core elements in the text, such as people, things, locations, etc., which are important materials for constructing the reply content. The target of the reply is determined according to the results of user intention recognition, such as providing information, seeking confirmation, or giving suggestions, etc. Then, in-depth analysis is carried out on the converted text to grasp the overall semantics. When generating the reply text, these results are integrated. If entity recognition shows that the dialogue focuses on "tourist attractions", the user intention is to obtain scenic spot recommendations, and context tracking shows that information such as budget and preferences was mentioned previously, then the reply text can be "Based on the budget and your preference for natural landscapes mentioned previously, I recommend [specific scenic spot], where the scenery is beautiful and the ticket price meets your budget." Using natural language generation technology, relevant entities, responses to intentions, etc. are combined according to grammar rules and language habits, while ensuring that the reply is semantically coherent with the entire dialogue context, with appropriate words and clear expressions, so as to achieve effective interaction in the multi-round dialogue with the user.

[0033] Optionally, in the target mode, the steps of feature extraction from the target voice data to obtain the target feature and determining the reply voice according to the target feature include: S211: If the target mode is the hybrid mode, Mel-frequency cepstral coefficient features, voiceprint features, and context semantic features are extracted from the target voice data to obtain the target feature; Specifically, when the target mode is the hybrid mode, this means that the cloud 120 and the target device 110 cooperate with each other, combining the advantages of offline and online. Therefore, through the target device 110, Mel-frequency cepstral coefficient features, voiceprint features, and context semantic features are extracted from the target voice data to obtain the target feature.

[0034] S212: Entity recognition, user intention recognition, context tracking, and reply text generation based on the multi-round dialogue strategy are performed according to the target feature to determine entity data and first data; Specifically, entity recognition is performed according to the target feature, and the data obtained from entity recognition is used as entity data. User intention recognition, context tracking, and reply text generation based on the multi-round dialogue strategy are performed according to the target feature and the entity data, and the data generated when generating the reply text is used as the first data. That is to say, the first data can be the reply text or the intermediate data for reply text generation based on the multi-round dialogue strategy.

[0035] S213: Obtain second data from the cloud according to the target feature, or obtain second data from the cloud according to the target feature and the entity data; Specifically, send the target feature to the cloud, and obtain the data generated by the cloud based on the target feature to generate a response text (named second data) from the cloud, or send the target feature and the entity data to the cloud, and obtain the data generated by the cloud based on the target feature and the entity data to generate a response text (named second data) from the cloud.

[0036] S214: Fuse the first data and the second data to determine the target text; Specifically, fuse the first data, and determine the target text according to the fused data. That is to say, the first data and the second data are data with the same expressed meaning, so that they can be fused.

[0037] S215: Perform speech synthesis according to the target text to obtain the response speech.

[0038] In this embodiment, by supporting the collaborative work of multiple modes such as offline mode, online mode, and hybrid mode, the voice interaction is ensured to be always available. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be performed when the network is unavailable in the existing single online mode; while the online mode can make full use of the powerful computing resources and data resources of the cloud to improve the accuracy and diversity of the voice interaction; the hybrid mode combines the advantages of offline and online, further optimizing the interaction effect. In the process of obtaining the target voice data and performing feature extraction to determine the response speech, while ensuring always available, try to use online resources to improve the accuracy and diversity of the voice interaction, thereby enhancing the user experience. In the hybrid mode, the target feature is obtained by extracting Mel-frequency cepstral coefficient features, voiceprint features, and context semantic features from the target voice data. This multi-feature extraction method helps to accurately grasp the multi-faceted characteristics of the voice. Furthermore, operations such as entity recognition and intent recognition based on the target feature can accurately understand the user's needs and determine the entity data and the first data. Obtain the second data from the cloud (based on the target feature or the target feature and the entity data), then fuse the first data and the second data to determine the target text, and finally synthesize the response speech. This process can make full use of the advantages of cloud resources and local processing, ensuring both the accuracy and richness of the data, and adapting to the requirements of different application scenarios, effectively improving the intelligence, flexibility, and accuracy of the voice interaction system, thereby providing high-quality voice responses for users.

[0039] In one embodiment, the step of obtaining the target voice data includes: S11: Obtain the voice collected by the microphone array in the target device as the initial voice; Specifically, the voice collected by the microphone array can be directly obtained from the microphone array of the microphone array in the target device, and this voice is used as the initial voice; the voice collected by the microphone array in the target device is stored in the local storage space of the target device, and the voice collected by the microphone array in the target device is obtained from the local storage space, and this voice is used as the initial voice.

[0040] The microphone array can be set according to requirements and is not limited here.

[0041] S12: Perform noise reduction processing, frame segmentation processing, and endpoint detection on the initial voice to obtain the target voice data.

[0042] Specifically, perform noise reduction processing on the initial voice, perform frame segmentation processing on the data obtained from the noise reduction processing, and perform endpoint detection on the data obtained from the frame segmentation processing.

[0043] The purpose of noise reduction processing is to remove the noise components in the voice. The Wiener filter can be used. It minimizes the mean square error filtering based on the statistical characteristics of the signal and the noise. The spectral subtraction method assumes the composition relationship of the power spectrum of the noisy voice and subtracts the estimated noise power spectrum to reduce noise. These methods can improve the clarity of the voice, reduce the interference of noise on speech recognition, improve the recognition accuracy, and also enhance the intelligibility of the voice, which is of great significance in various speech applications.

[0044] Frame segmentation processing is to divide the voice into segments according to a certain frame length and frame shift. The frame length is often taken as 20 - 30 milliseconds, and the frame shift is 10 - 15 milliseconds. The specific number of frames is determined according to the sampling frequency. Windowing operations also need to be performed, and Hamming windows are commonly used. Frame segmentation processing is convenient for extracting features such as MFCC (Mel Frequency Cepstral Coefficients). And due to the short-term stationarity of the voice, it helps to better utilize this characteristic, thereby improving the overall performance of tasks such as speech recognition and speaker recognition.

[0045] The methods of endpoint detection include: energy-based methods and zero-crossing rate-based methods. For energy-based methods, calculate the energy of the speech frame. The speech energy is relatively low in the silent segment, while it is relatively high in the speech segment. Set an energy threshold. When the frame energy exceeds this threshold, it is considered the start of the speech. When the frame energy is lower than the threshold for a certain period of time, it is considered the end of the speech. For zero-crossing rate-based methods, the zero-crossing rate refers to the number of times the speech signal crosses the horizontal axis (zero level). The zero-crossing rate of the silent part is relatively low, while the zero-crossing rate of the speech part is relatively high due to more waveform changes. Set an appropriate zero-crossing rate threshold to detect the endpoints of the speech.

[0046] The noise reduction process in this embodiment improves speech clarity and intelligibility, reduces speech recognition errors; frame processing facilitates feature extraction and adapts to the speech characteristics to improve system performance; endpoint detection reduces unnecessary processing, saves resources, and improves the accuracy of tasks such as speech recognition and synthesis.

[0047] In one embodiment, the steps of extracting features from the target speech data to obtain target features and determining the response speech according to the target features in the target mode include: S221: If the target mode is the offline mode, perform context semantic feature extraction on the target speech data to obtain the target features; Specifically, if the target mode is the offline mode, this means that the target device independently completes the speech interaction. Therefore, through the target device, perform context semantic feature extraction on the target speech data to obtain the target features.

[0048] S222: Perform entity recognition, user intention recognition, context tracking, and generate a first text based on a multi-round dialogue strategy according to the target features to obtain the target text; Specifically, through the target device, perform entity recognition, user intention recognition, context tracking, and generate a response text based on a multi-round dialogue strategy according to the target features, use the response text as the first text, and use the first text as the target text.

[0049] S223: Perform speech synthesis on the target text to obtain the response speech.

[0050] Specifically, through the target device, use a model trained based on TTS (Text To Speech, speech synthesis technology) to perform text-to-speech conversion on the target text, and use the speech as the response speech.

[0051] In this embodiment, when the target mode is the offline mode, the target device independently completes the speech interaction, thus ensuring that the speech interaction is always available in the case of a poor network state.

[0052] In one embodiment, the steps of extracting features from the target speech data to obtain target features and determining the response speech according to the target features in the target mode further include: S231: If the target mode is the online mode, perform Mel-frequency cepstral coefficient feature, voiceprint feature extraction, and context semantic feature extraction on the target speech data to obtain the target features; Specifically, if the target mode is the online mode, it means that the cloud 120 is the main and the target device 110 is the auxiliary to complete voice interaction. Among them, through the target device 110, Mel Frequency Cepstral Coefficient (MFCC) features and voiceprint features are extracted according to the target voice data, context semantic features are extracted according to the target voice data, and the extracted MFCC features, voiceprint features and context semantic features are used as the target features.

[0053] First, pre-emphasis, framing, and windowing operations are performed on the voice signal. Then, the fast Fourier transform (FFT) is performed to obtain the spectrum. Next, the spectrum is converted to the Mel scale, and the output energy of the Mel filter bank is calculated. Then, the logarithm is taken to obtain the log energy. Finally, the discrete cosine transform (DCT) is performed, and the DCT coefficients are taken as the MFCC features. These coefficients can better reflect the characteristics of the voice signal and have a wide range of applications in the field of speech recognition and other fields.

[0054] S232: Perform voice metric evaluation according to the target features to obtain metric data, where the voice metrics are voice complexity metrics and / or voice confidence metrics; Specifically, perform voice metric evaluation on the target features, take the evaluated voice metrics and metric values as associated data, and take all associated data as metric data.

[0055] For voice complexity, in terms of acoustic features, signal processing techniques are used to analyze parameters such as spectrum, energy, fundamental frequency, and formants, statistics of speech rate and pauses are used to evaluate rhythm, and signal-to-noise ratio is measured to determine the impact of background noise. Linguistically, the lexical diversity is statistically analyzed, and the syntactic structure and semantic depth are analyzed. In the analysis of interaction behavior, the number of dialogue turns, coherence, error recovery ability, and user intention diversity are considered. It can also be tested through experiments such as benchmark test sets and A / B tests, combined with subjective scoring, while taking into account multi-modal factors such as emotional intonation and multilingual mixing.

[0056] When evaluating voice confidence, first determine the acoustic confidence, use an acoustic model (such as HMM, etc.) to calculate the posterior probability of phonemes or words, and compare the probability gap of competing paths. At the language model level, calculate the probability of the text in the language model and check the lexical coverage. When combining context, verify the rationality of the result based on the dialogue history and adapt to specific domain knowledge. For real-time speech, calculate the confidence through streaming processing and set a threshold. Verify the confidence through manual annotation comparison, A / B tests, and confusion matrix analysis, and adopt strategies such as rejection, multiple candidates, or active confirmation for low confidence.

[0057] The calculation methods of voice complexity and voice confidence can both be selected from the existing technologies and are not limited here.

[0058] For example, the speech confidence evaluation method: calculate the phoneme posterior probability and the optimal-competing path probability difference based on the acoustic model (HMM / DNN), calculate the word sequence probability and OOV vocabulary coverage analysis by the language model, and combine context verification with dialogue state tracking and domain knowledge graph matching. Dynamically calculate the confidence value under the streaming architecture, set a multi-level threshold trigger mechanism, optimize the confidence model through confusion matrix analysis and manually labeled data, and execute the rejection / rerank of multiple candidates / secondary confirmation strategy for low-confidence results to achieve end-to-end confidence quantification (the F1 value in typical scenarios is increased by 23%).

[0059] OOV vocabulary (Out-Of-Vocabulary) refers to words that are not included in the preset vocabulary of a speech recognition or natural language processing system, such as new words, technical terms, dialects, proper nouns, or rare expressions. Since the system cannot directly match or generate such words, it may lead to recognition errors or semantic deviations. Handling OOV usually relies on sub-word units (such as phonemes), dynamic vocabulary expansion, context speculation, or improvement of the language model generalization ability to mitigate its impact on the recognition accuracy.

[0060] For example, the calculation of speech complexity adopts a hierarchical quantization model: extract acoustic parameters such as spectrum, fundamental frequency, and formant through signal processing, statistically evaluate the rhythm complexity by speech rate / pause interval, and quantify the noise interference in combination with the signal-to-noise ratio; calculate the vocabulary diversity index, syntactic tree depth, and semantic embedding distance based on NLP; synchronously analyze the dialogue turn switching frequency, intention jump entropy value, and error correction rate; finally, fuse the multi-modal benchmark test scores (such as the ASR error rate) and subjective evaluation weights to construct a comprehensive complexity coefficient covering three-dimensional features of acoustics - language - interaction.

[0061] S233: Based on the data within the preset range, judge whether to call the cloud according to the metric data, and obtain a first result; Optionally, the speech metric is a speech confidence metric. If the speech confidence in the metric data is within the first range of the preset range data, it means that there is no need to call the cloud at this time, and the first result is determined to be no; if the speech confidence in the metric data is outside the first range of the preset range data, it means that the cloud needs to be called at this time, and the first result is determined to be yes.

[0062] Optionally, the speech metric is a speech complexity metric. If the speech complexity in the metric data is within the second range of the preset range data, it means that there is no need to call the cloud at this time, and the first result is determined to be no; if the speech complexity in the metric data is outside the second range of the preset range data, it means that the cloud needs to be called at this time, and the first result is determined to be yes.

[0063] Optionally, the voice metrics are a voice complexity metric and a voice confidence metric. If the voice confidence in the metric data is within the first range of the preset range data, and the voice complexity in the metric data is within the second range of the preset range data, this means that the cloud does not need to be called, and the first result is determined to be no. Otherwise, it means that the cloud needs to be called, and the first result is determined to be yes at this time.

[0064] S234: If the first result is yes, then according to the target feature, obtain the reply voice from the cloud, or obtain a second text from the cloud as the target text according to the target feature, perform voice synthesis on the target text to obtain the reply voice, or obtain the reply voice from the cloud according to the target feature and the entity data, or obtain a second text from the cloud as the target text according to the target feature and the entity data, perform voice synthesis on the target text to obtain the reply voice, where the entity data is the data determined by the target device according to the target feature; Specifically, if the first result is yes, there are multiple implementation methods: (1) Obtain the reply voice from the cloud according to the target feature; (2) Obtain a second text from the cloud as the target text according to the target feature, perform voice synthesis on the target text to obtain the reply voice; (3) Obtain the reply voice from the cloud according to the target feature and the entity data; (4) Obtain a second text from the cloud as the target text according to the target feature and the entity data, perform voice synthesis on the target text to obtain the reply voice. In (1) and (2), entity data is extracted from the cloud. The cloud can quickly process large-scale data and accurately identify entity data related to the target feature, which helps to obtain a more accurate reply voice or target text, reduce the computing burden of the target device, and improve the overall response speed and accuracy. In (3) and (4), entity data is extracted from the target device, which can better adapt to the device environment and perform targeted processing according to the device characteristics and locally stored data.

[0065] S235: If the first result is no, then perform entity recognition, user intention recognition, context tracking, and generate a first text based on a multi-round dialogue strategy according to the target feature to obtain the target text, and perform voice synthesis on the target text to obtain the reply voice.

[0066] Specifically, if the first result is no, then through the target device, perform entity recognition, user intention recognition, context tracking, and generate a first text based on a multi-round dialogue strategy according to the target feature to obtain the target text, and perform voice synthesis on the target text to obtain the reply voice.

[0067] In this embodiment, when the target mode is online mode, firstly, a plurality of features are extracted from the target voice data to obtain the target features, and then the voice index is evaluated to obtain the index data. This helps to accurately judge the complexity and confidence of the voice, so as to decide whether to call the cloud based on the preset range data. If the cloud is called, the reply voice or target text can be obtained in a variety of ways and then the voice can be synthesized, and the rich resources of the cloud can be utilized; if not called, a series of operations are performed locally based on the target features to generate the target text and synthesize the voice. This method can flexibly utilize local and cloud resources according to the actual voice situation, and provide reply voice efficiently and accurately.

[0068] In one embodiment, the step of fusing the first data and the second data to determine the target text includes: S2141: Based on a preset weight configuration, perform weighted summation on the first data and the second data to obtain third data, and determine the target text according to the third data; or, Based on a preset voting mechanism, the first data and the second data are fused to obtain fourth data, and the target text is determined according to the fourth data.

[0069] Specifically, the weighted summation in step S2341 is to perform weighted averaging on the output probability distribution (ie, the first data and the second data). For example, the preset weight configuration is that the first data accounts for 30% and the second data accounts for 70%.

[0070] Based on a preset voting mechanism, the first data and the second data are merged to obtain the fourth data, and the target text is determined according to the fourth data, which specifically includes: first, defining voting rules. Determine how each element (such as data points, features, semantic units, etc.) in the first data and the second data participates in voting. For example, each element can be set to have the same voting weight, or different weights can be set according to the type and source of the element. Then, the first data and the second data are compared and analyzed at the element level. For each corresponding element, a voting operation is performed according to the voting rules. For example, if the values ​​of the two data on a certain feature are the same, it is considered that a consensus vote is reached; if they are different, which value is more advantageous and obtains more "votes" is determined according to the set rules. Then, the voting results are counted. The voting results of each element are summarized to form the fourth data. The value of each element in the fourth data is the final value determined according to the voting results. Finally, the target text is constructed based on the fourth data. The elements in the fourth data are organized according to specific semantic and grammatical rules to form a logically coherent and clearly expressed target text.

[0071] In this embodiment, weighted summation is performed through preset weight configuration, which can reasonably allocate weights according to factors such as the importance or reliability of data, so that the fusion result (the third data) can more scientifically and accurately reflect the comprehensive value of the two types of data, thereby determining a target text that better meets the requirements. The application of the preset voting mechanism enables the first data and the second data to express their "opinions" like different voters during the fusion process. The fourth data obtained by fusion combines the "voting results" of both, and the target text determined in this way can make full use of the respective advantages of the two types of data, improving the effectiveness and rationality of data fusion.

[0072] In one embodiment, the method further includes: S31: Obtain a mode confirmation signal; The mode confirmation signal is a signal for updating the target mode.

[0073] Specifically, it can be a mode confirmation signal input by the user, or a mode confirmation signal actively triggered by the program implementing this application according to preset conditions. For example, when the current voice interaction type is successfully changed, a mode confirmation signal is actively triggered. Another example is when the network fluctuation between the target device and the cloud exceeds the fluctuation threshold, a mode confirmation signal is actively triggered.

[0074] S32: According to the mode confirmation signal, obtain the current voice interaction type; The current voice interaction type is the type of the task currently undergoing voice interaction, or the type of the task predicted to be carried out in the next voice interaction. The tasks of the current voice interaction include various types. On the one hand, the voice interaction type is the information query type, and the tasks include querying the weather, news, knowledge answering, etc., to meet the user's needs for various types of information. On the other hand, the voice interaction type is the operation instruction type, such as controlling smart home devices, and the tasks include the user turning on the light, adjusting the temperature, etc. through voice commands. There is also the voice interaction type that is entertainment-related, and the tasks include playing music, audiobooks, etc. In addition, the voice interaction type is the auxiliary social interaction type, and the tasks include making voice calls, sending voice messages, etc., to conveniently implement various functions through voice interaction and provide a more natural and efficient interaction experience for users.

[0075] Specifically, when the mode confirmation signal is obtained, obtain the current voice interaction type currently being processed S33: According to the current voice interaction type, determine the first mode set; Specifically, according to the current voice interaction type, the first mode set is determined by using the look-up table method.

[0076] The first mode set includes one or more of the following modes: offline mode, online mode, and hybrid mode.

[0077] For example, when the task corresponding to the current voice interaction type is a task with high privacy requirements, the first mode set only includes the offline mode.

[0078] S34: If the first mode set only includes the offline mode, determine that the target mode is the offline mode; Specifically, if the first mode set only includes the offline mode, which means that voice interaction can only be performed offline through the target device at this time, determine that the target mode is the offline mode.

[0079] S35: If the first mode set includes the online mode and / or the hybrid mode, obtain monitoring data, and determine a second mode set according to the monitoring data, where the monitoring data includes one or more of network connection status, local resource load data, user preference data, and energy consumption constraint data; Specifically, if the first mode set includes the online mode and / or the hybrid mode, which means that voice interaction may be performed between the target device and the cloud at this time, the monitoring data of the target device can be obtained.

[0080] Optionally, determine the second mode set by using a look-up table method according to the monitoring data.

[0081] Optionally, input the monitoring data into a pre-trained mode classification model for classification prediction, use the vector elements in the predicted vector whose values are greater than the preset probability as target elements, and use the modes corresponding to each target element as the second mode set.

[0082] The mode classification model is a pre-trained multi-classification model. The model structure and model training method of the mode classification model can be selected from the prior art.

[0083] The second mode set includes one or more of the offline mode, the online mode, and the hybrid mode.

[0084] Network connection status: Reflects the connection situation between the device and the network, including whether it is connected, connection type (such as WiFi, cellular network), connection stability (signal strength, packet loss rate, etc.), network speed (upload and download rates), etc., and is used to judge the availability and quality of network communication.

[0085] Local resource load data: Refers to the occupancy situation of various resources (CPU, memory, disk I / O, GPU, etc.) of the local device (such as a computer, mobile phone), and shows the degree of resource tension or idleness through quantitative indicators to assist in resource management and task scheduling.

[0086] User preference data: Information about users' preferences and inclinations towards various things. For example, users' operation habits in the application (such as frequently used functions), content preferences (such as favorite music types, news categories), etc., which helps to provide personalized services and experiences.

[0087] Energy consumption constraint data: Limiting information related to the energy consumption of the device. It clarifies the maximum energy consumption that the device can tolerate in different working modes, or the energy consumption upper limit set based on factors such as the remaining battery power, which can be used to optimize the device operation to extend the battery life or meet the energy consumption requirements.

[0088] S36: Perform a union calculation on the first mode set and the second mode set to obtain a third mode set; Specifically, perform a union calculation on the first mode set and the second mode set. If the union is not empty, use the union as the third mode set. If the union is empty, use the offline mode as the third mode set.

[0089] S37: Screen the modes according to a preset screening strategy based on the third mode set to obtain the target mode.

[0090] The preset screening strategy is a pre-set strategy.

[0091] Specifically, first, the preset screening strategy may be set based on priorities. If the network connection is stable and the energy consumption constraint is low, the online mode may have a high priority; if the local resource load is large, the priority of the offline mode is increased. For the hybrid mode, a comprehensive trade-off is required. From the third mode set, first check the network connection status. If it is good and there is no energy consumption concern, the online mode is preferred; if the local resources are tense, the offline mode is preferred. If each factor is relatively balanced, the hybrid mode will be considered. At the same time, combined with the user preference data, such as the user often selects the offline mode, it is preferentially determined as the offline mode under the condition of meeting the basic conditions, so as to obtain the target mode.

[0092] In another embodiment of the present application, steps S35 to S37 may be replaced by: If the first mode set includes the online mode and / or the hybrid mode, obtain the monitoring data, splice the first mode set and the monitoring data, and input them into a pre-trained target model for classification prediction. Select the vector element with the largest value from the predicted vectors, and use the classification category corresponding to the selected vector element as the target mode.

[0093] The target model is a pre-trained multi-classification model. The model structure and training method of the target model can be selected from the prior art.

[0094] In this embodiment, first, by obtaining a mode confirmation signal and thereby obtaining the current voice interaction type to determine the first mode set, the mode range can be initially delimited according to the task requirements. When the first mode set only contains the offline mode, it is directly determined as the target mode, which reflects the efficient processing of simple situations. When the first mode set includes the online mode and / or the hybrid mode, a variety of monitoring data is introduced to determine the second mode set, comprehensively considering factors such as the network connection status, local resource load, user preferences, and energy consumption constraints, making the mode selection more in line with the actual situation. The union calculation of the first and second mode sets is performed to obtain the third mode set, and then the target mode is obtained by screening according to the preset screening strategy, which can comprehensively and accurately determine the target mode that is most suitable for the current voice interaction type and meets the current multiple condition constraints, improving the stability, efficiency, and user satisfaction of the voice interaction experience.

[0095] In one embodiment, the step of screening the mode according to the preset screening strategy based on the third mode set to obtain the target mode includes: S371: Obtain the target mode as the old mode; S372: Screen the mode according to the preset screening strategy based on the third mode set to obtain a new mode; Specifically, screen the mode according to the preset screening strategy based on the third mode set, and use the screened mode as the new mode.

[0096] S373: Update the target mode according to the new mode, and if the old mode is different from the updated target mode, then based on the seamless switching strategy, perform switching processing according to the old mode and the updated target mode.

[0097] The seamless switching strategy aims to achieve a smooth transition between different states, systems, or modes. At the technical level, the switching delay is reduced by optimizing algorithms, preloading resources, etc. In terms of services, the continuity of data transmission and the continuous availability of functions are ensured. For users, the entire switching process is rapid and imperceptible, like a continuous and uninterrupted experience.

[0098] Specifically, if the old mode is different from the updated target mode, this means that a mode switch is required at this time.

[0099] Based on the seamless switching strategy, switching processing is performed according to the old mode and the updated target mode, which specifically includes: If switching from the offline mode to the online mode (offline → online), at this time, the intermediate results processed locally, such as speech feature vectors, etc., need to be uploaded to the cloud to continue processing tasks in the cloud to ensure the continuity of the tasks. If switching from the online mode to the offline mode (online → offline), then some parameters of the cloud model need to be cached locally. In the special case of network disconnection, the local degraded model can be enabled to maintain the operation of basic functions. In terms of service hot switching, for automatic speech recognition (ASR), when switching from the local Kaldi (an open-source speech recognition toolkit containing many speech processing algorithms) to the cloud Whisper model (the Whisper model is an OpenAI speech recognition model), the speech buffer needs to be retained to prevent data loss. For natural language understanding (NLU, Natural Language Understanding), in the end-cloud collaboration, the method of extracting entities locally and processing intent classification in the cloud is adopted. In terms of ensuring the user experience, during the voice conversation switch, streaming processing is used to maintain the speech coherence. For example, the "latency filling" technology similar to Duplex (an artificial intelligence technology) is adopted to achieve seamless switching without perception. At the same time, the screen display mode icon gives a visual hint. If the cloud processing fails, an error recovery operation is performed to fallback to the local fallback response, such as prompting "The network is unstable and has switched to the offline mode". Through these steps, the switching processing between the old mode and the target mode is achieved based on the seamless switching strategy.

[0100] In this embodiment, by obtaining the old mode as the initial state of the target mode, and then screening out the new mode from the third mode set according to the preset screening strategy, this method can flexibly update the target mode according to the actual situation. If the old and new modes are different, switching processing is performed based on the seamless switching strategy. This technical effect can ensure that there will be no interruption or lag during the mode switch, realizing a smooth transition of the mode, so that users cannot feel an obvious switching difference during use, thereby improving the overall user experience and the stability of system operation.

[0101] In one embodiment, the step of concatenating the first mode set and the monitoring data and inputting them into the pre-trained target model for classification prediction, selecting the vector element with the largest value from the predicted vectors, and using the classification category corresponding to the selected vector element as the target mode includes: Obtain the target mode as the old mode; Concatenate the first mode set and the monitoring data and input them into the pre-trained target model for classification prediction, select the vector element with the largest value from the predicted vectors, and use the classification category corresponding to the selected vector element as the new mode; Update the target mode according to the new mode, and if the old mode is different from the updated target mode, perform a switching process based on the seamless switching strategy according to the old mode and the updated target mode.

[0102] In one embodiment, the step of performing a switching process based on the seamless switching strategy according to the old mode and the updated target mode includes: S3731: If the old mode is the offline mode and the updated target mode is the online mode, perform data retention processing on the local voice buffer, and send the intermediate result in the target device to the cloud, where the cloud is used to determine a reply voice or a second text according to the intermediate result; The local voice buffer stores the original voice data collected by the voice input device. It includes acoustic feature information such as the frequency and amplitude of the sound, and may also contain some data that has been preliminarily processed (such as noise reduction and feature extraction in the preprocessing stage). These data provide the basic materials for subsequent voice recognition, semantic understanding, and other operations.

[0103] Among them, the data storage area of the local voice buffer can be marked, and a protection identifier can be set to prevent it from being overwritten or cleared during mode switching. Or a temporary copy storage can be established to copy the voice buffer data to a specific temporary storage area to ensure the complete preservation of the data.

[0104] Encapsulate the intermediate result in JSON format and transmit the encapsulated intermediate result to the cloud through the HTTPS protocol. JSON (JavaScript Object Notation) is a lightweight data interchange format. It organizes data in key-value pairs, is easy for humans to read and write, and is also easy for machines to parse and generate. It is widely used in network data transmission and configuration files. HTTPS (Hypertext Transfer Protocol Secure) adds the SSL / TLS encryption protocol on the basis of HTTP. By encrypting data transmission, it ensures network communication security, such as protecting the privacy information of website users, verifying the authenticity of websites, and is commonly used in scenarios such as finance and e-commerce that require secure communication.

[0105] The data included in the intermediate results is related to the speech processing process. When switching from the old offline mode to the online mode, if the target device has performed a speech processing operation, the intermediate results may include some data during the feature extraction process of the target speech data, such as some extracted speech features. It may also include the preliminary results related to semantic understanding after the preliminary analysis of the speech, such as some keywords or the preliminarily constructed semantic framework. These intermediate results are sent to the cloud, and the cloud can more efficiently determine the reply speech or the second text based on these data, enabling a smooth transition from the offline mode to the online mode and making full use of the phased results obtained from the previous offline processing.

[0106] S3732: If the old mode is the online mode and the updated target mode is the offline mode, obtain the target parameters from the cloud, update the parameters of the local model according to the target parameters, and based on the updated local model, perform the steps of extracting features according to the target speech data to obtain target features and determining the reply speech according to the target features.

[0107] Specifically, integrating the target parameters into the local model according to a predetermined rule may involve operations such as parameter replacement and weight adjustment to complete the parameter update of the local model.

[0108] In this embodiment, when switching from offline to online, the local voice buffer data is retained and the intermediate results are sent to the cloud to obtain a reply; when switching from online to offline, parameters are obtained from the cloud to update the local model, ensuring the coherence and effectiveness of speech processing under different mode switches.

[0109] Please refer to Figure 3 As shown, in one embodiment, a multi-mode collaborative offline-online voice interaction system is provided. The system includes: a cloud and a target device. The target device is communicatively connected to the cloud. The cloud is a remote server, and the target device is configured to implement the steps of the above multi-mode collaborative offline-online voice interaction method.

[0110] In this embodiment, by supporting the collaborative work of multiple modes such as offline mode, online mode, and hybrid mode, the voice interaction is ensured to be always available. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be performed when the network is unavailable in the existing single online mode. The online mode can make full use of the powerful computing resources and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. The hybrid mode combines the advantages of offline and online, further optimizing the interaction effect. In the process of obtaining the target voice data and performing feature extraction to determine the response voice, while ensuring always available, online resources are utilized as much as possible to improve the accuracy of voice interaction and the diversity of responses, thereby enhancing the user experience. In the hybrid mode, by performing Mel Frequency Cepstral Coefficient (MFCC) feature extraction, voiceprint feature extraction, and context semantic feature extraction on the target voice data, the target features are obtained. This multi-feature extraction method helps to accurately grasp various aspects of the voice characteristics. Furthermore, operations such as entity recognition and intent recognition based on the target features can accurately understand the user's needs and determine the entity data and the first data. The second data (based on the target features or the target features and entity data) is obtained from the cloud, and then the first data and the second data are fused to determine the target text. Finally, the response voice is synthesized. This process can make full use of the advantages of cloud resources and local processing, ensuring both the accuracy and richness of the data, and being able to adapt to the requirements of different application scenarios, effectively improving the intelligence, flexibility, and accuracy of the voice interaction system, thus providing high-quality voice responses for users.

[0111] In one embodiment, a target device is proposed. The target device is communicatively connected to the cloud, and the cloud is a remote server. The target device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtain target voice data; In the target mode, perform feature extraction according to the target voice data to obtain target features, and determine the response voice according to the target features, where the target mode is one of the offline mode, online mode, and hybrid mode.

[0112] In this embodiment, by supporting the collaborative work of multiple modes such as offline mode, online mode, and hybrid mode, the voice interaction is ensured to be always available. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be performed when the network is unavailable in the existing single online mode; while the online mode can make full use of the powerful computing resources and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses; the hybrid mode combines the advantages of offline and online, further optimizing the interaction effect. In the process of obtaining the target voice data and performing feature extraction to determine the response voice, while ensuring always available, online resources are utilized as much as possible to improve the accuracy of voice interaction and the diversity of responses, thereby enhancing the user experience. In the hybrid mode, by performing Mel-frequency cepstral coefficient feature, voiceprint feature, and context semantic feature extraction on the target voice data, the target features are obtained. This multi-feature extraction method helps to accurately grasp the multi-faceted characteristics of the voice. Furthermore, operations such as entity recognition and intent recognition based on the target features can accurately understand the user's needs and determine the entity data and the first data. Obtain the second data from the cloud (based on the target features or the target features and the entity data), then fuse the first data and the second data to determine the target text, and finally synthesize the response voice. This process can make full use of the advantages of cloud resources and local processing, ensuring both the accuracy and richness of the data, and being able to adapt to the requirements of different application scenarios, effectively improving the intelligence, flexibility, and accuracy of the voice interaction system, thereby providing high-quality voice responses for users.

[0113] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: Obtain target voice data; In the target mode, perform feature extraction according to the target voice data to obtain target features, and determine the response voice according to the target features, where the target mode is one of the offline mode, online mode, and hybrid mode.

[0114] In this embodiment, by supporting the collaborative work of multiple modes such as offline mode, online mode, and hybrid mode, the voice interaction is ensured to be always available. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, overcoming the drawback that voice interaction cannot be performed when the network is unavailable in the existing single online mode. The online mode can make full use of the powerful computing resources and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. The hybrid mode combines the advantages of offline and online, further optimizing the interaction effect. In the process of obtaining the target voice data and performing feature extraction to determine the response voice, while ensuring always available, online resources are utilized as much as possible to improve the accuracy of voice interaction and the diversity of responses, thereby enhancing the user experience. In the hybrid mode, the target features are obtained by extracting Mel Frequency Cepstral Coefficient (MFCC) features, voiceprint features, and context semantic features from the target voice data. This multi-feature extraction method helps to accurately grasp various aspects of the voice characteristics. Furthermore, operations such as entity recognition and intent recognition based on the target features can accurately understand the user's needs and determine the entity data and the first data. The second data (based on the target features or the target features and the entity data) is obtained from the cloud, and then the first data and the second data are fused to determine the target text. Finally, the response voice is synthesized. This process can make full use of the advantages of cloud resources and local processing, ensuring both the accuracy and richness of the data and adapting to the requirements of different application scenarios, effectively improving the intelligence, flexibility, and accuracy of the voice interaction system, and thus providing high-quality voice responses for users.

[0115] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions on the cloud 120 side and the target device 110 side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0116] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0117] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0118] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A multi-mode collaborative offline and online voice interaction method, characterized in that: The method is applied to a target device, the target device is connected to a cloud for communication, the cloud being a remote server, and the method comprises: Obtain target voice data; In a target mode, feature extraction is performed according to the target voice data to obtain target features, and a reply voice is determined according to the target features, wherein the target mode is one of an offline mode, an online mode and a mixed mode; The step of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features comprises: If the target mode is a mixed mode, extracting Mel frequency cepstral coefficient features, voiceprint features, and context semantic features according to the target speech data to obtain the target features; Perform entity recognition, user intent recognition, context tracking, and response text generation based on a multi-round dialogue strategy according to the target features to determine entity data and first data; Acquire second data from the cloud according to the target feature, or acquire second data from the cloud according to the target feature and the entity data; fusing the first data and the second data to determine a target text; Speech synthesis is performed according to the target text to obtain the reply speech.

2. The multi-mode collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of acquiring target voice data comprises: Acquire the speech collected by the microphone array in the target device as the initial speech; The initial speech is subjected to noise reduction processing, frame processing and endpoint detection to obtain the target speech data.

3. The multi-mode collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features comprises: If the target mode is an offline mode, extracting contextual semantic features according to the target speech data to obtain the target features; Perform entity recognition, user intent recognition, context tracking and generate a first text based on a multi-round dialogue strategy according to the target features to obtain a target text; Speech synthesis is performed according to the target text to obtain the reply speech.

4. The multi-mode collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features, further includes: If the target mode is the online mode, extracting Mel frequency cepstral coefficient features, voiceprint features and context semantic features according to the target speech data to obtain the target features; Performing speech index evaluation according to the target feature to obtain index data, wherein the speech index is a speech complexity index and / or a speech confidence index; Based on the preset range data, judging whether to call the cloud according to the indicator data to obtain a first result; If the first result is yes, then according to the target feature, the reply voice is obtained from the cloud, or, according to the target feature, a second text is obtained from the cloud as the target text, and speech synthesis is performed according to the target text to obtain the reply voice, or, according to the target feature and the entity data, the reply voice is obtained from the cloud, or, according to the target feature and the entity data, a second text is obtained from the cloud as the target text, and speech synthesis is performed according to the target text to obtain the reply voice, wherein the entity data is data determined by the target device according to the target feature; If the first result is no, entity recognition, user intent recognition, context tracking and a first text are generated based on the target features to obtain a target text, and speech synthesis is performed based on the target text to obtain the reply speech.

5. The multi-mode collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of fusing the first data and the second data to determine the target text includes: Based on a preset weight configuration, weighted sum is performed on the first data and the second data to obtain third data, and the target text is determined according to the third data; or, Based on a preset voting mechanism, the first data and the second data are fused to obtain fourth data, and the target text is determined according to the fourth data.

6. The multi-mode collaborative online and offline voice interaction method according to claim 1, characterized in that: The method further comprises: Get mode confirmation signal; Acquire the current voice interaction type according to the mode confirmation signal; Determining a first mode set according to the current voice interaction type; If the first mode set only includes the offline mode, determining the target mode to be the offline mode; If the first mode set includes an online mode and / or a hybrid mode, acquiring monitoring data, and determining a second mode set according to the monitoring data, wherein the monitoring data includes one or more data selected from the group consisting of a network connection status, local resource load data, user preference data, and energy consumption constraint data; Performing a union calculation on the first pattern set and the second pattern set to obtain a third pattern set; The target pattern is obtained by performing pattern screening according to the third pattern set and a preset screening strategy.

7. The multi-mode collaborative online and offline voice interaction method according to claim 6, characterized in that: The step of performing pattern screening according to the third pattern set according to a preset screening strategy to obtain the target pattern includes: Obtaining the target mode as the old mode; Performing pattern screening according to the third pattern set according to a preset screening strategy to obtain a new pattern; The target mode is updated according to the new mode, and if the old mode is different from the updated target mode, a switching process is performed according to the old mode and the updated target mode based on a seamless switching strategy.

8. The multi-mode collaborative online and offline voice interaction method according to claim 7, characterized in that: The step of performing switching processing according to the old mode and the updated target mode based on the seamless switching strategy includes: If the old mode is an offline mode, and the updated target mode is an online mode, data retention processing is performed on the local voice buffer, and the intermediate result in the target device is sent to the cloud, wherein the cloud is used to determine the reply voice or the second text according to the intermediate result; If the old mode is an online mode and the updated target mode is an offline mode, the target parameters are obtained from the cloud, the parameters of the local model are updated according to the target parameters, and based on the updated local model, the steps of extracting features according to the target voice data, obtaining target features, and determining the reply voice according to the target features are performed.

9. A multi-mode collaborative online and offline voice interaction system, characterized in that: The system includes: a cloud and a target device, the target device is communicatively connected to the cloud, the cloud is a remote server, and the target device is configured to implement the steps of the multi-modal collaborative offline and online voice interaction method as described in any one of claims 1 to 8.

10. A target device, characterized in that: The target device is communicatively connected to the cloud, which is a remote server. The target device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-modal collaborative offline and online voice interaction method as described in any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multi-modal collaborative offline and online voice interaction method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Voice conversation method and system

    CN111833880A

  • Processing method, terminal equipment and storage medium

    CN113326018A

  • Voice system based on off-line mode and on-line mode for intelligent furniture

    CN113643711A

  • Voice interaction method, voice interaction device, vehicle and readable storage medium

    CN115394300A

  • Dialogue system off-line and on-line fusion application method and system

    CN115795017A

Cited By

  • Automatic speech recognition method and device, computer equipment and medium

    CN120673762A

  • Wearable device, semantic recognition method and device thereof, storage medium and computer program product

    CN121187530A