Multi-mode collaborative online and offline voice interaction method, system, device and medium
Through a multi-mode collaborative offline and online voice interaction method, combining offline and online modes, utilizing cloud resources and local processing, it solves the availability problem of traditional voice interaction systems when the network is unstable, and achieves high-accuracy and diverse voice interaction effects.
Patent Information
- Application Number
- CN202510539496.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Traditional voice interaction systems have low voice recognition accuracy and limited response content in pure offline mode, while pure online mode is highly dependent on network stability, resulting in the inability to guarantee availability when the network is interrupted or the latency is high.
A multi-modal collaborative offline and online voice interaction method is adopted, combining offline mode, online mode and hybrid mode. It works collaboratively with the target device and the cloud, utilizing the powerful computing and data resources of the cloud to perform feature extraction and reply voice generation, ensuring that voice interaction is always available.
When network conditions are poor or there is no network connection, offline mode enables voice interaction. Hybrid mode combines the advantages of offline and online modes, improving the accuracy of voice interaction and the diversity of responses, thereby enhancing the user experience.
Smart Images

Figure CN120071937B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of speech recognition technology and artificial intelligence technology, and in particular to a multi-modal collaborative offline and online speech interaction method, system, device and medium. Background Art
[0002] Traditional voice interaction systems often use a single, fixed mode (offline or online) for response processing. Limited by local computing power and knowledge base size, the purely offline mode suffers from the following drawbacks: low voice recognition accuracy and limited response content, making it difficult to meet the needs of complex scenarios. The purely online mode, on the other hand, is highly dependent on network stability and suffers from the drawback of being unable to guarantee constant availability in the event of network outages or high latency. Existing technologies address the drawbacks of both offline and online modes when constant availability cannot be guaranteed. Summary of the Invention
[0003] Based on this, it is necessary to implement voice interaction in a single mode of the existing technology, which cannot guarantee that it is always available, to overcome the technical problems of the disadvantages of the separate existence of offline mode and online mode. A multi-mode collaborative offline and online voice interaction method, system, equipment and medium are proposed.
[0004] In a first aspect, a multi-modal collaborative offline and online voice interaction method is provided. The method is applied to a target device, the target device is in communication with a cloud, and the cloud is a remote server. The method includes:
[0005] Obtain target voice data;
[0006] In the target mode, feature extraction is performed based on the target voice data to obtain target features, and a reply voice is determined based on the target features, wherein the target mode is one of an offline mode, an online mode, and a mixed mode.
[0007] In the second aspect, a multi-mode collaborative offline-online voice interaction system is provided, which includes: a cloud and a target device, the target device is communicatively connected to the cloud, the cloud is a remote server, and the target device is configured to implement the steps of the above-mentioned multi-mode collaborative offline-online voice interaction method.
[0008] In a third aspect, a target device is provided, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor, and when the processor executes the computer program, the steps of the above-mentioned multi-modal collaborative offline and online voice interaction method are implemented.
[0009] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned multi-modal collaborative offline and online voice interaction method are implemented.
[0010] The multimodal collaborative offline and online voice interaction method, system, device, and medium of this application ensures that voice interaction is always available by supporting the coordinated operation of multiple modes: offline, online, and hybrid. When network conditions are poor or no network connection is available, offline mode can be enabled for voice interaction, overcoming the drawback of the existing single online mode, which prevents voice interaction when the network is unavailable. The online mode, on the other hand, fully utilizes the powerful computing and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. The hybrid mode further optimizes the interaction effect by combining the advantages of both offline and online resources. In the process of acquiring target voice data and performing feature extraction to determine the response voice, while ensuring that it is always available, online resources are utilized to improve the accuracy and diversity of voice interaction, thereby enhancing the user experience. In the hybrid mode, target features are obtained by extracting Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features from the target voice data. This multi-feature extraction approach helps accurately grasp the various characteristics of the voice. Entity recognition and intent recognition based on the target features can then accurately understand user needs and determine entity data and primary data. The second data (based on target features or target features and entity data) is obtained from the cloud, and the first and second data are integrated to determine the target text, and finally the reply voice is synthesized. This process can fully utilize the advantages of cloud resources and local processing, which not only ensures the accuracy and richness of the data, but also can adapt to the needs of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thereby providing users with high-quality voice responses. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] in:
[0013] Figure 1 This is a diagram illustrating an application environment of a multi-modal collaborative offline and online voice interaction method according to an embodiment;
[0014] Figure 2 Flowchart of a multi-modal collaborative offline and online voice interaction method in one embodiment;
[0015] Figure 3 Flowchart of a multi-modal collaborative offline and online voice interaction method in one embodiment;
[0016] Figure 4 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0018] The multi-mode collaborative online and offline voice interaction method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, a target device 110 communicates with a cloud 120 via a network. The target device is in communication with the cloud 120, which is a remote server.
[0019] The target device 110 is configured to implement the multi-modal collaborative offline-online voice interaction method of the present application, specifically including: obtaining target voice data; in the target mode, performing feature extraction based on the target voice data to obtain target features, and determining the reply voice based on the target features, wherein the target mode is one of the offline mode, the online mode and the mixed mode.
[0020] This application ensures that voice interaction is always available by supporting multiple modes such as offline mode, online mode, and hybrid mode to work together. When the network condition is poor or there is no network connection, the offline mode can be enabled for voice interaction, which overcomes the disadvantage of the existing single online mode where voice interaction cannot be performed when the network is unavailable; the online mode can make full use of the powerful computing resources and data resources of the cloud to improve the accuracy of voice interaction and the diversity of responses; the hybrid mode combines the advantages of offline and online to further optimize the interaction effect. In the process of obtaining the target voice data and performing feature extraction to determine the reply voice, while ensuring that it is always available, try to use online resources to improve the accuracy of voice interaction and the diversity of responses, thereby improving the user experience.
[0021] It can be understood that in offline mode, the target device 110 completes the voice interaction independently; in online mode, the cloud 120 is the main and the target device 110 is the auxiliary to complete the voice interaction; in hybrid mode, the cloud 120 and the target device 110 cooperate with each other, combining the advantages of offline and online.
[0022] The target device 110 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. Figure 4 The computer device.
[0023] The cloud 120 can be implemented using an independent server or a server cluster consisting of multiple servers.
[0024] The present invention is described in detail below through specific examples.
[0025] See also Figure 2 As shown, Figure 2 A flowchart of a multi-modal collaborative offline and online voice interaction method provided in an embodiment of the present invention is provided. The method is applied to a target device, the target device is in communication with a cloud, and the cloud is a remote server. The method includes:
[0026] S1: Acquire target voice data;
[0027] Specifically, the user inputs voice into the target device through a microphone, and the voice is directly used as the target voice data, or the voice is processed (for example, noise reduction) and the processed data is used as the target voice data.
[0028] S2: In a target mode, feature extraction is performed based on the target voice data to obtain target features, and a reply voice is determined based on the target features, wherein the target mode is one of an offline mode, an online mode, and a mixed mode.
[0029] Specifically, a predetermined target mode is obtained; when the target mode is an offline mode, the steps of extracting features based on the target voice data, obtaining target features, and determining the reply voice based on the target features are implemented only through the target device; when the target mode is an online mode, the cloud 120 is used as the main method and the target device 110 is used as the auxiliary method to complete the voice interaction; in a hybrid mode, the cloud 120 and the target device 110 cooperate with each other, combining the advantages of offline and online.
[0030] The target mode is adaptively determined based on monitoring data such as network connectivity, local resource load, user preferences, energy consumption constraints, and the current voice interaction type. In other words, the entire process dynamically adjusts and determines the target mode based on multiple factors, thereby achieving multi-mode collaboration.
[0031] When extracting features from the target speech data, at least contextual semantic features are extracted, and all extracted features are used as target features. A model trained using ASR (speech Recognition) is used to perform speech-to-text conversion on the target speech data, and a reply text is generated based on the converted text. A model trained using TTS (Text To Speech) is used to convert the reply text into speech, and the resulting speech is used as the reply speech.
[0032] The method further includes: performing feature extraction according to the target voice data in the target mode to obtain target features, and determining a reply voice according to the target features, wherein the target mode is one of an offline mode, an online mode and a mixed mode, and further includes: controlling the target device to play the reply voice.
[0033] Optionally, the step of generating a reply text based on the converted text specifically includes: performing entity recognition based on the converted text, performing user intent recognition based on the result of entity recognition and the converted text, splicing the result of user intent recognition, the result of entity recognition and the converted text, and inputting the result into the question-answering model to generate a reply text.
[0034] The question-answering model is based on a deep neural network architecture and achieves end-to-end intelligent interaction through natural language understanding and generation technology: after inputting the user question (that is, the converted text), the model first performs semantic analysis and context modeling, using the attention mechanism to capture the core intent; then it integrates the local knowledge base with real-time cloud data, and combines multi-source information to generate a set of candidate answers; then, through confidence assessment and logical consistency verification, it dynamically selects the optimal answer (that is, the reply text) for output.
[0035] Optionally, the step of generating a reply text based on the converted text specifically includes: performing entity recognition based on the converted text, performing user intent recognition based on the result of entity recognition and the converted text, performing context tracking based on the result of user intent recognition, the result of entity recognition and the converted text, and based on a multi-round dialogue strategy, generating a reply text based on the result of context tracking, the result of user intent recognition, the result of entity recognition and the converted text.
[0036] Methods for entity recognition, user intent recognition, context tracking, and generating response text based on multi-round dialogue strategies can all be selected from existing technologies.
[0037] Entity recognition based on the converted text involves first building an annotated corpus containing various entities (such as names of people, places, and organizations). Next, machine learning algorithms, such as feature-based classification algorithms or neural network models, are used. Feature-based methods extract features such as part of speech, word frequency, and context from the text and compare these features with entity features in the annotated corpus. If a neural network model is used, the converted text is vectorized and input, and the semantic and grammatical information in the text is learned through a multi-layer network to identify the entities. Rules and constraints can also be used to improve accuracy during the recognition process, such as defining entity boundaries and excluding interfering words.
[0038] User intent is identified based on the entity recognition results and the converted text. Specifically, the following steps are performed: first, the content of the entity recognition is analyzed to identify the key entities and their relationships. Then, combined with the converted complete text, information such as vocabulary and grammatical structure is examined. For example, if entities such as "hotel" and "date" are identified, and the text also contains words such as "booking," this may imply the intention to book a hotel. An intent classification model can be constructed and trained using a large amount of labeled corpus. The model categorizes user intent into different categories based on entities, keywords, and semantic information in the text. Techniques such as semantic role labeling can also be used to further analyze the semantics of the text to determine whether the user is querying, requesting an action, or expressing a certain attitude, thereby accurately identifying user intent.
[0039] Context tracking is performed based on the results of user intent recognition, entity recognition, and the converted text. Specifically, the following steps are performed: First, key elements in the text are identified based on the entity recognition results. These elements are crucial components of the context. For example, in a conversation scenario, recognized entities such as people and places can form a basic framework for the topic. Then, combined with the user intent recognition results, the direction and purpose of the conversation are clarified. If the intent is to query information about a specific product, the subsequent text content should focus on product-related attributes and reviews. For the converted text, logical relationships between sentences are analyzed, such as causal relationships, progression, and parallelism. If the intent shifts from querying product information to comparing different products, entity recognition, such as product names, becomes key evidence for tracking this change. A data structure can be established to store entities, intents, and text snippets from each round for easy retrieval. Furthermore, semantic understanding technology ensures that even if the entity representation changes, the context can be accurately tracked as long as the semantic associations still point to the same topic. For example, different representations of the same place name can be associated with the correct geographical context. By continuously updating this storage structure and comprehensively analyzing the above results, effective context tracking is achieved.
[0040] Based on a multi-turn conversation strategy, responses are generated based on context tracking, user intent recognition, entity recognition, and converted text. Specifically, context tracking clarifies the flow and direction of the conversation and identifies key information from previous exchanges. Entity recognition provides key elements in the text, such as people, objects, and locations, which are crucial for constructing responses. User intent recognition determines the goal of the response, such as providing information, seeking confirmation, or offering advice. The converted text is then analyzed in depth to grasp the overall semantics. These findings are integrated when generating responses. For example, if entity recognition indicates that the conversation revolves around "tourist attractions," the user's intent is to obtain recommendations, and context tracking indicates that budget and preferences were previously mentioned, the response could be, "Based on your previously mentioned budget and preference for natural landscapes, I recommend [specific attraction]; it has beautiful scenery and the ticket price is within your budget." Natural language generation technology is used to combine relevant entities and responses to the intent according to grammatical rules and linguistic conventions. The response is also ensured to be semantically coherent with the overall conversation context, with appropriate wording and clear expression, enabling effective interaction with the user throughout the multi-turn conversation.
[0041] Optionally, in the target mode, the steps of performing feature extraction based on the target voice data to obtain target features, and determining the reply voice based on the target features include:
[0042] S211: If the target mode is a mixed mode, extract the Mel-frequency cepstral coefficient feature, voiceprint feature, and context semantic feature based on the target speech data to obtain the target feature;
[0043] Specifically, the target mode is a hybrid mode, which means that the cloud 120 and the target device 110 cooperate with each other, combining the advantages of offline and online. Therefore, through the target device 110, the Mel-frequency cepstral coefficient feature, voiceprint feature extraction and context semantic feature extraction are performed according to the target voice data to obtain the target feature.
[0044] S212: Perform entity recognition, user intent recognition, context tracking, and response text generation based on a multi-round dialogue strategy according to the target feature to determine entity data and first data;
[0045] Specifically, entity recognition is performed based on the target features, and the data obtained from entity recognition is used as entity data. User intent recognition and context tracking are then performed based on the target features and entity data, and reply text is generated based on a multi-turn dialogue strategy. The data generated from the reply text is used as the first data. In other words, the first data can be the reply text or intermediate data generated from the reply text based on the multi-turn dialogue strategy.
[0046] S213: Acquire second data from the cloud based on the target feature, or acquire second data from the cloud based on the target feature and the entity data;
[0047] Specifically, the target feature is sent to the cloud, and data (named as second data) generated by the cloud when the reply text is generated based on the target feature is obtained from the cloud, or the target feature and the entity data are sent to the cloud, and data (named as second data) generated by the cloud when the reply text is generated based on the target feature and the entity data is obtained from the cloud.
[0048] S214: Fusing the first data and the second data to determine a target text;
[0049] Specifically, the first data and the second data are fused, and the target text is determined based on the fused data. That is, the first data and the second data are data that express the same meaning, so they can be fused.
[0050] S215: Perform speech synthesis according to the target text to obtain the reply speech.
[0051] This embodiment ensures that voice interaction is always available by supporting multiple modes—offline, online, and hybrid—working together. When network conditions are poor or there's no connection, offline mode can be enabled for voice interaction, overcoming the drawback of the existing single online mode, which prevents voice interaction when the network is unavailable. Online mode, on the other hand, fully utilizes the powerful computing and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. Hybrid mode further optimizes the interaction experience by combining the advantages of both offline and online modes. While ensuring consistent availability during the acquisition of target voice data and feature extraction to determine the response voice, online resources are utilized to maximize the accuracy and diversity of voice interaction, thereby enhancing the user experience. In hybrid mode, target features are extracted from the target voice data using Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features. This multi-feature extraction approach helps accurately capture multiple aspects of speech characteristics. Entity recognition and intent recognition based on the target features can then accurately understand user needs and determine entity data and primary data. The second data (based on target features or target features and entity data) is obtained from the cloud, and the first and second data are integrated to determine the target text, and finally the reply voice is synthesized. This process can fully utilize the advantages of cloud resources and local processing, which not only ensures the accuracy and richness of the data, but also can adapt to the needs of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thereby providing users with high-quality voice responses.
[0052] In one embodiment, the step of obtaining target voice data includes:
[0053] S11: Acquire the speech collected by the microphone array in the target device as the initial speech;
[0054] Specifically, the voice collected by the microphone array in the target device can be directly obtained from the microphone array of the microphone array in the target device, and the voice can be used as the initial voice; the voice collected by the microphone array in the target device can be stored in the local storage space of the target device, and the voice collected by the microphone array in the target device can be obtained from the local storage space, and the voice can be used as the initial voice.
[0055] The microphone array can be configured as required and is not limited here.
[0056] S12: performing noise reduction processing, frame processing and endpoint detection on the initial speech to obtain the target speech data.
[0057] Specifically, the initial speech is subjected to noise reduction processing, the data obtained by the noise reduction processing is subjected to frame processing, and the data obtained by the frame processing is subjected to endpoint detection.
[0058] Noise reduction aims to remove noise from speech. A typical approach is to use a Wiener filter, which minimizes the mean square error (MSE) based on the statistical characteristics of the signal and noise. Spectral subtraction assumes a relationship between the power spectra of noisy speech and subtracts the estimated noise power spectrum to achieve noise reduction. These methods can improve speech clarity, reduce noise interference on speech recognition, increase recognition accuracy, and enhance speech intelligibility, making them crucial in a variety of speech applications.
[0059] Framing involves segmenting speech into segments based on a specific frame length and frame shift. Frame lengths are typically 20-30 milliseconds, with frame shifts of 10-15 milliseconds. The specific number of frames is determined by the sampling frequency. Windowing is also performed, often using a Hamming window. Framing facilitates the extraction of features such as MFCCs (Mel-Frequency Cepstral Coefficients). Furthermore, due to the short-term stationarity of speech, it helps to better utilize this property, thereby improving the overall performance of tasks such as speech recognition and speaker identification.
[0060] Endpoint detection methods include energy-based and zero-crossing rate-based methods. The energy-based method calculates the energy of a speech frame. Speech energy is lower in silent segments and higher in speech segments. An energy threshold is set. When the frame energy exceeds this threshold, it is considered the beginning of speech. When the frame energy remains below the threshold for a period of time, it is considered the end of speech. The zero-crossing rate-based method refers to the number of times a speech signal crosses the horizontal axis (zero level). The zero-crossing rate is relatively low in silent segments, while higher in speech segments due to more waveform changes. The endpoints of speech are detected by setting an appropriate zero-crossing rate threshold.
[0061] The noise reduction processing in this embodiment improves speech clarity and intelligibility, and reduces speech recognition errors; the frame processing facilitates feature extraction, adapts to speech characteristics and improves system performance; endpoint detection reduces unnecessary processing, saves resources and improves the accuracy of tasks such as speech recognition and synthesis.
[0062] In one embodiment, in the target mode, the steps of extracting features based on the target voice data to obtain target features, and determining the reply voice based on the target features include:
[0063] S221: If the target mode is the offline mode, extracting contextual semantic features based on the target speech data to obtain the target features;
[0064] Specifically, if the target mode is an offline mode, it means that the target device completes the voice interaction independently. Therefore, the target device extracts contextual semantic features based on the target voice data to obtain the target features.
[0065] S222: Perform entity recognition, user intent recognition, context tracking, and generate a first text based on a multi-round dialogue strategy according to the target features to obtain a target text;
[0066] Specifically, through the target device, entity recognition, user intent recognition, context tracking and reply text generation based on the multi-round dialogue strategy are performed according to the target features, and the reply text is used as the first text and the first text is used as the target text.
[0067] S223: Perform speech synthesis according to the target text to obtain the reply speech.
[0068] Specifically, through the target device, a model trained based on TTS (Text To Speech) is used to convert the target text into speech, and the speech is used as the reply speech.
[0069] In this embodiment, the target mode is an offline mode, and voice interaction is completed independently through the target device, thereby ensuring that voice interaction is always available when the network status is poor.
[0070] In one embodiment, the step of extracting features based on the target voice data in the target mode to obtain target features, and determining the reply voice based on the target features, further includes:
[0071] S231: If the target mode is the online mode, extracting Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features based on the target speech data to obtain the target features;
[0072] Specifically, if the target mode is an online mode, it means that the cloud 120 is used as the main mode and the target device 110 is used as the auxiliary mode to complete the voice interaction. The target device 110 extracts the Mel-frequency cepstral coefficient features and voiceprint features based on the target voice data, and extracts the contextual semantic features based on the target voice data. The extracted Mel-frequency cepstral coefficient features, voiceprint features and contextual semantic features are used as the target features.
[0073] First, the speech signal is pre-emphasized, framed, and windowed. A fast Fourier transform (FFT) is then performed to obtain the spectrum. The spectrum is then converted to the Mel scale, and the output energy of the Mel filter bank is calculated. The logarithm of this energy is then taken to obtain the logarithmic energy. Finally, a discrete cosine transform (DCT) is performed, and the DCT coefficients are the Mel-frequency cepstral coefficient features. These coefficients can effectively reflect the characteristics of the speech signal and have a wide range of applications in fields such as speech recognition.
[0074] S232: performing speech index evaluation according to the target feature to obtain index data, wherein the speech index is a speech complexity index and / or a speech confidence index;
[0075] Specifically, a speech index evaluation is performed on the target feature, the speech index and index value obtained by the evaluation are used as associated data, and all associated data are used as index data.
[0076] Regarding speech complexity, acoustic features include signal processing techniques to analyze parameters such as spectrum, energy, fundamental frequency, and formants. Speech rate and pauses are statistically analyzed to assess rhythm, and the signal-to-noise ratio is measured to determine the impact of background noise. Linguistically, vocabulary diversity is measured, and syntactic structure and semantic depth are analyzed. Interactive behavior analysis considers conversational turn length, coherence, error resilience, and the diversity of user intent. Experimental testing, such as benchmark test sets and A / B testing, can also be combined with subjective scoring, taking into account multimodal factors such as emotional intonation and multilingual mixing.
[0077] When assessing speech confidence, we first determine acoustic confidence. Using acoustic models (such as HMMs), we calculate the posterior probability of phonemes or words and compare the probability gaps between competing paths. At the language model level, we calculate the probability of the text within the language model and check vocabulary coverage. When contextualizing speech, we verify the rationality of the results based on conversation history and adapt to specific domain knowledge. For real-time speech, we calculate confidence and set thresholds through streaming processing. We verify confidence through manual annotation comparison, A / B testing, and confusion matrix analysis. For low confidence levels, we employ strategies such as rejection, multiple candidate validation, or active confirmation.
[0078] The calculation methods of speech complexity and speech confidence can be selected from existing technologies and are not limited here.
[0079] For example, speech confidence assessment methods use acoustic models (HMM / DNN) to calculate phoneme posterior probabilities and the optimal-competitive path probability difference. Language models calculate word sequence probabilities and analyze out-of-scope vocabulary coverage. Context verification combines dialogue state tracking with domain knowledge graph matching. Confidence values are dynamically calculated within a streaming architecture, with a multi-level threshold trigger mechanism. Confidence models are optimized through confusion matrix analysis and manually annotated data. Low-confidence results are subject to rejection, multiple candidate re-ranking, and secondary confirmation strategies, achieving end-to-end confidence quantification (F1 score improvement of 23% in typical scenarios).
[0080] Out-of-vocabulary (OOV) refers to words not included in the pre-defined vocabulary of speech recognition or natural language processing systems, such as new words, technical terms, dialects, proper nouns, or rare expressions. Because the system cannot directly match or generate such words, this can lead to recognition errors or semantic bias. Addressing OOV typically relies on sub-word units (such as phonemes), dynamic vocabulary expansion, context inference, or improved language model generalization to mitigate its impact on recognition accuracy.
[0081] For example, the calculation of speech complexity adopts a hierarchical quantification model: through signal processing, acoustic parameters such as spectrum, fundamental frequency, and resonance peak are extracted, the speaking rate / pause interval is statistically evaluated to evaluate the rhythm complexity, and the noise interference is quantified by combining the signal-to-noise ratio; based on NLP, the vocabulary diversity index, syntactic tree depth and semantic embedding distance are calculated; the frequency of dialogue turn switching, intention jump entropy and error correction rate are simultaneously analyzed; finally, the multimodal benchmark test scores (such as ASR error rate) and subjective evaluation weights are integrated to construct a comprehensive complexity coefficient, covering the three-dimensional features of acoustics, language and interaction.
[0082] S233: Based on the preset range data, determine whether to call the cloud according to the indicator data to obtain a first result;
[0083] Optionally, the speech indicator is a speech confidence indicator. If the speech confidence in the indicator data is within the first range of the preset range data, it means that there is no need to call the cloud, and the first result is determined to be no; if the speech confidence in the indicator data is outside the first range of the preset range data, it means that there is a need to call the cloud, and the first result is determined to be yes.
[0084] Optionally, the speech indicator is a speech complexity indicator. If the speech complexity in the indicator data is within the second range of the preset range data, it means that there is no need to call the cloud, and the first result is determined to be no; if the speech complexity in the indicator data is outside the second range of the preset range data, it means that there is a need to call the cloud, and the first result is determined to be yes.
[0085] Optionally, the speech indicators are a speech complexity indicator and a speech confidence indicator. If the speech confidence in the indicator data is within the first range of the preset range data, and the speech complexity in the indicator data is within the second range of the preset range data, it means that the cloud does not need to be called, and the first result is determined to be no. Otherwise, it means that the cloud needs to be called, and the first result is determined to be yes.
[0086] S234: If the first result is yes, then obtaining the reply voice from the cloud based on the target feature, or obtaining a second text from the cloud based on the target feature as the target text, performing speech synthesis based on the target text to obtain the reply voice, or obtaining the reply voice from the cloud based on the target feature and the entity data, or obtaining a second text from the cloud based on the target feature and the entity data as the target text, performing speech synthesis based on the target text to obtain the reply voice, wherein the entity data is data determined by the target device based on the target feature;
[0087] Specifically, if the first result is yes, there are several implementation methods: (1) obtaining the reply voice from the cloud based on the target feature; (2) obtaining a second text from the cloud based on the target feature as the target text, performing speech synthesis based on the target text to obtain the reply voice; (3) obtaining the reply voice from the cloud based on the target feature and the entity data; (4) obtaining a second text from the cloud based on the target feature and the entity data as the target text, performing speech synthesis based on the target text to obtain the reply voice. (1) and (2) extract entity data in the cloud. The cloud can quickly process large-scale data and accurately identify entity data related to the target feature, which helps to obtain more accurate reply voice or target text, reduce the computing burden of the target device, and improve the overall response speed and accuracy. (3) and (4) extract entity data on the target device, which can better adapt to the device environment and perform targeted processing based on device characteristics and locally stored data.
[0088] S235: If the first result is no, entity recognition, user intent recognition, context tracking and a first text are generated based on the target features to obtain a target text, and speech synthesis is performed based on the target text to obtain the reply speech.
[0089] Specifically, if the first result is no, the target device is used to perform entity recognition, user intent recognition, context tracking, and generate a first text based on a multi-round dialogue strategy according to the target features to obtain a target text, and speech synthesis is performed based on the target text to obtain the reply speech.
[0090] In this embodiment, when the target mode is online, multiple feature extractions are first performed on the target speech data to obtain target features, and then speech index evaluation is performed to obtain index data. This helps to accurately determine the complexity and confidence of the speech, so that the decision on whether to call the cloud can be made based on the preset range of data. If the cloud is called, the reply speech or target text can be obtained through various methods and then synthesized into speech, which can take advantage of the rich resources of the cloud. If not, a series of operations are performed locally based on the target features to generate the target text and synthesize speech. This method can flexibly utilize local and cloud resources according to the actual situation of the speech, and provide reply speech efficiently and accurately.
[0091] In one embodiment, the step of fusing the first data and the second data to determine the target text includes:
[0092] S2141: Based on a preset weight configuration, perform weighted summation on the first data and the second data to obtain third data, and determine the target text according to the third data; or
[0093] Based on a preset voting mechanism, the first data and the second data are fused to obtain fourth data, and the target text is determined based on the fourth data.
[0094] Specifically, the weighted summation in step S2341 is a weighted average of the output probability distribution (ie, the first data and the second data). For example, the preset weight configuration is that the first data accounts for 30% and the second data accounts for 70%.
[0095] Based on a preset voting mechanism, the first and second data are fused to generate fourth data. The target text is then determined based on the fourth data. Specifically, the process involves: first, defining voting rules. This determines how each element (e.g., data points, features, semantic units, etc.) in the first and second data participates in the voting. For example, each element can be assigned the same voting weight, or different weights can be assigned based on element type, source, etc. Next, an element-level comparative analysis is performed on the first and second data. For each corresponding element, a vote is performed according to the voting rules. For example, if the two data have the same value for a certain feature, the vote is considered unanimous; if they differ, the predefined rules determine which value has the greater advantage and receives more votes. Next, the voting results are tallied. The voting results for each element are aggregated to generate the fourth data. The value of each element in the fourth data is the final value determined based on the voting results. Finally, the target text is constructed based on the fourth data. The elements in the fourth data are organized according to specific semantic and grammatical rules to form a logically coherent and clearly expressed target text.
[0096] This embodiment uses preset weights to configure a weighted summation, rationally assigning weights based on factors such as the importance and reliability of the data. This allows the fusion result (third data) to more scientifically and accurately reflect the combined value of the two data types, thereby determining a target text that better meets the needs. Furthermore, the use of a preset voting mechanism allows the first and second data to express their "opinions" like different voters during the fusion process. The resulting fourth data combines the "voting results" of both data types. This determines the target text by leveraging the respective strengths of both data types, improving the effectiveness and rationality of data fusion.
[0097] In one embodiment, the method further comprises:
[0098] S31: Acquisition mode confirmation signal;
[0099] The mode confirmation signal is a signal for updating the target mode.
[0100] Specifically, the mode confirmation signal can be input by the user, or it can be triggered by the program implementing the present application according to preset conditions. For example, when the current voice interaction type is successfully changed, the mode confirmation signal is triggered. For another example, when the network fluctuation between the target device and the cloud exceeds the fluctuation threshold, the mode confirmation signal is triggered.
[0101] S32: Acquire the current voice interaction type according to the mode confirmation signal;
[0102] The current voice interaction type is the type of task currently being performed, or the type of task predicted to be performed next. Current voice interaction tasks include multiple types. On the one hand, the voice interaction type is information query type, with tasks including querying the weather, news, and knowledge questions, meeting users' needs for various types of information. On the other hand, the voice interaction type is operation command type, such as controlling smart home devices, with tasks including users turning on lights and adjusting the temperature through voice commands. There are also entertainment-related types, with tasks including playing music and audiobooks. Furthermore, the voice interaction type is auxiliary social interaction type, with tasks including making voice calls and sending voice messages. Various functions can be conveniently implemented through voice interaction, providing users with a more natural and efficient interactive experience.
[0103] Specifically, when the mode confirmation signal is obtained, the current voice interaction type currently being processed is obtained.
[0104] S33: Determine a first mode set according to the current voice interaction type;
[0105] Specifically, according to the current voice interaction type, a table lookup method is used to determine the first mode set.
[0106] The first mode set includes one or more modes of an offline mode, an online mode, and a mixed mode.
[0107] For example, when the task corresponding to the current voice interaction type is a task with high privacy requirements, the first mode set only includes the offline mode.
[0108] S34: If the first mode set only includes the offline mode, determining that the target mode is the offline mode;
[0109] Specifically, if the first mode set only includes the offline mode, which means that voice interaction can only be performed offline through the target device, the target mode is determined to be the offline mode.
[0110] S35: If the first mode set includes the online mode and / or the hybrid mode, acquiring monitoring data, and determining a second mode set based on the monitoring data, wherein the monitoring data includes one or more data selected from the group consisting of network connection status, local resource load data, user preference data, and energy consumption constraint data;
[0111] Specifically, if the first mode set includes an online mode and / or a hybrid mode, it means that the target device may interact with the cloud through voice, and the monitoring data of the target device may be obtained.
[0112] Optionally, based on the monitoring data, a table lookup method is used to determine the second pattern set.
[0113] Optionally, the monitoring data is input into a pre-trained pattern classification model for classification prediction, and vector elements in the predicted vector whose values are greater than a preset probability are used as target elements, and the patterns corresponding to each target element are used as the second pattern set.
[0114] The pattern classification model is a pre-trained multi-classification model. The model structure and model training method of the pattern classification model can be selected from existing technologies.
[0115] The second mode set includes one or more modes of an offline mode, an online mode, and a mixed mode.
[0116] Network connection status: reflects the connection status between the device and the network, including whether it is connected, connection type (such as WiFi, cellular network), connection stability (signal strength, packet loss rate, etc.), network speed (upload and download rates), etc., used to determine the availability and quality of network communication.
[0117] Local resource load data: refers to the usage of various resources (CPU, memory, disk I / O, GPU, etc.) of local devices (such as computers and mobile phones). It uses quantitative indicators to show the degree of resource shortage or idleness to assist in resource management and task scheduling.
[0118] User preference data: This is information about users' preferences and tendencies in various areas. For example, users' operating habits in the app (such as frequently used features) and content preferences (such as favorite music genres and news categories) help provide personalized services and experiences.
[0119] Energy Consumption Constraints: This information specifies the device's maximum energy consumption under different operating modes, or sets energy consumption caps based on factors such as the remaining battery charge. This can be used to optimize device operation to extend battery life or meet energy consumption requirements.
[0120] S36: performing a union calculation on the first pattern set and the second pattern set to obtain a third pattern set;
[0121] Specifically, a union calculation is performed on the first mode set and the second mode set. If the union is not empty, the union is used as the third mode set. If the union is empty, the offline mode is used as the third mode set.
[0122] S37: Perform pattern screening according to the third pattern set and a preset screening strategy to obtain the target pattern.
[0123] Preset filtering strategies are pre-set strategies.
[0124] Specifically, first, the preset screening strategy may be based on priority settings. If the network connection is stable and the energy consumption constraint is low, the online mode may have a high priority; if the local resource load is large, the offline mode priority is increased. For the hybrid mode, a comprehensive trade-off is required. From the third mode set, first check the network connection status. If it is good and there is no energy consumption concern, the online mode is preferred; if the local resources are tight, the offline mode is preferred. If the various factors are relatively balanced, the hybrid mode will be considered. At the same time, combined with user preference data, if the user often chooses offline mode, the offline mode will be prioritized when the basic conditions are met, thereby obtaining the target mode.
[0125] In another embodiment of the present application, steps S35 to S37 can be replaced as follows: if the first pattern set includes an online mode and / or a hybrid mode, monitoring data is obtained, the first pattern set and the monitoring data are spliced together and input into a pre-trained target model for classification prediction, the vector element with the largest value is selected from the predicted vector, and the classification category corresponding to the selected vector element is used as the target pattern.
[0126] The target model is a pre-trained multi-classification model. The model structure and training method of the target model can be selected from existing technologies.
[0127] This embodiment first determines the first mode set by obtaining a mode confirmation signal and obtaining the current voice interaction type based on it, and can preliminarily frame the mode range according to the task requirements. When the first mode set contains only an offline mode, it is directly determined as the target mode, which reflects the efficient processing of simple situations. When the first mode set contains an online mode and / or a mixed mode, a variety of monitoring data are introduced to determine the second mode set, and factors such as network connection status, local resource load, user preferences and energy consumption constraints are comprehensively considered to make the mode selection more in line with the actual situation. The third mode set is obtained by performing a union calculation on the first and second mode sets, and then the target mode is obtained by filtering according to the preset filtering strategy. It can comprehensively and accurately determine the target mode that best suits the current voice interaction type and meets the current multiple condition constraints, thereby improving the stability, efficiency and user satisfaction of the voice interaction experience.
[0128] In one embodiment, the step of performing pattern screening according to a preset screening strategy based on the third pattern set to obtain the target pattern includes:
[0129] S371: Acquire the target mode as the old mode;
[0130] S372: Perform pattern screening according to the third pattern set according to a preset screening strategy to obtain a new pattern;
[0131] Specifically, pattern screening is performed according to the third pattern set according to a preset screening strategy, and the pattern obtained by screening is used as a new pattern.
[0132] S373: Update the target mode according to the new mode, and if the old mode is different from the updated target mode, perform switching processing according to the old mode and the updated target mode based on a seamless switching strategy.
[0133] A seamless transition strategy aims to achieve smooth transitions between different states, systems, or modes. Technically, this reduces transition latency through algorithm optimization and resource preloading. Service-wise, it ensures data transmission consistency and continuous functional availability. For users, the entire transition process is swift and imperceptible, creating a continuous, uninterrupted experience.
[0134] Specifically, if the old mode is different from the updated target mode, it means that mode switching is required.
[0135] Based on a seamless switching strategy, the switch is performed based on the old mode and the updated target mode. Specifically, when switching from offline mode to online mode (offline → online), intermediate results from local processing, such as speech feature vectors, need to be uploaded to the cloud so that task processing can continue in the cloud, ensuring task continuity. When switching from online mode to offline mode (online → offline), some parameters of the cloud model need to be cached locally. In the event of a network outage, a local degraded model can be used to maintain basic functionality. Regarding service hot switching, for speech recognition (ASR), when switching from local Kaldi (an open source speech recognition toolkit containing numerous speech processing algorithms) to the cloud-based Whisper model (Whisper is OpenAI's speech recognition model), the speech buffer must be preserved to prevent data loss. For natural language understanding (NLU), local entity extraction and cloud-based intent classification are used in end-to-end collaboration. To ensure user experience, streaming processing is used to maintain speech continuity during voice conversation switching, for example, using "delayed filling" technology similar to Duplex (an artificial intelligence technology) to achieve seamless switching. At the same time, the screen displays a mode icon to provide a visual reminder. If cloud processing fails, error recovery is performed, and a fallback to local response is provided, such as the prompt "Network is unstable, switched to offline mode." These steps implement a seamless transition between the old mode and the target mode.
[0136] This embodiment uses the old mode as the initial state of the target mode, then selects the new mode from the third mode set according to a preset filtering strategy. This approach allows for flexible updating of the target mode based on actual conditions. If the old and new modes differ, the switch is performed based on a seamless switching strategy. This technical effect ensures that there will be no interruption or lag during mode switching, achieving a smooth transition between modes. Users will not notice any noticeable difference in switching during use, thereby improving the overall user experience and system stability.
[0137] In one embodiment, the step of concatenating the first pattern set and the monitoring data and inputting them into a pre-trained target model for classification prediction, selecting a vector element with the largest value from the predicted vector, and using the classification category corresponding to the selected vector element as the target pattern includes:
[0138] Obtaining the target mode as the old mode;
[0139] The first pattern set and the monitoring data are concatenated and input into the pre-trained target model for classification prediction. The vector element with the largest value is selected from the predicted vector, and the classification category corresponding to the selected vector element is used as the new pattern;
[0140] The target mode is updated according to the new mode, and if the old mode is different from the updated target mode, a switching process is performed according to the old mode and the updated target mode based on a seamless switching strategy.
[0141] In one embodiment, the step of performing switching processing according to the old mode and the updated target mode based on the seamless switching strategy includes:
[0142] S3731: If the old mode is the offline mode and the updated target mode is the online mode, performing data retention processing on the local voice buffer and sending the intermediate result in the target device to the cloud, wherein the cloud is configured to determine a reply voice or a second text according to the intermediate result;
[0143] The local speech buffer stores the raw speech data collected by the voice input device. This includes acoustic characteristics such as frequency and amplitude, and may also include data that has undergone preliminary processing (such as noise reduction and feature extraction preprocessing). This data provides the basis for subsequent operations such as speech recognition and semantic understanding.
[0144] The data storage area of the local voice buffer can be marked and protected to prevent it from being overwritten or cleared when switching modes. Alternatively, a temporary copy storage can be established to copy the voice buffer data to a specific temporary storage area to ensure data integrity.
[0145] Intermediate results are encapsulated in JSON format and transmitted to the cloud via HTTPS. JSON (JavaScript Object Notation) is a lightweight data exchange format. It organizes data in key-value pairs, making it easy for humans to read and write, as well as for machines to parse and generate. It is widely used for network data transmission and configuration files. HTTPS (Hypertext Transfer Protocol Secure) adds the SSL / TLS encryption protocol to HTTP. By encrypting data transmission, it ensures network communication security, such as protecting website users' privacy and verifying website authenticity. It is commonly used in scenarios requiring secure communication, such as finance and e-commerce.
[0146] The data contained in the intermediate results is related to the voice processing process. When the old mode is switched from offline mode to online mode, if the target device performs voice processing operations, the intermediate results may contain some data from the feature extraction process of the target voice data, such as some voice features that have been extracted. It may also contain preliminary results related to semantic understanding after preliminary analysis of the voice, such as some keywords or preliminary constructed semantic frameworks. These intermediate results are sent to the cloud, and the cloud can more efficiently determine the reply voice or second text based on this data, so that the switch from offline mode to online mode can be a smooth transition and make full use of the interim results obtained from the previous offline processing.
[0147] S3732: If the old mode is an online mode and the updated target mode is an offline mode, the target parameters are obtained from the cloud, the parameters of the local model are updated according to the target parameters, and based on the updated local model, the steps of extracting features according to the target voice data, obtaining target features, and determining the reply voice according to the target features are performed.
[0148] Specifically, the target parameters are integrated into the local model according to predetermined rules, which may involve operations such as parameter replacement and weight adjustment to complete the parameter update of the local model.
[0149] In this embodiment, when switching from offline to online, the local voice buffer data is retained and the intermediate results are sent to the cloud for response; when switching from online to offline, parameters are obtained from the cloud to update the local model, ensuring the consistency and effectiveness of voice processing when switching between different modes.
[0150] See also Figure 3 As shown, in one embodiment, a multi-mode collaborative offline-online voice interaction system is provided, the system including: a cloud and a target device, the target device is communicatively connected to the cloud, the cloud is a remote server, and the target device is configured to implement the steps of the above-mentioned multi-mode collaborative offline-online voice interaction method.
[0151] This embodiment ensures that voice interaction is always available by supporting multiple modes—offline, online, and hybrid—working together. When network conditions are poor or there's no connection, offline mode can be enabled for voice interaction, overcoming the drawback of the existing single online mode, which prevents voice interaction when the network is unavailable. Online mode, on the other hand, fully utilizes the powerful computing and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. Hybrid mode further optimizes the interaction experience by combining the advantages of both offline and online modes. While ensuring consistent availability during the acquisition of target voice data and feature extraction to determine the response voice, online resources are utilized to maximize the accuracy and diversity of voice interaction, thereby enhancing the user experience. In hybrid mode, target features are extracted from the target voice data using Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features. This multi-feature extraction approach helps accurately capture multiple aspects of speech characteristics. Entity recognition and intent recognition based on the target features can then accurately understand user needs and determine entity data and primary data. The second data (based on target features or target features and entity data) is obtained from the cloud, and the first and second data are integrated to determine the target text, and finally the reply voice is synthesized. This process can fully utilize the advantages of cloud resources and local processing, which not only ensures the accuracy and richness of the data, but also can adapt to the needs of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thereby providing users with high-quality voice responses.
[0152] In one embodiment, a target device is provided. The target device is communicatively connected to a cloud, which is a remote server. The target device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0153] Obtain target voice data;
[0154] In the target mode, feature extraction is performed based on the target voice data to obtain target features, and a reply voice is determined based on the target features, wherein the target mode is one of an offline mode, an online mode, and a mixed mode.
[0155] This embodiment ensures that voice interaction is always available by supporting multiple modes—offline, online, and hybrid—working together. When network conditions are poor or there's no connection, offline mode can be enabled for voice interaction, overcoming the drawback of the existing single online mode, which prevents voice interaction when the network is unavailable. Online mode, on the other hand, fully utilizes the powerful computing and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. Hybrid mode further optimizes the interaction experience by combining the advantages of both offline and online modes. While ensuring consistent availability during the acquisition of target voice data and feature extraction to determine the response voice, online resources are utilized to maximize the accuracy and diversity of voice interaction, thereby enhancing the user experience. In hybrid mode, target features are extracted from the target voice data using Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features. This multi-feature extraction approach helps accurately capture multiple aspects of speech characteristics. Entity recognition and intent recognition based on the target features can then accurately understand user needs and determine entity data and primary data. The second data (based on target features or target features and entity data) is obtained from the cloud, and the first and second data are integrated to determine the target text, and finally the reply voice is synthesized. This process can fully utilize the advantages of cloud resources and local processing, which not only ensures the accuracy and richness of the data, but also can adapt to the needs of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thereby providing users with high-quality voice responses.
[0156] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0157] Obtain target voice data;
[0158] In the target mode, feature extraction is performed based on the target voice data to obtain target features, and a reply voice is determined based on the target features, wherein the target mode is one of an offline mode, an online mode, and a mixed mode.
[0159] This embodiment ensures that voice interaction is always available by supporting multiple modes—offline, online, and hybrid—working together. When network conditions are poor or there's no connection, offline mode can be enabled for voice interaction, overcoming the drawback of the existing single online mode, which prevents voice interaction when the network is unavailable. Online mode, on the other hand, fully utilizes the powerful computing and data resources in the cloud, improving the accuracy of voice interaction and the diversity of responses. Hybrid mode further optimizes the interaction experience by combining the advantages of both offline and online modes. While ensuring consistent availability during the acquisition of target voice data and feature extraction to determine the response voice, online resources are utilized to maximize the accuracy and diversity of voice interaction, thereby enhancing the user experience. In hybrid mode, target features are extracted from the target voice data using Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features. This multi-feature extraction approach helps accurately capture multiple aspects of speech characteristics. Entity recognition and intent recognition based on the target features can then accurately understand user needs and determine entity data and primary data. The second data (based on target features or target features and entity data) is obtained from the cloud, and the first and second data are integrated to determine the target text, and finally the reply voice is synthesized. This process can fully utilize the advantages of cloud resources and local processing, which not only ensures the accuracy and richness of the data, but also can adapt to the needs of different application scenarios, effectively improving the intelligence, flexibility and accuracy of the voice interaction system, thereby providing users with high-quality voice responses.
[0160] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be referred to the relevant descriptions on the cloud 120 side and the target device 110 side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0161] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0162] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0163] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A multi-modal collaborative offline and online voice interaction method, characterized in that: The method is applied to a target device, the target device is in communication with a cloud, and the cloud is a remote server. The method includes: Obtain target voice data; In a target mode, feature extraction is performed based on the target voice data to obtain target features, and a reply voice is determined based on the target features, wherein the target mode is one of an offline mode, an online mode, and a hybrid mode; The step of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features includes: If the target mode is a mixed mode, extracting Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features based on the target speech data to obtain the target features; Perform entity recognition, user intent recognition, context tracking, and response text generation based on a multi-round dialogue strategy according to the target features to determine entity data and first data; Acquiring second data from the cloud based on the target feature, or acquiring second data from the cloud based on the target feature and the entity data; fusing the first data and the second data to determine a target text; Perform speech synthesis based on the target text to obtain the reply speech; The step of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features, further includes: If the target mode is the online mode, extracting Mel-frequency cepstral coefficient features, voiceprint features, and contextual semantic features based on the target speech data to obtain the target features; Performing speech index evaluation according to the target feature to obtain index data, wherein the speech index is a speech complexity index and / or a speech confidence index; Based on the preset range data, determining whether to call the cloud according to the indicator data to obtain a first result; If the first result is yes, then obtaining the reply voice from the cloud based on the target feature, or obtaining a second text from the cloud based on the target feature as the target text, performing speech synthesis based on the target text to obtain the reply voice, or obtaining the reply voice from the cloud based on the target feature and the entity data, or obtaining a second text from the cloud based on the target feature and the entity data as the target text, performing speech synthesis based on the target text to obtain the reply voice, wherein the entity data is data determined by the target device based on the target feature; If the first result is no, performing entity recognition, user intent recognition, context tracking, and generating a first text based on the target features and a multi-round dialogue strategy to obtain a target text, and performing speech synthesis based on the target text to obtain the reply speech; The method further comprises: Get mode confirmation signal; Acquire a current voice interaction type according to the mode confirmation signal, where the current voice interaction type is a type of a task currently undergoing voice interaction or a type of a task predicted to be performed next; Determining a first mode set according to the current voice interaction type; If the first mode set only includes the offline mode, determining that the target mode is the offline mode; If the first mode set includes the online mode and / or the hybrid mode, acquiring monitoring data, and determining the second mode set based on the monitoring data, wherein the monitoring data includes one or more data selected from the group consisting of network connection status, local resource load data, user preference data, and energy consumption constraint data; performing a union calculation on the first pattern set and the second pattern set to obtain a third pattern set; Performing pattern screening according to the third pattern set and a preset screening strategy to obtain the target pattern; The step of fusing the first data and the second data to determine the target text includes: Based on a preset weight configuration, performing a weighted summation on the first data and the second data to obtain third data, and determining the target text according to the third data; or Based on a preset voting mechanism, the first data and the second data are merged to obtain fourth data, and the target text is determined according to the fourth data; The first data is a reply text or the first data is intermediate data generated by generating a reply text based on a multi-round dialogue strategy.
2. The multi-modal collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of obtaining target voice data includes: Acquire the speech collected by the microphone array in the target device as the initial speech; The initial speech is subjected to noise reduction processing, frame processing and endpoint detection to obtain the target speech data.
3. The multi-modal collaborative online and offline voice interaction method according to claim 1, characterized in that: The steps of extracting features according to the target voice data in the target mode to obtain target features, and determining the reply voice according to the target features include: If the target mode is an offline mode, extracting contextual semantic features based on the target speech data to obtain the target features; Performing entity recognition, user intent recognition, context tracking, and generating a first text based on a multi-round dialogue strategy according to the target features to obtain a target text; Speech synthesis is performed according to the target text to obtain the reply speech.
4. The multi-modal collaborative online and offline voice interaction method according to claim 1, characterized in that: The step of performing pattern screening according to a preset screening strategy based on the third pattern set to obtain the target pattern includes: Obtaining the target mode as the old mode; Performing pattern screening according to the third pattern set and a preset screening strategy to obtain a new pattern; The target mode is updated according to the new mode, and if the old mode is different from the updated target mode, a switching process is performed according to the old mode and the updated target mode based on a seamless switching strategy.
5. The multi-modal collaborative online and offline voice interaction method according to claim 4, characterized in that: The step of performing switching processing according to the old mode and the updated target mode based on the seamless switching strategy includes: If the old mode is an offline mode and the updated target mode is an online mode, performing data retention processing on the local voice buffer and sending the intermediate result in the target device to the cloud, wherein the cloud is configured to determine a reply voice or a second text according to the intermediate result; If the old mode is an online mode and the updated target mode is an offline mode, the target parameters are obtained from the cloud, the parameters of the local model are updated according to the target parameters, and based on the updated local model, the steps of extracting features according to the target voice data, obtaining target features, and determining the reply voice according to the target features are performed.
6. A multi-modal collaborative online and offline voice interaction system, characterized by: The system includes: a cloud and a target device, the target device is communicatively connected to the cloud, the cloud is a remote server, and the target device is configured to implement the steps of the multi-modal collaborative offline and online voice interaction method as described in any one of claims 1 to 5.
7. A target device, characterized in that: The target device is communicatively connected to the cloud, which is a remote server. The target device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the multi-modal collaborative offline and online voice interaction method as described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the multimodal collaborative offline and online voice interaction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Voice conversation method and system
CN111833880A
Processing method, terminal equipment and storage medium
CN113326018A
Voice system based on off-line mode and on-line mode for intelligent furniture
CN113643711A
Dialogue system off-line and on-line fusion application method and system
CN115795017A