End-side voice interaction method and device

By combining compressed sensing and spiking neural networks with a lightweight model, speech emotion recognition and feedback are achieved on edge devices. This solves the problems of high latency, high power consumption and weak semantic understanding in existing technologies, and realizes low power consumption and real-time emotion recognition and speech feedback, which is suitable for resource-constrained edge deployments.

CN120977295APending Publication Date: 2025-11-18E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511244567.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively combine compressed sensing and spiking neural networks, resulting in large latency, high energy consumption, and weak semantic understanding in edge voice emotion recognition, making it difficult to achieve real-time voice emotion interaction on edge devices.

Method used

Compressed sensing technology is used to reconstruct speech signals through subsampling. Emotion-related pulse features are extracted using a spiking neural network and combined with a lightweight classification network for emotion classification. Emotional speech is generated using FastSpeech2-Lite and HiFi-GAN Mini, and cloud-assisted complex emotion recognition is provided.

Benefits of technology

It achieves low-power, real-time emotion recognition and voice feedback, balancing the real-time performance and accuracy of edge devices while protecting user privacy, making it suitable for resource-constrained edge deployments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977295A_ABST
    Figure CN120977295A_ABST
Patent Text Reader

Abstract

The invention relates to an end-side voice interaction method, and belongs to the technical field of voice interaction, and the method comprises the steps: carrying out the sub-sampling reconstruction of a voice signal at a voice collection end through employing the compressed sensing technology on end-side equipment; inputting the reconstructed voice signal into a pulse neural network module to extract emotion-related pulse features; inputting the emotion-related pulse features into a lightweight classification network for classification; the automatic speech recognition model transwrites the speech signals obtained through reconstruction into text content, the text content serves as input of a natural language processing large model, and semantic emotion tags are output through emotion cross modeling after semantic analysis and classification by means of a pre-training language model or access to a large model platform; and a FastSpeech 2-Lite and HiFi-GAN Mini combined method is adopted, and the semantic emotion tag and the text content are converted into voice output with corresponding emotion. According to the invention, off-line and low-power-consumption emotional speech recognition and synthesis are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of voice interaction, and particularly relates to an end-side voice interaction method and device. BACKGROUND

[0002] Traditional voice emotion recognition methods usually use continuous sampling methods to collect voice and use deep learning models for feature extraction and classification. This approach requires high sampling rates and computing resources, making it difficult to directly meet the low-power and real-time needs of edge devices. The theory of compressed sensing shows that when a signal has sparsity in a certain basis domain, it can obtain sufficient information with a much lower sampling rate than the Nyquist theorem requires. For voice signals, there is also sparse representation under a suitable transformation basis, so compressed sensing technology can be used to achieve low-speed sampling and reconstruction, reducing data volume and energy consumption. On the other hand, spiking neural networks process information through event-driven methods, and their calculations are only updated at the time of neuron firing pulses, making them naturally suitable for low-power and asynchronous signal processing. For example, previous studies have shown that spiking neural networks for audio processing significantly reduce energy consumption and have higher power efficiency than traditional methods. However, compressed sensing and spiking neural networks have not been combined in the field of voice emotion recognition for end-side devices.

[0003] Traditional voice emotion interaction methods mostly use cloud-based centralized reasoning, and the end-side only completes data upload and basic preprocessing. Although this architecture has strong computing power, it brings obvious network dependence, data delay, and privacy leakage problems. Especially in user interaction sensitive scenarios (such as emotional companionship and voice assistants), the end-cloud split causes inconsistency between emotion recognition and response, affecting user experience. With the development of edge computing and end-cloud collaboration systems, the system can complete preliminary identification and rapid response on the end-side, and transfer complex semantic analysis and personalized modeling tasks to the cloud, effectively combining local real-time performance with remote computing power. In recent years, natural language processing pre-training models (such as BERT, ChatGLM, etc.) have been widely used in sentiment classification, semantic understanding, and other fields, and their multi-modal capabilities combined with voice input have become increasingly prominent. Integrating automatic speech recognition modules and natural language processing large models into the emotion recognition process can enhance the analysis of voice content semantics, allowing for accurate judgment of complex and multi-level emotional states.

[0004] Meanwhile, in terms of voice synthesis, although models such as FastSpeech2 and HiFi-GAN can generate high-quality voice, they are large and not easy to run on resource-constrained edge devices. Therefore, lightweight models (such as FastSpeech2-Lite and HiFi-GAN Mini) are needed to support end-side deployment.

[0005] In summary, the prior art has not yet provided a complete voice emotion interaction method that integrates low-bandwidth voice perception, pulse event coding, semantic emotion reasoning, and real-time voice feedback, and is suitable for end-side actual deployment. SUMMARY

[0006] In view of the above deficiencies of the prior art, the purpose of the application is to provide an end-side voice interaction method, device, electronic equipment and storage medium, to propose an end-side voice emotion interaction method with novel structure, high energy efficiency and strong deployability, effectively solving the technical pain points of large delay, high energy consumption and weak semantic understanding existing in existing solutions.

[0007] In a first aspect of the application, an end-side voice interaction method is proposed, comprising:

[0008] On the end-side device, the compressed sensing technology is used to perform sub-sampling reconstruction of the voice signal at the voice collection end;

[0009] The reconstructed voice signal is input into the pulse neural network module to extract emotion-related pulse features;

[0010] The emotion-related pulse features are input into a lightweight classification network for classification;

[0011] The automatic speech recognition model transcribes the reconstructed voice signal into text content, which is input into a natural language processing large model, and the pre-trained language model or the large model platform is used for semantic analysis and classification after the emotion cross-modeling outputs semantic emotion labels;

[0012] The joint method of FastSpeech2-Lite and HiFi-GAN Mini is adopted to convert the semantic emotion labels and text content into voice output with corresponding emotions.

[0013] Further, in the above-mentioned end-side voice interaction method, the compressed sensing technology is used to perform sub-sampling reconstruction of the voice signal at the voice collection end, comprising:

[0014] The measurement matrix linearly projects the voice signal collected at the voice collection end to obtain an observation vector: y=Φx

[0015] The voice signal is sparsely represented in a sparse basis: x=Ψs

[0016] The sparse representation of the voice signal in the sparse basis is substituted into the observation vector to obtain an observation model:

[0017] y=Φx=ΦΨs=Θs

[0018] The sparse coefficient vector is recovered by solving the following l1 norm minimization problem using a sparse reconstruction algorithm to obtain the optimal sparse coefficient:

[0019]

[0020] By reconstructing an approximate original speech signal;

[0021] wherein, Φ represents a measurement matrix, x represents a speech signal collected by a speech collection end, y represents an observation vector, s represents a sparse coefficient vector, Ψ represents a transform basis matrix, Θ = ΦΨ represents a comprehensive measurement matrix, represents an optimal sparse coefficient, represents an approximate original speech signal.

[0022] Further, in the above-mentioned one end-side speech interaction method, the reconstructed speech signal is input into a pulse neural network module to extract emotion-related pulse features, comprising:

[0023] The membrane potential of each neuron in the pulse neural network module is updated according to discrete time, and the dynamic model is represented as follows:

[0024]

[0025] wherein: u j [t] represents the membrane potential of the jth neuron at time t, λ represents a voltage decay constant, w ji represents the weight of the i-th input pulse connection of the j-th neuron; x i [t] represents the input pulse of the i-th at time t; z j [t]∈{0,1} represents whether the jth neuron fires a pulse at time t, V th represents the pulse firing threshold, when the membrane potential u j [t+1] exceeds the pulse firing threshold V th , that is, z j [t+1] = 1 neuron will fire a pulse, and reset the membrane potential.

[0026] Further, in the above-mentioned one end-side speech interaction method, the emotion-related pulse features are input into a lightweight classification network for classification, comprising:

[0027] The emotion-related pulse features are input into the lightweight classification network, and the emotion classification output is represented by using a Softmax function as:

[0028]

[0029] wherein, z i represents the score of the i-th emotion, and K represents the total number of categories.

[0030] Further, in the above-mentioned end-side voice interaction method, a joint method of FastSpeech2-Lite and HiFi-GAN Mini is adopted to convert semantic emotion labels and text content into voice output with corresponding emotions, including:

[0031] The pitch predictor and the energy predictor output the pitch and the energy of each frame, respectively;

[0032] The pitch and the energy of each frame are added to the hidden representation of the Transformer as supplementary information, the duration predictor predicts the duration of each phoneme, and adjusts the length of the hidden layer sequence according to the duration of each phoneme;

[0033] During the training process, the pitch, energy and duration prediction are optimized by using mean square error or a custom loss function;

[0034] FastSpeech2-Lite outputs emotion-perceived mel-spectrogram;

[0035] HiFi-GAN Mini converts the emotion-perceived mel-spectrogram into waveform audio as a neural vocoder.

[0036] Further, the above-mentioned end-side voice interaction method further includes:

[0037] When the emotion classification confidence is less than a preset threshold or a complex emotional state is encountered, the intermediate feature or the emotion classification result is sent to the cloud;

[0038] The cloud-side deep model further classifies the emotion;

[0039] The automatic speech recognition model of the cloud side transcribes the voice signal into text content, and the text content is input into the natural language processing large model as an input, and the pre-trained language model or the access to the large model platform is used for semantic analysis and classification of the emotion cross-modeling output semantic emotion labels.

[0040] Further, in the above-mentioned end-side voice interaction method, the loss function during the training process is represented as follows: the intermediate feature or the emotion classification result is sent to the cloud through low-latency protocol communication.

[0041] The second aspect of the present application also proposes an end-side voice interaction device, including:

[0042] The reconstruction module is used for sub-sampling and reconstructing the voice signal at the voice acquisition end on the end-side device by using the compressed sensing technology;

[0043] The extraction module is used for inputting the reconstructed voice signal into the pulse neural network module to extract emotion-related pulse features;

[0044] The classification module is used for inputting the emotion-related pulse features into a lightweight classification network for classification.

[0045] The modeling output module is used for transcribing the reconstructed voice signal into text content through an automatic speech recognition model, taking the text content as an input of a natural language processing large model, and outputting a semantic emotion label through semantic analysis and classification of a pre-trained language model or an access to a large model platform.

[0046] The voice output module is used for converting the semantic emotion label and the text content into voice output with corresponding emotions by using a joint method of FastSpeech2-Lite and HiFi-GAN Mini.

[0047] The third aspect of the present application further provides an electronic device, comprising a processor and a memory.

[0048] The processor is used for executing the end-side voice interaction method according to any one of the above aspects by calling programs or instructions stored in the memory.

[0049] The fourth aspect of the present application further provides a computer readable storage medium storing programs or instructions, which make a computer execute the end-side voice interaction method according to any one of the above aspects.

[0050] The present application has the following beneficial effects:

[0051] 1) The compressed sensing technology is used to perform sub-sampling at the voice collection end, so that most information of the original voice is obtained by using fewer observation values, the sampling rate and the data amount are reduced, and the device power and the storage burden are reduced.

[0052] 2) The event-driven pulse neural network module is used for voice activity detection and feature extraction, and the calculation is triggered only when valid voice is detected, so that the energy consumption is significantly reduced.

[0053] 3) The lightweight emotion classification task is realized at the end side, the complex model inference is performed at the cloud side, the ASR and NLP modules are introduced for semantic understanding and emotion modeling, and the two modules work together to balance the real-time performance, the robustness and the accuracy, and the user privacy is protected.

[0054] 4) The FastSpeech2-Lite is used to generate the mel spectrogram of the emotional voice, and the HiFi-GAN Mini is used to generate the final waveform, so that the high-fidelity emotional voice synthesis is realized, and the model size is suitable for end-side deployment. BRIEF DESCRIPTION OF DRAWINGS

[0055] The accompanying drawings are included to provide a further understanding of the embodiments of the application, and are incorporated in and constitute a part of this specification. The drawings illustrate the embodiments of the application and, together with the description, serve to explain the principles of the application. In the drawings:

[0056] Figure 1 An end-side voice interaction method provided for an embodiment of the application Figure 1 ;

[0057] Figure 2 A method for converting a semantic emotion label and text content into a voice output with a corresponding emotion provided for an embodiment of the application

[0058] Figure 3 An end-side voice interaction method provided for an embodiment of the application Figure 2 ;

[0059] Figure 4 An end-side voice interaction device diagram provided for an embodiment of the application

[0060] Figure 5 A schematic block diagram of an electronic device provided for an embodiment of the application. DETAILED DESCRIPTION

[0061] In order to make the personnel in the art better understand the technical solutions in the embodiments of the application, the technical solutions of the application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are some of the embodiments of the application, rather than all the embodiments of the application. It should be understood that these descriptions are only exemplary and are not used to limit the scope of the application. Based on the embodiments of the application, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the scope of protection of the application.

[0062] In addition, in the following description, the description of well-known structures and techniques is omitted to avoid unnecessary confusion of the concepts disclosed in the application.

[0063] In the description of the application, the terms "first", "second", "third" are only used for the purpose of description, and cannot be understood as indicating or implying relative importance. The terms "mounting", "connecting", "connecting" should be broadly understood, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication inside two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the application can be understood according to the specific circumstances.

[0064] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, like reference numerals refer to like elements throughout the description. The following exemplary embodiments are not representative of all embodiments consistent with the present invention. Rather, they are merely examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.

[0065] The present application provides an end-side voice interaction method and device, and proposes an end-side voice emotion interaction method with novel structure, high energy efficiency and strong deployability, which effectively solves the technical problems of large delay, high energy consumption and weak semantic understanding existing in the prior art.

[0066] Before introducing the embodiments of the present application, first introduce the professional terms related to the present application.

[0067] Compressed sensing (CS): a signal sampling technology that utilizes the sparsity of a signal in a certain transform domain, and can reconstruct the signal using far fewer sampling points than required by the Nyquist sampling theorem.

[0068] Spiking neural network (SNN): the third generation of artificial neural networks, which simulates the spiking and dynamic timing characteristics of neurons, and is used for efficient event-driven computing.

[0069] Method embodiment

[0070] Figure 1 An end-side voice interaction method provided by the embodiments of the present application Figure 1 .

[0071] In a first aspect of the present application, an end-side voice interaction method is provided, which combines Figure 1 , and includes five steps S1 to S5:

[0072] S1: On the end-side device, use the compressed sensing technology to perform sub-sampling reconstruction of the voice signal at the voice collection end.

[0073] Specifically, in the embodiments of the present application, the method of using the compressed sensing technology to perform sub-sampling reconstruction of the voice signal at the voice collection end on the end-side device is described in detail below.

[0074] S2: Input the reconstructed voice signal into the spiking neural network module to extract emotion-related pulse features.

[0075] Specifically, in the embodiments of the present application, the method of inputting the reconstructed voice signal into the spiking neural network module to extract emotion-related pulse features is described in detail below.

[0076] S3: inputting the emotion-related pulse feature into a lightweight classification network for classification.

[0077] Specifically, in the embodiments of the present application, the method of inputting the emotion-related pulse feature into a lightweight classification network for classification is described in detail below.

[0078] S4: The automatic speech recognition model converts the reconstructed speech signal into text content, and the text content is input into a natural language processing large model. The pre-trained language model or the large model platform is used for semantic analysis and classification, and the cross-model output semantic emotion label.

[0079] Specifically, in the embodiments of the present application, the automatic speech recognition model such as ASR is introduced for end-side deployment. The speech signal is converted into text content in real time, and the text content is input into a natural language processing large model. The natural language processing large model can be NLP. The pre-trained language model or the large model platform is used for semantic analysis and emotion cross-modeling. The pre-trained language model can be BERT, ChatGLM, etc. After modeling, more accurate semantic emotion labels, user intentions, etc. can be output.

[0080] S5: A joint method of FastSpeech2-Lite and HiFi-GAN Mini is used to convert the semantic emotion label and the text content into speech output with corresponding emotions.

[0081] Specifically, in the embodiments of the present application, the method of using the joint method of FastSpeech2-Lite and HiFi-GAN Mini to convert the semantic emotion label and the text content into speech output with corresponding emotions is described in detail below.

[0082] Further, in the above-mentioned one end-side voice interaction method, the compressed sensing technology is used to perform sub-sampling and reconstruct the speech signal at the voice collection end, comprising:

[0083] The measurement matrix linearly projects the speech signal collected at the voice collection end to obtain an observation vector: y = Φx

[0084] The speech signal is sparsely represented in a sparse basis: x = Ψs

[0085] The sparse representation of the speech signal in the sparse basis is substituted into the observation vector to obtain an observation model:

[0086] y = Φx = ΦΨs = Θs

[0087] The sparse coefficient vector is recovered by solving the following l1 norm minimization problem using a sparse reconstruction algorithm to obtain the optimal sparse coefficient:

[0088]

[0089] By reconstructing an approximate original speech signal;

[0090] wherein, Φ represents a measurement matrix, x represents a speech signal collected by a speech collection end, y represents an observation vector, s represents a sparse coefficient vector, Ψ represents a transform basis matrix, Θ = ΦΨ represents a comprehensive measurement matrix, represents an optimal sparse coefficient, represents an approximate original speech signal.

[0091] Specifically, in the embodiment of the present application, using the compressed sensing technology to reconstruct the speech signal at the speech collection end can significantly reduce the requirement for the sampling rate, thereby reducing the data volume and energy consumption, and is particularly suitable for resource-limited end-side devices. A random Gaussian matrix or a structured sparse matrix can be selected as the measurement matrix Φ, and a sparse reconstruction algorithm (such as BP, OMP, etc.) can be combined to realize fast reconstruction.

[0092] Further, in the above-mentioned end-side speech interaction method, the reconstructed speech signal is input into a spiking neural network module to extract emotion-related spiking features, comprising:

[0093] The membrane potential of each neuron in the spiking neural network module is updated according to discrete time, and the dynamic model is represented as follows:

[0094]

[0095] wherein, u j [t] represents the membrane potential of the jth neuron at time t, λ represents a voltage decay constant, w ji represents the weight of the i-th input pulse connection of the j-th neuron; x i [t] represents the input pulse of the i-th at time t; z j [t] ∈ {0, 1} represents whether the j-th neuron fires a pulse at time t, V th represents a pulse firing threshold, when the membrane potential u j [t+1] exceeds the pulse firing threshold V th , that is, z j [t+1] = 1 neuron will fire a pulse, and the membrane potential will be reset.

[0096] Specifically, in the embodiment of the present application, the pulse neural network module effectively simulates the discharge behavior of biological neurons, so that the network only produces energy consumption when an event occurs. Compared with traditional artificial neural networks, the power consumption can be significantly reduced when processing sparse events such as speech activity. The pulse neural network module not only can detect the start and end positions of speech online, but also can extract timing features synchronously; through statistical feature analysis of the output pulse sequence, such as pulse emission rate and pulse time coding, expression features reflecting high-level information such as speech emotion can be further extracted.

[0097] Further, in the above-mentioned end-side voice interaction method, the emotion-related pulse feature is input into a lightweight classification network for classification, comprising:

[0098] The emotion-related pulse feature is input into a lightweight classification network, and a Softmax function is used for emotion classification output, represented as:

[0099]

[0100] Wherein, z i represents the score of the i-th emotion, and K represents the total number of categories.

[0101] Specifically, in the embodiment of the present application, the lightweight classification network completes preliminary emotion recognition on the end side to meet the low delay response requirement.

[0102] Figure 2 A method for converting semantic emotion labels and text content into voice output with corresponding emotions is provided.

[0103] Further, in the above-mentioned end-side voice interaction method, a joint method of FastSpeech2-Lite and HiFi-GANMini is adopted to convert semantic emotion labels and text content into voice output with corresponding emotions. The conversion of semantic emotion labels and text content into voice output with corresponding emotions includes five steps of S21 to S25:

[0104] S21: The pitch predictor and the energy predictor output the pitch and the energy of each frame, respectively;

[0105] S22: The pitch and the energy of each frame are added to the hidden representation of the Transformer as supplementary information, the duration predictor predicts the duration of each phoneme, and the hidden layer sequence length is adjusted according to the duration of each phoneme;

[0106] S23: In the training process, the pitch, energy and duration prediction are optimized by using mean square error or a custom loss function;

[0107] S24: FastSpeech2-Lite outputs emotion-aware mel-spectrogram;

[0108] S25: HiFi-GAN Mini converts emotion-aware mel-spectrogram into high-quality waveform audio as a neural vocoder.

[0109] Specifically, in the embodiment of the present application, the FastSpeech2-Lite internal pitch predictor and energy predictor output the pitch and energy

[0110] wherein,

[0111] F 0,t represents the fundamental frequency of the frame of speech, i.e. the pitch period frequency of vocal cord vibration, and ε represents a small constant to avoid numerical instability, and s t [n] represents the speech sampling point of the t-th frame, and N represents the number of sampling points per frame.

[0112] The pitch and energy of each frame are added to the hidden representation of the Transformer as supplementary information to enhance the prosody modeling capability of the speech, and the duration predictor predicts the duration of each phoneme and adjusts the length of the hidden layer sequence according to the duration of each phoneme so that the output frame number meets the speed requirement, and in the training process, the pitch, energy and duration prediction are optimized using mean square error or a custom loss function; FastSpeech2-Lite outputs emotion-aware mel-spectrogram; HiFi-GAN Mini converts emotion-aware mel-spectrogram into high-quality waveform audio as a neural vocoder.

[0113] In the training process, the training target includes the following loss functions:

[0114] Adversarial loss:

[0115] Mel-spectrogram reconstruction loss:

[0116] Discriminator feature matching loss:

[0117] The final generator total loss function is: L G = L adv + λ fm L fm + λ mel L mel

[0118] wherein the weight hyperparameter takes the value: λfm = 2, l mel = 45.

[0119] In some embodiments, the above-mentioned compressed sensing technology, pulse neural network module, etc. can be deployed to an edge AI chip platform such as RK3588 by means of RKNN Toolkit 2. RK3588 adopts an 8nm process, an eight-core CPU, and a built-in 6TOPS NPU, has high computing power and low power consumption characteristics, and is suitable for offline inference of the present method. FastSpeech2-Lite and HiFi-GAN Mini models are converted into RKNN format by RKNN tool, and loaded into RK3588 to realize end-side emotional synthesis, thereby meeting the real-time and energy-saving requirements of intelligent terminals.

[0120] It should be understood that the model is deployed on a low-power AI chip such as RK3588 by means of RKNN Toolkit, etc. to realize offline running on the end side to meet the engineering landing requirements of intelligent terminal devices.

[0121] Figure 3 An end-side voice interaction method provided for the embodiments of the present application Figure 2 .

[0122] Further, the above-mentioned end-side voice interaction method, in combination with Figure 3 further comprises three steps S31 to S33:

[0123] S31: When the emotion classification confidence is less than a preset threshold or a complex emotional state is encountered, the intermediate feature or emotion classification result is sent to the cloud;

[0124] S32: The deep model of the cloud further classifies the emotion;

[0125] S33: The automatic speech recognition model of the cloud transcribes the speech signal into text content, the text content is input into the natural language processing large model, and the pre-trained language model or the access to the large model platform is used for semantic analysis and classification of the emotion cross-modeling output semantic emotion label.

[0126] Specifically, in the embodiment of the present application, when the emotion classification confidence is less than the preset threshold or a complex emotion state is encountered, the intermediate feature or emotion classification result is uploaded to the cloud, and a more complex deep model of the cloud, such as BiLSTM or Transformer, is further analyzed to improve the overall recognition accuracy. Here, an automatic speech recognition model deployed in the cloud, such as ASR, is introduced to convert the speech signal into text content in real time. The text content is used as the input of the natural language processing large model. The natural language processing large model can be NLP, which uses a pre-trained language model or accesses a large model platform for semantic analysis and emotion cross-modeling. The pre-trained language model can be BERT or ChatGLM. After modeling, more accurate semantic emotion labels, user intentions and other information can be output.

[0127] Further, in the above-mentioned end-side voice interaction method, the intermediate feature or emotion classification result is sent to the cloud through low-latency protocol communication.

[0128] Specifically, in the embodiment of the present application, low-latency protocol communication can be used between the end side and the cloud to realize the cooperation of fast edge response and high-precision cloud decision. The low-latency protocol can be MQTT or gRPC.

[0129] Device embodiment

[0130] Figure 4 An end-side voice interaction device provided by the embodiment of the present application is shown in the figure.

[0131] The second aspect of the present application also proposes an end-side voice interaction device, which combines Figure 4 , comprising:

[0132] The reconstruction module 41 is configured to reconstruct the speech signal at the speech acquisition end using the compressed sensing technology on the end-side device.

[0133] Specifically, in the embodiment of the present application, the reconstruction module 41 reconstructs the speech signal at the speech acquisition end using the compressed sensing technology on the end-side device.

[0134] The extraction module 42 is configured to input the reconstructed speech signal into the pulse neural network module to extract emotion-related pulse features.

[0135] Specifically, in the embodiment of the present application, the extraction module 42 inputs the reconstructed speech signal into the pulse neural network module to extract emotion-related pulse features.

[0136] The classification module 43 is configured to input the emotion-related pulse features into a lightweight classification network for classification.

[0137] Specifically, in the embodiment of the present application, the classification module 43 inputs the emotion-related pulse features into a lightweight classification network for classification.

[0138] The modeling output module 44 is configured to transcribe the reconstructed speech signal into text content by an automatic speech recognition model, and the text content is input into a natural language processing large model, and a pre-trained language model or a large model platform is used for semantic analysis and emotion cross-modeling to output a semantic emotion label.

[0139] Specifically, in the embodiment of the present application, an automatic speech recognition model such as ASR is introduced on the edge side to transcribe the speech signal into text content in real time, and the text content is input into a natural language processing large model, and the natural language processing large model can be NLP, and a pre-trained language model or a large model platform is used for semantic analysis and emotion cross-modeling, and the pre-trained language model can be BERT, ChatGLM, etc., and the modeling output module 44 can output more accurate semantic emotion labels, user intentions, etc.

[0140] The speech output module 45 is configured to convert the semantic emotion label and the text content into speech output with corresponding emotions by using a joint method of FastSpeech2-Lite and HiFi-GAN Mini.

[0141] Specifically, in the embodiment of the present application, the speech output module 45 converts the semantic emotion label and the text content into speech output with corresponding emotions by using a joint method of FastSpeech2-Lite and HiFi-GAN Mini.

[0142] The third aspect of the present application also provides an electronic device, comprising: a processor and a memory;

[0143] The processor is configured to execute the edge-side voice interaction method according to any one of the above embodiments by calling the program or instructions stored in the memory.

[0144] The fourth aspect of the present application also provides a computer readable storage medium, which stores programs or instructions, and the programs or instructions make the computer execute the edge-side voice interaction method according to any one of the above embodiments.

[0145] Figure 5 is a schematic block diagram of an electronic device provided by the embodiment of the present application.

[0146] As Figure 5As shown, the electronic device includes at least one processor 501, at least one memory 502, and at least one communication interface 503. The various components in the electronic device are coupled together by a bus system 504. The communication interface 503 is configured to perform information transmission with external devices. It can be understood that the bus system 505 is used to realize the connection communication between the components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all the buses are marked as the bus system 504 in the Figure 5

[0147] It can be understood that the memory 502 in the embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0148] In some embodiments, the memory 502 stores the following elements, executable units or data structures, or a subset of them, or an extended set of them: an operating system and an application program.

[0149] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program includes various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program for implementing any of the methods of the method for terminal-side voice interaction provided by the embodiment of the application can be included in the application program.

[0150] In the embodiment of the application, the processor 501 processes the steps of each embodiment of the method for terminal-side voice interaction provided by the embodiment of the application by calling the program or instruction stored in the memory 502, specifically, the program or instruction stored in the application program.

[0151] At the terminal device, the compressed sensing technology is used to perform sub-sampling and reconstruct the voice signal at the voice collection end;

[0152] The reconstructed voice signal is input into a pulse neural network module to extract emotion-related pulse features;

[0153] The emotion-related pulse features are input into a lightweight classification network for classification;

[0154] The automatic speech recognition model transcribes the reconstructed voice signal into text content, and the text content is input into a natural language processing large model. The pre-trained language model or the large model platform is used for semantic analysis and classification, and the cross-model output of the semantic emotion label is output;

[0155] ​The joint method of FastSpeech2-Lite and HiFi-GAN Mini is used to convert the semantic emotion label and the text content into voice output with corresponding emotions.

[0156] Any of the end-side voice interaction methods provided by the embodiments of the present application can be applied in the processor 501 or implemented by the processor 501. The processor 501 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 501 or the instruction in the form of software. The above processor 501 can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a ready programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general processor can be a microprocessor or the processor can also be any conventional processor.

[0157] The steps of any of the end-side voice interaction methods provided by the embodiments of the present application can be directly embodied as hardware decoding processor execution completion or combined execution completion by hardware and software units in the decoding processor. The software unit can be located in a random memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register and other mature storage media in the field. The storage medium is located in the memory 502, and the processor 501 reads the information in the memory 502, and combines the hardware to complete the steps of the method.

[0158] Those skilled in the art can understand that although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means within the scope of the present application and forms different embodiments.

[0159] Those skilled in the art can understand that the description of each embodiment is focused on, and the part not described in detail in a certain embodiment can refer to the related description of other embodiments.

[0160] Although the embodiments of the present application have been described with reference to the accompanying drawings, various modifications and changes can be suggested to one skilled in the art, and it is intended that the present application encompass such modifications and changes as fall within the scope of the appended claims. The embodiments of the present application described above are combinations of elements and features of the present application. The elements and features can be used alone or in any combination(s). Each of the elements and features can be used alone or in any combination(s). Thus, the foregoing embodiments of the present application are broad interpretations of the present application and are intended to be examples of embodiments of the present application. Other embodiments of the present application will be obvi ous to those of ordinary skill in the art.

[0161] The embodiments of the present application described above are combinations of elements and features of the present application. The elements and features can be used alone or in any combination(s). Each of the elements and features can be used alone or in any combination(s). Thus, the foregoing embodiments of the present application are broad interpretations of the present application and are intended to be examples of embodiments of the present application. Other embodiments of the present application will be obvi ous to those of ordinary skill in the art.

Claims

1. A terminal-side voice interaction method, characterized in that, include: On the edge device, compressed sensing technology is used to perform subsampling and reconstruct the speech signal at the speech acquisition end; The reconstructed speech signal is input into a spiking neural network module to extract emotion-related pulse features; The emotion-related pulse features are input into a lightweight classification network for classification. The automatic speech recognition model transcribes the reconstructed speech signal into text content. The text content serves as the input to a large natural language processing model. After semantic parsing and classification using a pre-trained language model or by connecting to a large model platform, the model outputs semantic emotion labels through cross-modeling. We employ a combined approach of FastSpeech2-Lite and HiFi-GAN Mini to transform semantic emotion tags and text content into speech output with corresponding emotions.

2. The terminal-side voice interaction method according to claim 1, characterized in that, Using compressed sensing technology to perform subsampling and reconstruct speech signals at the speech acquisition end, including: The measurement matrix is ​​linearly projected onto the speech signal acquired by the speech acquisition terminal to obtain the observation vector: y = Φx The sparse representation of a speech signal under a sparse basis is: x = Ψs Substituting the sparse representation of the speech signal under the sparse basis into the observation vector yields the observation model: y = Φx = ΦΨs = Θs The optimal sparse coefficients are obtained by reconstructing the sparse coefficient vector through solving the following l1 norm minimization problem using a sparse reconstruction algorithm: pass Reconstructing an approximate original speech signal; Where Φ represents the measurement matrix, x represents the speech signal acquired by the speech acquisition terminal, y represents the observation vector, s represents the sparse coefficient vector, Ψ represents the transformation basis matrix, and Θ=ΦΨ represents the comprehensive measurement matrix. Represents the optimal sparsity coefficients. This represents an approximate original speech signal.

3. The terminal-side voice interaction method according to claim 1, characterized in that, The reconstructed speech signal is input into a spiking neural network module to extract emotion-related pulse features, including: In the spiking neural network module, the membrane potential of each neuron is updated according to discrete time, and the dynamic model is represented as follows: Where: u j [t] represents the membrane potential of the j-th neuron at time t, λ represents the voltage decay constant, and w ji x represents the weight of the i-th input pulse connection of the j-th neuron; i [t] represents the input pulse of the i-th channel at time t; z j [t]∈{0,1} indicates whether the j-th neuron fires a pulse at time t, V th This represents the pulse emission threshold, when the membrane potential u j [t+1] Exceeds the pulse emission threshold V th At that time, i.e., z j The [t+1]=1 neuron will fire a pulse and reset the membrane potential.

4. The terminal-side voice interaction method according to claim 1, characterized in that, The emotion-related impulse features are input into a lightweight classification network for classification, including: The emotion-related impulse features are input into a lightweight classification network, and the emotion classification output using the Softmax function is represented as follows: Among them, z i Let K represent the score of the i-th emotion category, and K represent the total number of categories.

5. The terminal-side voice interaction method according to claim 1, characterized in that, A joint approach using FastSpeech2-Lite and HiFi-GAN Mini is employed to transform semantic emotion tags and text content into speech output with corresponding emotions, including: The pitch predictor and energy predictor output the pitch and energy of each frame, respectively. The pitch and energy of each frame are added as supplementary information to the hidden representation of the Transformer. The duration predictor predicts the duration of each phoneme and adjusts the hidden layer sequence length according to the duration of each phoneme. During training, pitch, energy, and duration predictions were all optimized using mean squared error or a custom loss function. FastSpeech2-Lite outputs a Mel-spectrum graph of emotion perception; HiFi-GAN Mini, as a neural vocoder, converts the Mel spectrogram of emotion perception into waveform audio.

6. The terminal-side voice interaction method according to claim 4, characterized in that, The method further includes: When the confidence level of the emotion classification is less than the preset threshold or when a complex emotional state is encountered, the intermediate features or emotion classification results will be sent to the cloud. Deep models in the cloud further classify emotions; The cloud-based automatic speech recognition model transcribes speech signals into text content. This text content serves as input to a large natural language processing model. The model then uses a pre-trained language model or connects to a large model platform to perform semantic parsing and classification, followed by cross-modeling of emotions to output semantic emotion labels.

7. The terminal-side voice interaction method according to claim 6, characterized in that, The loss function during training is represented as follows: intermediate features or sentiment classification results are sent to the cloud via low-latency protocol communication.

8. A terminal-side voice interaction device, characterized in that, include: Reconstruction module: Used on the edge device to perform subsampling and reconstruct speech signals at the speech acquisition end using compressed sensing technology; Extraction module: Used to input the reconstructed speech signal into the spiking neural network module to extract emotion-related pulse features; Classification module: used to input the emotion-related pulse features into a lightweight classification network for classification; Modeling output module: It is used to transcribe the reconstructed speech signal into text content through an automatic speech recognition model. The text content serves as the input of a large natural language processing model. After semantic parsing and classification using a pre-trained language model or by connecting to a large model platform, it outputs semantic emotion labels through cross-modeling. Voice output module: Used to convert semantic emotion tags and text content into voice output with corresponding emotions using a joint approach of FastSpeech2-Lite and HiFi-GAN Mini.

9. An electronic device, characterized in that, include: Processor and memory; The processor executes an end-side voice interaction method as described in any one of claims 1 to 7 by calling programs or instructions stored in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform an end-side voice interaction method as described in any one of claims 1 to 7.