Real-time voice processing method based on embedded development board

By building a real-time voice processing system on the ESP32 development board, using audio input and output devices and camera hardware resources, combined with innovative voice algorithm architecture, the shortcomings of traditional voice generation solutions in resource-constrained scenarios are solved, and low-power consumption, high adaptability, and personalized real-time voice generation is achieved, which improves the voice interaction experience.

CN120183378APending Publication Date: 2025-06-20LINKER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510311936.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional voice generation solutions are difficult to achieve immediate and personalized voice requirements in resource-constrained scenarios such as small smart home devices and portable smart wearables, and are limited by hardware performance and algorithm complexity.

Method used

The real-time voice processing method based on the ESP32 development board is adopted, and the audio input and output devices and camera hardware resources provided by the development board are fully utilized, combined with the innovative voice algorithm architecture, real-time voice generation with low power consumption, high adaptability, flexible customization.

Benefits of technology

It realizes real-time voice generation with low power consumption, high adaptability, flexible customization, and meets the immediate and personalized voice needs in scenarios such as small smart home devices and portable smart wearables, and provides a smoother and smart voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005314802400000037
    Figure BDA0005314802400000037
  • Figure BDA0005314802400000061
    Figure BDA0005314802400000061
  • Figure FDA0005314802390000013
    Figure FDA0005314802390000013
Patent Text Reader

Abstract

The invention discloses a real-time voice processing method based on an embedded development board, and the method comprises the following steps: 1, building a voice interaction system based on the embedded development board, and enabling the voice interaction system to be used for storing audio clips and capturing sound signals in real time; and 2, analyzing sound signals captured in real time, acquiring text information acquired from an external input source, analyzing the text information, screening dynamic samples, performing audio fusion generation based on an attention mechanism, completing voice processing, and finally performing model optimization and adaptive adjustment. According to the real-time voice processing method based on the embedded development board, the hardware resource potential of audio input and output equipment, a camera and the like of the ESP32 development board is fully excavated, the conventional limitation is broken through in combination with an innovative voice algorithm architecture, and a real-time voice generation function which is low in power consumption, high in adaptability and capable of being flexibly customized is realized; and smoother and more intelligent voice interaction experience is brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent voice interaction, and more specifically to a real-time voice processing method based on an embedded development board. Background Art

[0002] In today's era of rapid development of intelligence, various smart terminal devices have an increasingly urgent need for real-time, accurate and diversified voice interaction. However, traditional voice generation solutions are often limited by hardware performance, algorithm complexity or lack of adaptability to complex scenarios, and are difficult to meet the instant and personalized voice needs in resource-constrained scenarios such as small smart home devices and portable smart wearables.

[0003] For example, in the prior art, there is an invention patent with patent number 202310656368.2, entitled User Intent Recognition Method Based on Multimodal Fusion in Complex Human-Computer Interaction Scenarios, which discloses the use of multimodal fusion to realize user intent recognition for complex human-computer interactions. However, this method requires strong hardware support in the recognition process. Therefore, it is limited by hardware performance, algorithm complexity or lack of adaptability to complex scenarios, and it is difficult to meet the instant and personalized voice needs in resource-constrained scenarios such as small smart home devices and portable smart wearables. Summary of the invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a real-time speech generation function that can fully tap the potential of the hardware resources such as audio input and output devices and cameras of the ESP32 development board, combine with an innovative voice algorithm architecture, break conventional limitations, and achieve low power consumption, high adaptability, and flexible customization, so as to bring users a smoother and more intelligent voice interaction experience.

[0005] To achieve the above object, the present invention provides the following technical solution: a real-time speech processing method based on an embedded development board, characterized in that it comprises the following steps:

[0006] Step 1: Build a voice interaction system based on an embedded development board to store audio clips and capture sound signals in real time;

[0007] Step 2: Analyze the sound signal captured in real time, collect and analyze the text information obtained from the external input source, filter the dynamic samples, and then generate audio fusion based on the attention mechanism to complete the speech processing, and finally optimize the model and perform adaptive adjustment.

[0008] As a further improvement of the present invention, the specific steps of building a voice interaction system based on an embedded development board in step 1 are as follows:

[0009] Step 1: Build a voice interaction system with the ESP32 development board as the core. Its built-in audio input module captures sound signals in real time and converts them into digital audio streams after preprocessing. stream ; The camera synchronously collects visual image frames I frame , through the image recognition model M img Extract scene feature vector Used for subsequent voice scene adaptation;

[0010] Step 1 and 2: Build a local micro audio sample library L audio , stores audio clips of different types and styles, and each sample is associated with multi-dimensional metadata, including timbre tags T lobel 、Emotional tendency tag , semantic scene adaptation C score , convenient for quick retrieval and call.

[0011] As a further improvement of the present invention, the specific method of parsing the sound signal captured in real time in step 2 and collecting and parsing the text information obtained from the external input source is: obtaining the text information T from the external input source input , through the natural language processing module NLP module Perform word segmentation, part-of-speech tagging, and semantic understanding to generate structured text representation At the same time, combined with audio input stream A stream Analyze the current environmental noise level N level , using the camera image frame to determine the scene activity A activity ,The above environmental factors serve as regulatory parameters for subsequent speech generation.

[0012] As a further improvement of the present invention, the specific method of screening dynamic samples in step 2 is: first, according to the structured text and environmental factors, generate the initial sample screening conditions, and then continuously monitor the user feedback index F of the generated speech during the voice interaction process feedback , dynamically adjust the sample screening weight according to the feedback. If it is found that a certain type of sample causes more negative feedback, reduce its subsequent screening priority to achieve adaptive optimization. Finally, filter the dynamic samples through the optimized screening conditions.

[0013] As a further improvement of the present invention, the specific method of generating the initial sample screening condition in step 2 is: assuming that the key semantic feature vector in the text is Calculate the audio samples in the sample library and The semantic relevance R semantic (i) A method based on word vector cosine similarity combined with semantic weight allocation

[0014] Among them, wj is the semantic weight of keyword j is the semantic feature vector of keyword j corresponding to sample i. Combining the timbre, emotion label of the sample and the adaptation degree of the current environment, comprehensively sort and filter to obtain the candidate audio sample set C samples .

[0015] As a further improvement of the present invention, the specific steps of audio fusion generation based on the attention mechanism in the second step are as follows:

[0016] Step 2-1, introduce the multi-head attention mechanism M multihead to process the candidate audio sample set C samples ;

[0017] Step 2-2, fuse the weighted sample features to generate a fused feature vector and input it into the generator G lite in the lightweight generative adversarial network CAN net , combined with the text information to generate the final speech signal V output , expressed as:

[0018]

[0019] Step 2-3, the discriminator D of the generative adversarial network net is used to discriminate the authenticity of the generated speech, and continuously improve the quality of the generated speech through adversarial training to make it close to natural speech.

[0020] As a further improvement of the present invention, the specific steps of introducing the multi-head attention mechanism M multihead to process the candidate audio sample set C samples are as follows: taking the feature vector of each sample as the input, and calculating the attention weight W att (i) of each sample in the current speech generation task through the attention layer, so that the model can focus on the audio features that best match the text and the scene

[0021] As a further improvement of the present invention, the specific method of model optimization and adaptive adjustment in the second step is: regularly collect a small amount of new speech samples and corresponding interaction feedback data, and use the online learning algorithm O learn to fine-tune the model parameters. For example, when updating the parameters θ net of the generator G G , use the mini-batch stochastic gradient descent variant algorithm

[0022]

[0023] Among them, η is the learning rate, Loss GAN To generate adversarial loss, Loss adapt is the adaptive adjustment loss, and λ is the balance coefficient.

[0024] The beneficial effect of the present invention is that the speech processing method of the present embodiment is adopted to fully tap the potential of the hardware resources such as the audio input and output devices and the camera of the ESP32 development board, and the innovative speech algorithm architecture is combined to break the conventional limitations, realize the real-time speech generation function with low power consumption, high adaptability and flexible customization, and bring users a smoother and more intelligent speech interaction experience. The algorithm is optimized for the resource characteristics of the ESP32 development board: the storage capacity and computing power of the ESP32 chip are fully considered, and a lightweight and efficient speech processing process is designed to reduce unnecessary data redundancy and complex calculations, and ensure stable operation in an embedded environment. At the same time, multimodal fusion perception enhances speech generation: innovatively integrates audio input and camera visual information, uses visual scenes to assist in understanding user intentions and environmental atmosphere, and dynamically adjusts the speech generation strategy to make the speech output more in line with the actual situation. Finally, an adaptive dynamic sample matching algorithm is used: abandoning the traditional static audio sample retrieval method, dynamically updating the sample selection weight according to real-time speech interaction feedback, optimizing the selection of synthetic audio samples, and improving the accuracy and flexibility of speech generation. DETAILED DESCRIPTION

[0025] The present invention will be further described in detail with reference to the given embodiments below.

[0026] A real-time speech processing method based on an embedded development board in this embodiment is mainly applied to the ESP32 development board. The speech processing method constructed for the development board fully considers the storage capacity and computing power of the ESP32 chip, designs a lightweight and efficient speech processing flow, reduces unnecessary data redundancy and complex calculations, and ensures stable operation in an embedded environment; therefore, the specific steps of the method are as follows:

[0027] First, build the system architecture based on the ESP32 development board, which mainly consists of the following two steps:

[0028] (1) The voice interaction system is built with the ESP32 development board as the core. Its built-in audio input module captures sound signals in real time and converts them into digital audio streams after preprocessing. stream ; The camera synchronously collects visual image frames I frame , through the image recognition model M img Extract scene feature vector Used for subsequent voice scene adaptation.

[0029] (2) Build a local micro audio sample library L audio, store audio segments of different types and styles, with each sample associated with multi-dimensional metadata, including timbre label T label , emotional tendency E tag , semantic scene adaptability C score etc., facilitating quick retrieval and invocation.

[0030] Next, parse and process the speech data through the core speech generation algorithm, including the following four aspects:

[0031] The first aspect is to perform real-time information collection and parsing, specifically:

[0032] Obtain text information T from external input sources (such as user instructions, sensor triggers, etc.) input , and through the natural language processing module NLP module perform word segmentation, part-of-speech tagging, and semantic understanding to generate a structured text representation At the same time, combine the audio input stream A stream to analyze the current ambient noise level N level , and use the camera image frame to judge the scene activity A activity , and these environmental factors will be used as adjustment parameters for subsequent speech generation.

[0033] The second aspect is to perform dynamic sample screening, specifically:

[0034] Based on the structured text and environmental factors, generate initial sample screening conditions. Let the key semantic feature vector in the text be Calculate the semantic correlation degree R between each audio sample in the sample library and semantic (i), using a method based on the cosine similarity of word vectors combined with semantic weight assignment

[0035]

[0036] where w j is the semantic weight of keyword j, is the semantic feature vector of sample i corresponding to keyword j. Combine the timbre, emotional label of the sample and the current environmental adaptability, and comprehensively sort and screen to obtain the candidate audio sample set C samples .

[0037] During the speech interaction process, continuously monitor the user feedback index F feedback (such as the frequency of user repeated instructions, changes in speech recognition accuracy, etc.), dynamically adjust the sample screening weight according to the feedback. If it is found that a certain type of sample causes more negative feedback, reduce its subsequent screening priority to achieve adaptive optimization.

[0038] In the third aspect, audio fusion generation based on the attention mechanism is as follows:

[0039] Introduce the multi-head attention mechanism M multihead Process the candidate audio sample set C samples Each sample's feature vector is used as the input, and the attention weights W (i) of each sample in the current speech generation task are calculated through the attention layer, enabling the model to focus on the audio features that best match the text and scene att (i), enabling the model to focus on the audio features that best match the text and scene

[0040] Fuse the weighted sample features to generate a fused feature vector And input it into the generator G lite in the lightweight generative adversarial network GAN net , combined with the text information Generate the final speech signal V output , which can be expressed as

[0041] The discriminator D of the generative adversarial network net is used to discriminate the authenticity of the generated speech, and continuously improve the quality of the generated speech through adversarial training to make it close to natural speech

[0042] In the fourth aspect, after the above speech generation is completed, in order to further optimize the model's speech generation, the model is optimized and adaptively adjusted as follows:

[0043] Considering the limited resources of the ESP32, an incremental training strategy is adopted to optimize the model. Periodically collect a small amount of new speech samples and corresponding interaction feedback data, and use the online learning algorithm O learn to fine-tune the model parameters. For example, when updating the parameters θ net of the generator G G during backpropagation, use a variant algorithm of mini-batch stochastic gradient descent

[0044]

[0045] where η is the learning rate, Loss GAN is the generative adversarial loss, Loss adapt is the adaptive adjustment loss (calculated based on user feedback), and λ is the balance coefficient to ensure that the model can adapt to new data without overly increasing the computational burden

[0046] In summary, the real-time speech processing method based on the embedded development board of the present invention can achieve the following technical effects:

[0047] Optimize the algorithm according to the resource characteristics of the ESP32 development board: fully consider the storage capacity and computing power of the ESP32 chip, design a lightweight and efficient voice processing flow, reduce unnecessary data redundancy and complex operations, and ensure stable operation in the embedded environment.

[0048] Enhance voice generation through multi-modal fusion perception: innovatively fuse audio input and camera visual information, use the visual scene to assist in understanding the user's intention and environmental atmosphere, and dynamically adjust the voice generation strategy to make the voice output more in line with the actual situation.

[0049] Adaptive dynamic sample matching algorithm: abandon the traditional static audio sample retrieval method, dynamically update the sample selection weight according to the real-time voice interaction feedback, optimize the selection of synthesized audio samples, and improve the accuracy and flexibility of voice generation.

[0050] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A real-time speech processing method based on an embedded development board, characterized in that: The steps include: Step 1: Build a voice interaction system based on an embedded development board to store audio clips and capture sound signals in real time; Step 2: Analyze the sound signal captured in real time, collect and analyze the text information obtained from the external input source, filter the dynamic samples, and then generate audio fusion based on the attention mechanism to complete the speech processing, and finally optimize the model and perform adaptive adjustment.

2. The real-time speech processing method based on the embedded development board according to claim 1, characterized in that: The specific steps of building a voice interaction system based on an embedded development board in step 1 are as follows: Step 1: Build a voice interaction system with the ESP32 development board as the core. Its built-in audio input module captures sound signals in real time and converts them into digital audio streams after preprocessing. stream ; The camera synchronously collects visual image frames I frame , through the image recognition model M img Extract scene feature vector Used for subsequent voice scene adaptation; Step 1 and 2: Build a local micro audio sample library L audio , stores audio clips of different types and styles, and each sample is associated with multi-dimensional metadata, including timbre tags T label 、Emotional tendency tag , semantic scene adaptation C score , convenient for quick retrieval and call.

3. The real-time speech processing method based on the embedded development board according to claim 2 is characterized in that: The specific method of analyzing the sound signal captured in real time and collecting text information obtained from the external input source and then analyzing it is as follows: obtaining text information T from the external input source input , through the natural language processing module NLP module Perform word segmentation, part-of-speech tagging, and semantic understanding to generate structured text representation At the same time, combined with audio input stream A stream Analyze the current environmental noise level N level , using the camera image frame to determine the scene activity A activity ,The above environmental factors serve as regulatory parameters for subsequent speech generation.

4. The real-time speech processing method based on the embedded development board according to claim 2 or 3, characterized in that: The specific method of screening dynamic samples in step 2 is: first, based on the structured text and environmental factors, generate the initial sample screening conditions, and then continuously monitor the user feedback index F of the generated speech during the voice interaction process feedback , dynamically adjust the sample screening weight according to the feedback. If it is found that a certain type of sample causes more negative feedback, reduce its subsequent screening priority to achieve adaptive optimization. Finally, filter the dynamic samples through the optimized screening conditions.

5. The real-time speech processing method based on the embedded development board according to claim 4 is characterized in that: The specific method of generating the initial sample screening condition in step 2 is: assuming that the key semantic feature vector in the text is Calculate the audio samples in the sample library and The semantic relevance R semantic (i) A method based on word vector cosine similarity combined with semantic weight allocation Among them, w j is the semantic weight of keyword j, is the semantic feature vector of keyword j corresponding to sample i. Combined with the timbre, emotional label and adaptability of the sample to the current environment, the candidate audio sample set C is selected through comprehensive sorting samples .

6. The real-time speech processing method based on the embedded development board according to claim 1, 2 or 3, characterized in that: The specific steps of performing audio fusion generation based on the attention mechanism in step 2 are as follows: Step 21: Introduce the multi-head attention mechanism M multihead For the candidate audio sample set C sample to process; Step 22: Fuse the weighted sample features to generate a fused feature vector And input to the lightweight generative adversarial network GAN lite The generator G in net , combined with text information Generate the final speech signal V output , expressed as: Step 2 and 3: Generate the discriminator D of the adversarial network net It is used to determine the authenticity of generated speech and continuously improve the quality of generated speech through adversarial training to make it close to natural speech.

7. The real-time speech processing method based on the embedded development board according to claim 6 is characterized in that: In step 21, the multi-head attention mechanism M is introduced multihead For the candidate audio sample set C samples The specific steps of processing are: the feature vector of each sample As input, the attention weight W of each sample in the current speech generation task is calculated through the attention layer. att (i) This enables the model to focus on the audio features that best match the text and scene.

8. The real-time speech processing method based on the embedded development board according to claim 1, 2 or 3, characterized in that: The specific method of performing model optimization and adaptive adjustment in step 2 is: regularly collecting a small amount of new voice samples and corresponding interactive feedback data, and using online learning algorithm 0 learn Fine-tune model parameters, such as updating the generator G in back-propagation net Parameter θ G When , a mini-batch stochastic gradient descent variant is used Among them, η is the learning rate, Loss GAN To generate adversarial loss, Loss adapt is the adaptive adjustment loss, and λ is the balance coefficient.

Citation Information

Patent Citations

  • Multi-modal fusion user intention recognition method in complex human-computer interaction scene

    CN116661603A