Ultra-low bit rate speech communication system for internet of things applications

By utilizing a speech semantic extraction model and narrowband IoT protocol, an ultra-low bit rate voice communication system for the Internet of Things (IoT) solves the problems of low coding efficiency and poor voice quality in IoT, achieving efficient voice transmission at extremely low bit rates, and is suitable for resource-constrained IoT environments.

CN120853593BActive Publication Date: 2025-11-28JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511323847.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-28
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Traditional voice communication coding methods suffer from low coding efficiency and poor voice quality in the Internet of Things (IoT), making them unsuitable for resource-constrained scenarios, especially under low bandwidth and low power consumption conditions, where they struggle to meet the actual needs of voice communication.

Method used

Design an ultra-low bit rate voice communication system for Internet of Things (IoT) applications, including voice signal acquisition, preprocessing, semantic feature extraction, and communication modules. Utilize a voice semantic extraction model and a narrowband IoT protocol, perform data compression and semantic feature extraction through a cloud-based model, and output structured semantic commands.

Benefits of technology

It achieves efficient voice transmission at extremely low bit rates, improving the accuracy and real-time performance of voice recognition. It is suitable for IoT environments with limited resources, power consumption sensitivity, and limited communication bandwidth, and has good real-time performance and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853593B_ABST
    Figure CN120853593B_ABST
Patent Text Reader

Abstract

The application relates to an ultra-low-bit-rate speech communication system for Internet of Things application, and belongs to the technical field of speech signal processing. The technical problem that the existing technical conventional speech communication coding mode has low coding efficiency, poor speech quality and cannot adapt to the resource-limited scene of the Internet of Things is solved. The ultra-low-bit-rate speech communication system for Internet of Things application comprises a speech signal acquisition module, a speech signal preprocessing module, a semantic feature extraction module, a communication module and a remote instruction execution module which are sequentially connected. The ultra-low-bit-rate speech communication system for Internet of Things application can realize effective expression and transmission of speech information under the condition of ultra-low bit rate by combining speech compression, feature extraction and semantic understanding, is especially suitable for the Internet of Things application environment which is resource-limited, power-sensitive and communication bandwidth-limited, and has good real-time performance, stability and deployment flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of signal processing, and particularly relates to an extremely low bit rate speech communication system for Internet of Things application. BACKGROUND

[0002] With the rapid development of Internet of Things (IoT) technology, speech as a natural and intuitive human-computer interaction method has been widely used in smart home, wearable devices, remote monitoring, industrial automation and other scenarios. Speech interaction not only improves the intelligent degree of devices, but also significantly enhances user experience. However, IoT terminal devices generally have limited computing power, tight storage space, limited communication bandwidth and strict power consumption requirements, and other problems, especially in large-scale deployment and battery-powered application environments, these resource limitations are particularly prominent. Traditional speech communication coding methods, such as G.711, G.729, AMR-NB, etc., although perform well in general communication systems, but they usually require high bit rates, which are difficult to adapt to the transmission requirements of ultra-low bandwidth and low power consumption in Internet of Things. For example, the bit rate of G.711 is 64kbps, the bit rate of G.729 is 8kbps, and even the relatively efficient AMR-NB has a minimum bit rate of more than 4.75kbps, while in low-power wide-area network (LPWAN) technologies such as LoRa, NB-IoT, etc., the actual available data bandwidth is usually only a few hundred bits per second, far from meeting the requirements of traditional speech communication coding methods. This has greatly limited the application of speech communication in Internet of Things.

[0003] In view of the above problems, some existing researches have proposed low bit rate speech coding technologies, such as Codec2 and MELPe, etc. These codecs have improved compression efficiency, with a minimum of 1.2kbps or even lower, but their speech quality still decreases significantly at extremely low bit rates, with severe speech distortion and reduced intelligibility, which is difficult to meet the actual application requirements of speech communication. More importantly, most of these schemes are designed for traditional speech communication, and lack adaptation and optimization for the unique communication mode of Internet of Things. In addition, the communication of Internet of Things system usually has the characteristics of short-time burst, strong event triggering, non-continuous speech transmission, etc., and the traditional continuous speech coding mechanism has low resource utilization efficiency in such scenarios, and there is a risk of waste of computing and communication resources. Therefore, it is difficult to build an Internet of Things speech transmission system that takes into account performance and practicality by relying only on traditional low bit rate coding technology.

[0004] In conclusion, there is an urgent need for a communication system specifically designed for IoT scenarios that can achieve voice transmission at extremely low bit rates. This system should ensure basic intelligibility and interactivity of the voice while minimizing data transmission load, thereby enhancing the practical application value of voice functionality in resource-constrained IoT systems and expanding its application space in areas such as smart terminals, remote voice control, and edge computing. Summary of the Invention

[0005] This invention addresses the technical problems of low coding efficiency, poor voice quality, and inability to adapt to resource-constrained scenarios in traditional voice communication encoding methods, providing an ultra-low bit rate voice communication system for IoT applications. The voice communication system of this invention enables efficient voice transmission under conditions of low bandwidth, low power consumption, and low computing resources.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0007] An ultra-low bit rate voice communication system for Internet of Things (IoT) applications includes a voice signal acquisition module, a voice signal preprocessing module, a semantic feature extraction module, a communication module, and a remote command execution module connected in sequence.

[0008] The voice signal acquisition module is used to acquire the user's raw voice signal and convert the raw voice signal into a digital voice signal;

[0009] The speech signal preprocessing module is used to perform noise reduction, gain adjustment, framing, and window function processing on the digital speech signals acquired by the speech signal acquisition module.

[0010] The semantic feature extraction module is used to extract high-level semantic information from the speech signal preprocessed by the speech signal preprocessing module. This semantic feature extraction module decodes the compressed speech features based on the speech semantic extraction model, semantic classifier or graph neural network in the cloud, extracts high-level semantic information, and outputs structured semantic commands.

[0011] The communication module is used to send the high-level semantic information extracted by the semantic feature extraction module to the remote instruction execution module in the form of data packets with an extremely low bit rate via a narrowband IoT communication protocol;

[0012] The remote instruction execution module is used to receive data packets sent by the communication module and then perform corresponding control operations or recording tasks based on their content.

[0013] In the above technical solution, the voice signal acquisition module uses a single microphone to acquire the user's original voice signal using the sampling theorem. The formula for the sampling theorem is as follows:

[0014] ;

[0015] where, fs is the sampling frequency, fmax is the highest frequency in the signal;

[0016] The sampled speech signal is quantized into discrete digital values, quantization error is represented as:

[0017] ;

[0018] where, and are the maximum and minimum values of the analog signal, respectively, is the number of quantization bits.

[0019] In the above technical solution, the speech signal preprocessing module:

[0020] First, a digital filter is used to denoise the digital speech signal, including removing low-frequency noise and high-frequency interference;

[0021] The frequency response of the digital filter is represented as:

[0022] ;

[0023] where, is the frequency, is the cutoff frequency, is the order of the filter;

[0024] Then, the automatic gain control (AGC) technique is used to adjust the volume of the digital speech signal, ensuring that the signal is within an appropriate range. The gain of the automatic gain control technique is calculated by the following formula:

[0025] ;

[0026] where, is the target amplitude, is the amplitude of the current signal, is the gain factor;

[0027] Next, the digital speech signal is divided into multiple frames, with each frame serving as an independent processing unit. Within each frame, the local characteristics of the signal are extracted through short-time Fourier transform, whose formula is:

[0028] ;

[0029] where, is the time-domain signal, is the window function, is the frequency, is the result of the short-time Fourier transform, is a unit of imaginary number, is a time delay, is a differential operator.

[0030] In the above technical solution, the voice signal acquisition module uses multiple microphone arrays to collect the original voice signal of the user, and needs to ensure that the signals collected by each microphone are synchronized. The signal synchronization is realized by calculating the time delay difference through the cross-correlation method. The time delay difference calculation formula is as follows:

[0031] ;

[0032] wherein, and are the sequences of two signals, is a time delay value, is a cross-correlation function, is a discrete time index, and the best synchronization point is found by the maximum value.

[0033] In the above technical solution, the voice semantic extraction model is a BERT model trained by an Internet of Things data set, and a self-supervised model wav2vec2.0 after fine-tuning is combined to constitute.

[0034] In the above technical solution, the semantic feature extraction module:

[0035] The self-supervised model wav2vec2.0 extracts context-related features from the spectrogram, including:

[0036] First, the preprocessed original voice signal is extracted by a convolutional neural network (CNN) to extract local features , which is represented as:

[0037] ;

[0038] Then, the local features learn long-term dependencies through the Transformer layer of the BERT model, and output a series of context-related acoustic feature vectors for the self-supervised model wav2vec2.0, which is represented as:

[0039] ;

[0040] The loss function is used to fine-tune the self-supervised model wav2vec2.0 to obtain accurate text information;

[0041] The loss function is represented as follows:

[0042] ;

[0043] where, is the conditional probability of the target label sequence given the input sequence , is the input sequence, is the target label sequence, is the alignment path, is the compression mapping function, is the probability that the model outputs an alignment path given the input sequence ;

[0044] Input text information: where each is a word in the text, and each word is passed through the embedding layer of the BERT model to get a word embedding representation

[0045] ;

[0046] where, is the word vector at the th position, is the embedding function that maps discrete symbols to a dense vector space, is the th discrete symbol, is the total length of the target sequence, is the index of the time step;

[0047] The BERT model uses a Transformer layer to encode the input word embeddings bidirectionally to generate a context-dependent feature representation for each position ;

[0048] ;

[0049] where, denotes bidirectional encoding of the word embeddings;

[0050] The context-dependent feature representation generated by the BERT model is used for the Named Entity Recognition (NER) task, which identifies key entities in the text. The output of the Named Entity Recognition task is a label for each position, indicating the entity at that position.

[0051] ;

[0052] where, denotes the Performing named entity recognition.

[0053] The present application has the following advantages:

[0054] The low-bit-rate speech communication system for Internet of Things application of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate, by optimizing the compression and transmission process of speech features, i.e. the semantic feature extraction module performs data compression and coding through the cloud-based speech semantic extraction model, extracts high-level semantic information, and outputs structured semantic commands; the decoding and semantic extraction of the speech signal are completed at the terminal. Therefore, the speech communication system of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate.

[0055] The low-bit-rate speech communication system for Internet of Things application of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate, by optimizing the compression and transmission process of speech features, i.e. the semantic feature extraction module performs data compression and coding through the cloud-based speech semantic extraction model, extracts high-level semantic information, and outputs structured semantic commands; the decoding and semantic extraction of the speech signal are completed at the terminal. Therefore, the speech communication system of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate.

[0056] The low-bit-rate speech communication system for Internet of Things application of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate, by optimizing the compression and transmission process of speech features, i.e. the semantic feature extraction module performs data compression and coding through the cloud-based speech semantic extraction model, extracts high-level semantic information, and outputs structured semantic commands; the decoding and semantic extraction of the speech signal are completed at the terminal. Therefore, the speech communication system of the present application can effectively reduce the amount of data transmission, improve the accuracy and real-time performance of speech recognition, and ensure efficient voice interaction under the condition of extremely low bit rate. BRIEF DESCRIPTION OF DRAWINGS

[0057] The present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0058] Figure 1 The structure block diagram of an embodiment of the low-bit-rate speech communication system for Internet of Things application of the present application.

[0059] Figure 2 The working flowchart of the low-bit-rate speech communication system for Internet of Things application of the present application. DETAILED DESCRIPTION

[0060] The present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0061] Reference Figure 1As shown, the Internet of Things-oriented extremely low bit rate voice transmission system of the present application comprises a voice signal acquisition module, a voice signal preprocessing module, a semantic feature extraction module, a communication module and a remote instruction execution module connected in sequence. The voice signal acquisition module is used to collect the original voice signal of the user in the environment in real time through a sensor and convert it into a digital voice signal. The voice signal acquisition module can include a microphone, an analog-to-digital converter (ADC) and a basic signal conditioning circuit, which is used to stably acquire voice input under different noise environments and ensure that the sampling rate and resolution during the signal acquisition process meet the subsequent processing requirements. The voice signal preprocessing module is responsible for performing preliminary processing such as denoising, gain adjustment, framing and window function processing on the digital voice signal collected by the voice signal acquisition module, such as denoising, echo cancellation, endpoint detection, pre-emphasis, frame segmentation and window function processing, to enhance the clarity of the voice signal and reduce the interference of background noise, provide clean and structured voice feature input for semantic extraction, and provide a clear signal for subsequent semantic extraction. The semantic feature extraction module is used to extract high-level semantic information from the voice signal preprocessed by the voice signal preprocessing module, including command word recognition, voice intent understanding or keyword extraction operations, such as extracting key information (such as the emotion of the voice, the content of the instruction, etc.) according to the features in the voice signal to reduce the amount of data transmitted and improve efficiency; the semantic feature extraction module can decode the compressed voice features based on a small cloud (i.e. cloud server side) side voice semantic extraction model, semantic classifier or graph neural network, extract high-level semantic information, and output structured semantic commands or control commands for driving the target device to perform corresponding tasks. The semantic feature extraction module can dynamically adjust the recognition strategy according to system resources to achieve semantic recognition under different precision and complexity. The communication module is used to send the high-level semantic information (i.e. key information) extracted by the semantic feature extraction module to the remote instruction execution module in the form of an extremely low bit rate data packet through a narrowband Internet of Things communication protocol. The remote instruction execution module is used to complete corresponding control operations or record tasks according to the content of the data packet sent by the communication module after receiving the data packet. The remote instruction execution module controls related hardware or services by receiving instructions from remote devices to complete system target tasks.

[0062] The various modules included in the Internet of Things application-oriented extremely low bit rate voice communication system of the present application will be described more clearly and in detail below.

[0063] The voice signal acquisition module collects the original voice signal of the user through the microphone of the terminal device mobile phone or computer, etc. The Python programming software calls the hardware devices such as the microphone, and the transmitted instruction signal is text information.

[0064] When the voice signal acquisition module uses a single microphone to collect the original voice signal, the sound wave signal is first converted into an electrical signal, and the analog signal is converted into a digital voice signal through a signal acquisition card (such as DAQ). In this process, the sampling theorem is applied for sampling, i.e. according to the formula:

[0065]

[0066] wherein, is the sampling frequency, is the highest frequency in the signal, and the sampling frequency is ensured to meet the Nyquist theorem to avoid aliasing. The sampling frequency is 16 kHz. The sampled signal will be quantized into discrete digital values, and the quantization error can be represented as:

[0067]

[0068] wherein, and are the maximum and minimum values of the analog signal, is the quantization bit number.

[0069] The voice signal preprocessing module first uses a digital filter to remove low-frequency noise and high-frequency interference. The frequency response of the digital filter is usually:

[0070]

[0071] wherein, is the frequency, is the cutoff frequency, is the order of the filter.

[0072] Then, the automatic gain control (AGC) technique is used to adjust the signal volume to ensure that the signal is within an appropriate range. The gain of AGC can be calculated by the following formula:

[0073]

[0074] wherein, is the target amplitude, is the amplitude of the current signal, is the gain factor.

[0075] Next, the digital voice signal is divided into multiple frames (usually 20-40 milliseconds), and each frame is treated as an independent processing unit. In each frame, the local features of the signal can be extracted through short-time Fourier transform (STFT), and the formula is:

[0076]

[0077] wherein,​​​​​ is a time domain signal, is a window function, is a frequency, is a result of short-time Fourier transform, is an imaginary unit, is a time delay, is a differential operator.

[0078] If the voice signal acquisition module uses a multi-microphone array to acquire signals, it is also necessary to ensure that the signals acquired by each microphone are synchronized to ensure that the signals acquired by each microphone are effectively combined at the same time. Synchronization can be achieved by calculating the time delay difference. The common method is cross-correlation (Cross-correlation). The formula for calculating the time delay difference is as follows:

[0079] ;

[0080] wherein, and are sequences of two signals, is a time delay value, is a cross-correlation function, is a discrete time index, and the best synchronization point is found by the maximum value.

[0081] The semantic feature extraction module decodes the compressed voice features based on the cloud-based voice semantic extraction model, extracts high-level semantic information, and outputs structured semantic commands. The training of the voice semantic extraction model is described in Figure 1 , which is carried out on the cloud server. The main steps are to first build the data set of the Internet of Things to which it belongs, then train the BERT model, and combine the fine-tuned self-supervised model wav2vec 2.0 to form the voice semantic extraction model, and download it to the terminal device.

[0082] The semantic feature extraction module uses the self-supervised model wav2vec2.0 to extract context-related features from the spectrogram. First, the pre-processed original voice signal (typically a 16kHz or higher sampling rate audio file) is extracted through a convolutional neural network (CNN) to extract local features :

[0083] ;

[0084] Then, the local features are learned by the Transformer layer of the BERT model to learn long-term dependencies, and a series of context-related acoustic feature vectors are output by the self-supervised model wav2vec2.0. These vectors can reflect the voice content and context information in different time periods of the voice signal.

[0085] ;

[0086] The self-supervised model wav2vec2.0 is fine-tuned using a loss function to obtain accurate text information.

[0087] loss function It is expressed as follows:

[0088] ;

[0089] in, To in a given input sequence Under the condition of target label sequence The conditional probability of occurrence Given the input sequence, For the target label sequence, For alignment paths, which are possible paths in the model's output sequence, each path includes a combination of target symbols and whitespace characters. For compression mapping functions, it will align the path Mapping to target label sequence , In the input sequence Below, the model output alignment path is The probability of;

[0090] First, the input text information Each of them It is a word in the text. Each word goes through the embedding layer of the BERT model to obtain the word embedding representation as follows:

[0091] ;

[0092] in, For the first Word vectors at each position, To make discrete symbols Embedding functions mapped to dense vector spaces. For the first A discrete symbol, The total length of the target sequence. The index is the time step index, indicating which symbol in the sequence is currently being processed;

[0093] The BERT model uses Transformer layers to bidirectionally encode the input word embeddings, generating each position... Context-dependent feature representation ,in The context of the word was taken into consideration.

[0094] ;

[0095] wherein, denotes bi-directional encoding of word embeddings;

[0096] Contextual feature representation generated by BERT model BERT model can perform Named Entity Recognition (NER) task to identify key entities in the text, such as device name, operation type, etc. The output of NER is a label for each position, indicating the entity at that position.

[0097] ;

[0098] wherein, denotes performing Named Entity Recognition on .

[0099] As in the text corresponding to the voice is "turn up the living room temperature to 22 degrees", NER may identify the following entities.

[0100] ;

[0101] Finally generate structured semantic instructions , containing device name, operation type and numerical parameter information.

[0102] ;

[0103] The communication module transmits the compressed voice data stream (i.e. semantic instructions) between IoT devices through low bit rate, ensuring fast and stable transmission of voice data in bandwidth-limited network environment.

[0104] The remote instruction execution module first parses the structured instructions to identify the device name, operation type and numerical parameter. The parsed result is: device name = living room temperature, operation type = turn up, numerical parameter = 22°C. According to the parsing result, the instruction is converted into a control command. In the IoT system, the control command can be a device API request or a hardware control instruction, depending on the communication protocol of the device and platform. The form of control command is as follows:

[0105] .

[0106] Referring to Figure 2The working process of the low-bit-rate voice communication system for Internet of Things application of the present application is as follows (only the working process steps are shown in the figure): the voice signal of a user is collected by a voice signal collection module in real time through a sensor (i.e. the voice is input through a microphone) and converted into a digital voice signal by an analog-to-digital converter (ADC); then, the digital voice signal enters a voice signal preprocessing module to perform noise reduction (noise suppression), gain adjustment, and frame processing, etc. to ensure the signal quality; the processed signal is transmitted to a semantic feature extraction module to extract key information and reduce the data volume; subsequently, the extracted semantic features are recognized and mapped by an instruction recognition module to obtain structured instruction information; the extracted compressed signal is packaged and transmitted between devices at a low bit rate by a communication module; finally, the signal reaches a remote instruction execution module which decodes the structured instruction and executes the corresponding hardware operation or task according to the execution instruction, and the synthesized voice is fed back to the user.

[0107] In summary, the low-bit-rate voice communication system for Internet of Things application of the present application can effectively express and transmit voice information under the condition of a very low bit rate by combining voice compression, feature extraction, and semantic understanding, and is particularly suitable for the Internet of Things application environment with limited resources, sensitive power consumption, and limited communication bandwidth, and has good real-time performance, stability, and deployment flexibility.

[0108] Obviously, the above embodiments are only examples for the purpose of clear illustration, and are not intended to limit the embodiments. Based on the above description, other different forms of changes or variations can be made by those of ordinary skill in the art. All the embodiments do not need to be exhausted, and the obvious changes or variations derived therefrom are still within the protection scope of the present application.

Claims

1. An ultra-low bit rate voice communication system for Internet of Things (IoT) applications, characterized in that, It includes a voice signal acquisition module, a voice signal preprocessing module, a semantic feature extraction module, a communication module, and a remote command execution module connected in sequence; The voice signal acquisition module is used to acquire the user's raw voice signal and convert the raw voice signal into a digital voice signal; The speech signal preprocessing module is used to perform noise reduction, gain adjustment, framing, and window function processing on the digital speech signals acquired by the speech signal acquisition module. The semantic feature extraction module is used to extract high-level semantic information from the speech signal preprocessed by the speech signal preprocessing module. Based on the cloud-based speech semantic extraction model, the semantic feature extraction module decodes the compressed speech features, extracts high-level semantic information, and outputs structured semantic commands. The communication module is used to send the high-level semantic information extracted by the semantic feature extraction module to the remote instruction execution module in the form of data packets with an extremely low bit rate via a narrowband IoT communication protocol; The remote command execution module is used to receive data packets sent by the communication module and then perform corresponding control operations or recording tasks based on their content. The semantic feature extraction module: The self-supervised model wav2vec2.0 is used to extract context-related features from the spectrogram, including: First, the preprocessed raw speech signal Local features are extracted using a convolutional neural network. , represented as: ; Then, local features By learning long-term dependencies through the Transformer layer of the BERT model, a series of context-dependent acoustic feature vectors are output for the self-supervised model wav2vec2.

0. , represented as: ; The self-supervised model wav2vec2.0 is fine-tuned using a loss function to obtain accurate text information; loss function It is expressed as follows: ; in, To in a given input sequence Under the condition of target label sequence The conditional probability of occurrence Given the input sequence, For the target label sequence, To align the path, For compression mapping functions, In the input sequence Below, the model output is the alignment path. The probability of; Enter text information: Each of them It is a word in the text. Each word goes through the embedding layer of the BERT model to obtain the word embedding representation as follows: ; in, For the first Word vectors at each position, To map discrete symbols Embedding functions to dense vector spaces For the first A discrete symbol, The total length of the target sequence. For the index of the time step; The BERT model uses Transformer layers to bidirectionally encode the input word embeddings, generating each position... Context-dependent feature representation ; ; in, This indicates bidirectional encoding of word embeddings; Context-dependent feature representations generated by the BERT model The task of named entity recognition is to identify key entities in text. The output of the named entity recognition task is a label for each position. , representing the entity at that location; ; in, Indicates to Perform named entity recognition.

2. The ultra-low bit rate voice communication system for IoT applications according to claim 1, characterized in that, The speech signal acquisition module uses a single microphone to acquire the user's raw speech signal using the sampling theorem, the formula of which is as follows: ; in, Sampling frequency, The highest frequency in the signal; The sampled speech signal is quantized into discrete digital values, and the quantization error... Represented as: ; in, and These are the maximum and minimum values ​​of the analog signal, respectively. The number of bits used for quantization.

3. The ultra-low bit rate voice communication system for IoT applications according to claim 1, characterized in that, The speech signal preprocessing module: First, digital filters are used to denoise the digital speech signal, including removing low-frequency noise and high-frequency interference. Frequency response of digital filters Represented as: ; in, It's frequency. It is the cutoff frequency. It is the order of the filter; Then, automatic gain control (AGDC) technology is used to adjust the volume of the digital voice signal to ensure that the signal is within an appropriate range. The gain of AGDC is calculated using the following formula: ; in, It is the target range. It is the amplitude of the current signal. It is the gain factor; Next, the digital speech signal is divided into multiple frames, each serving as an independent processing unit. Within each frame, the local features of the signal are extracted using a short-time Fourier transform, with the following formula: ; in, For time-domain signals, For window functions, For frequency, This is the result of the short-time Fourier transform. The imaginary unit, For time delay, It is a differential operator.

4. The ultra-low bit rate voice communication system for IoT applications according to claim 1, characterized in that, The voice signal acquisition module uses multiple microphone arrays to collect the user's raw voice signal. It is necessary to ensure that the signals collected by each microphone are synchronized. Signal synchronization is achieved by calculating the time delay difference using a cross-correlation method. The formula for calculating the time delay difference is as follows: ; in, and It is a sequence of two signals. This is the delay value. It is a cross-correlation function. For discrete-time indexing, the optimal synchronization point is found by using the maximum value.

5. The ultra-low bit rate voice communication system for Internet of Things applications according to any one of claims 1-4, characterized in that, The speech semantic extraction model can also decode compressed speech features based on semantic classifiers or graph neural networks, extract high-level semantic information, and output structured semantic commands.

Citation Information

Patent Citations

  • Remote speech enhancement transmission method and system based on semantic communication

    CN119296566A

  • Speech recognition method and system based on artificial intelligence

    CN120564724A