Audio transmission methods, devices, and storage media based on hierarchical compression

By monitoring network status and timbre characteristics, the compression level of audio signals is dynamically adjusted, solving the problem of low transmission efficiency of long audio signals and achieving efficient audio signal transmission and feedback response.

CN121811891BActive Publication Date: 2026-05-26EARDA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
EARDA TECH CO LTD
Filing Date
2026-03-06
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In home terminal devices, the transmission of long audio signals consumes a lot of bandwidth, resulting in low transmission efficiency and affecting subsequent feedback and response efficiency, especially in question-and-answer scenarios such as parent-child education.

Method used

By monitoring network status, audio signals are compressed in stages, and the compression level is dynamically adjusted based on network status, chat information, and timbre characteristics to achieve adaptive transmission of audio signals, balancing resource consumption and speech recognition accuracy.

Benefits of technology

It improves the transmission efficiency of audio signals, maintains the feedback and response efficiency of terminal devices, and adapts to the needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811891B_ABST
    Figure CN121811891B_ABST
Patent Text Reader

Abstract

This invention provides an audio transmission method, device, and storage medium based on hierarchical compression. The method includes: monitoring the network status between the device and a terminal device; upon receiving a target audio signal transmitted from the terminal device, decompressing the target audio signal into an original audio signal according to the current compression level; encoding the original audio signal into basic audio features; extracting timbre features from the basic audio features; decoding the basic audio features according to the current compression level to obtain chat information; generating a new compression level based on the network status, chat information, and timbre features; and transmitting the new compression level to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level. This embodiment effectively improves the efficiency of audio signal transmission, maintains the feedback efficiency towards the terminal device, and maintains the user-facing response efficiency of the terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, and in particular relates to an audio transmission method, device and storage medium based on hierarchical compression. Background Technology

[0002] Currently, smart terminal devices in homes, such as speakers, televisions, and lamps, are gradually being integrated with Large Language Models (LLMs) to provide users with various voice services.

[0003] The terminal device collects the audio signal emitted by the user and transmits the raw audio signal directly to the cloud. The cloud performs speech recognition on the audio signal to obtain text information, calls a large language model to understand the text information, and feeds back to the terminal device so that the terminal device can respond as quickly as possible.

[0004] However, in scenarios such as Q&A (especially parent-child education), audio signals are long and large in size, which consumes more bandwidth and takes longer to transmit, resulting in lower transmission efficiency and affecting the efficiency of subsequent feedback and response. Summary of the Invention

[0005] In view of this, the present invention provides an audio transmission method, device and storage medium based on hierarchical compression to improve the efficiency of transmitting audio signals.

[0006] A first aspect of the present invention provides an audio transmission method based on hierarchical compression, comprising:

[0007] Monitor the network status between the device and the terminal equipment;

[0008] When receiving the target audio signal transmitted by the terminal device, the target audio signal is decompressed into the original audio signal according to the current compression level;

[0009] Encode the original audio signal into basic audio features;

[0010] Extract the timbre features of the current original audio signal from the basic audio features;

[0011] The basic audio features are decoded according to the current compression level to obtain the chat information;

[0012] A new compression level is generated based on the network status, the chat information, and the timbre characteristics;

[0013] The new compression level is transmitted to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level.

[0014] A second aspect of the present invention provides an audio transmission apparatus based on hierarchical compression, comprising:

[0015] The network status monitoring module is used to monitor the network status between the device and the terminal device.

[0016] An audio signal decompression module is used to decompress the target audio signal into the original audio signal according to the current compression level when receiving the target audio signal transmitted by the terminal device.

[0017] The audio encoding module is used to encode the current original audio signal into basic audio features;

[0018] The timbre feature extraction module is used to extract the timbre features of the current original audio signal from the basic audio features;

[0019] An audio decoding module is used to decode the basic audio features according to the current compression level to obtain chat information;

[0020] A compression level generation module is used to generate a new compression level based on the network status, the chat information, and the timbre characteristics.

[0021] A compression level transmission module is used to transmit a new compression level to the terminal device so as to compress the new original audio signal into a new target audio signal according to the new compression level.

[0022] A third aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the audio transmission method based on hierarchical compression as described in the first aspect above.

[0023] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio transmission method based on hierarchical compression as described in the first aspect above.

[0024] A fifth aspect of the present invention provides a computer program product that, when run on a computer, causes the computer to perform the audio transmission method based on hierarchical compression as described in the first aspect above.

[0025] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0026] In this embodiment, the network status between the terminal device and the device is monitored. Upon receiving the target audio signal transmitted by the terminal device, the target audio signal is decompressed into the original audio signal according to the current compression level. The original audio signal is encoded into basic audio features. The timbre features of the original audio signal are extracted from the basic audio features. The basic audio features are decoded according to the current compression level to obtain chat information. A new compression level is generated based on the network status, chat information, and timbre features. The new compression level is transmitted to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level. This embodiment determines the compression level of the audio signal based on adaptive factors such as network status and dynamic scene of the conversation, and dynamically adjusts speech recognition according to the compression level. It seeks a balance between the resource consumption of the terminal device in transmitting audio signals and the accuracy of speech recognition, effectively improving the efficiency of audio signal transmission, maintaining the feedback efficiency to the terminal device, and maintaining the response efficiency of the terminal device to the user. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of an audio transmission method based on hierarchical compression provided in an embodiment of the present invention;

[0029] Figure 2 This is a schematic diagram of the structure of a composite speech model provided in an embodiment of the present invention;

[0030] Figure 3 This is a schematic diagram of the structure of a timbre extraction module provided in an embodiment of the present invention;

[0031] Figure 4 This is a schematic diagram of the structure of a decoding module provided in an embodiment of the present invention;

[0032] Figure 5 This is a schematic diagram of an audio transmission device based on hierarchical compression provided in an embodiment of the present invention;

[0033] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0034] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the present invention. However, those skilled in the art will recognize that the present application may be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted to avoid unnecessary detail that could obscure the description of the present application.

[0035] The technical solution of the present invention will be illustrated below through specific embodiments.

[0036] Reference Figure 1 The diagram illustrates an audio transmission method based on hierarchical compression provided by an embodiment of the present invention, which may specifically include the following steps:

[0037] Step 101: Monitor the network status between the terminal device and the terminal device.

[0038] In this embodiment, the manufacturer can deploy a server cluster in the cloud and provide AI (artificial intelligence) products. The AI ​​products may include cloud modules, audio boards and other components, which can be easily integrated into other terminal devices, or may include integrated terminal devices such as Rubik's Cubes.

[0039] The cloud module supports Wi-Fi and Bluetooth communication, making it easy to embed into various devices and achieve intelligent upgrades. It provides multiple interfaces, such as I2C (Inter-Integrated Circuit), SPI (Serial Peripheral Interface), and UART (Universal Asynchronous Receiver / Transmitter), which facilitates connection and communication with other hardware (such as microphones and speakers) when integrated into other terminal devices.

[0040] The audio board integrates interfaces such as speakers, microphones, lithium battery management, touch switches, and LEDs (light-emitting diodes) on the basis of the cloud module. It can support lighting linkage and control functions, realize voice control of lighting effects, and also provide accessories such as speakers, microphones, and batteries.

[0041] The Rubik's Cube features a portable design with a built-in battery to improve battery life, making it suitable for mobile use. It is equipped with dual microphones and speakers to ensure clear voice interaction and sound quality, meeting daily voice interaction needs.

[0042] In this embodiment, a dynamic hierarchical compression audio signal transmission mode is used between the terminal device and the cloud server. In this transmission mode, the compression level is agreed upon between the terminal device and the cloud server.

[0043] The compression level can be mapped to different types of compression algorithms (such as Opus, G.729, AAC, etc.). Different types of compression algorithms provide different compression ratios. The compression level can also be mapped to compression parameters in the same type of compression algorithm (such as bit rate, bandwidth, frame rate, VBR (variable bit rate), CBR (constant bit rate), coding complexity (such as the compression parameter OPUS_SET_COMPLEXITY() in Opus), etc.). Different compression parameters provide different compression ratios, and this embodiment does not impose any restrictions on this.

[0044] For example, compression levels include Level 1 compression, Level 2 compression, and Level 3 compression.

[0045] Level 1 compression achieves high bit rates (e.g., 48-64kbps) and full bandwidth (48kHz), approaching the quality of the original speech signal and preserving ambient sound details. When used for regular speech recognition, WER (word error rate) is <5%, but bandwidth overhead increases by 2-4 times.

[0046] Secondary compression uses a medium bit rate (e.g., 16-32kbps) and medium bandwidth (16kHz), preserving the main audio segments and making consonants distinguishable. When used for regular speech recognition, the WER is 5%~10%, making it suitable for most scenarios.

[0047] Level 3 compression uses low bit rates (e.g., 6-8kbps) and narrow bandwidths (8kHz), significantly reducing high-frequency information (e.g., / s / , / f / , / th / , etc.). When used for regular speech recognition, the WER is 15%~30%, and the WER increases particularly noticeably in noisy environments.

[0048] In practical applications, users trigger a session using wake words, key presses, or other methods. At this time, a long connection is established between the terminal device and the cloud server. The cloud server collects specified indicator values ​​at multiple times to monitor the network status between the terminal device and the cloud server.

[0049] For example, the indicator value includes at least one of the following:

[0050] Download speed, upload speed, network latency, packet loss rate, CPU (central processing unit) utilization of the terminal device, and memory utilization of the terminal device.

[0051] In general, algorithms such as Max-Min can be used to normalize the index values ​​to facilitate subsequent calculations.

[0052] Step 102: When receiving the target audio signal transmitted by the terminal device, decompress the current target audio signal into the original audio signal according to the current compression level.

[0053] After being woken up by the user, the terminal device can use the microphone to collect the original audio signal, perform preprocessing such as noise reduction and enhancement on the original audio signal, and use technologies such as Voice Activity Detection (VAD) to remove silent segments from the original audio signal.

[0054] At this point, the terminal device can query the current compression level, configure the compression algorithm according to the current compression level, use the compression algorithm to compress the current original audio signal into the target audio signal, and use encrypted transmission (such as TLS (Transport Layer Security Protocol), SSL (Secure Sockets Layer Protocol) and other methods) to transmit the current target audio signal to the cloud server.

[0055] Initially, the current compression level is the default compression level (such as level 2 compression). During subsequent interactions, the current compression level is set by the cloud server.

[0056] When the cloud server receives the target audio signal transmitted by the terminal device, it queries the current compression level, configures the compression algorithm according to the current compression level, and uses the compression algorithm to decompress the current target audio signal into the original audio signal.

[0057] Step 103: Encode the current original audio signal into basic audio features.

[0058] In this embodiment, a composite speech model can be pre-built and trained, which has functions such as speech recognition and timbre recognition.

[0059] Among them, such as Figure 2 As shown, the composite speech model includes an encoder module. The encoder module can reuse third-party pre-trained model structures, such as Conformer and Transformer, or it can define its own model structure. This embodiment does not impose any restrictions on this.

[0060] The current original audio signal is input into the encoding module, which encodes the current original audio signal into a set of high-dimensional speech feature vectors, denoted as the basic audio features.

[0061] The encoding module can be trained in a self-supervised manner to provide neutral (task-unbiased) audio features for tasks such as speech recognition and timbre recognition.

[0062] During training, a generator can be configured for the encoding module. The encoding module is responsible for extracting basic speech features from the sample speech signal, while the generator is responsible for generating a reconstructed speech signal based on the basic speech features with the aim of recovering the sample speech signal. The parameters in the encoding module and the generator are updated based on the loss value between the reconstructed speech signal and the sample speech signal. When training ends, the generator is discarded and the parameters in the encoding module are fixed.

[0063] Step 104: Extract the timbre features of the current original audio signal from the basic audio features.

[0064] like Figure 2 As shown, the composite speech model includes a timbre extraction module, which is responsible for extracting the timbre features of the current original audio signal from the basic audio features.

[0065] Among them, timbre features are the characteristics that represent the speaker's timbre.

[0066] Generally, a terminal device uses one account to log in. Subsequently, in different application scenarios, different users may wake up the terminal device and trigger a session.

[0067] In this embodiment, under conditions such as account authorization upon login on the terminal device, the speaker's context can be tracked from the speaker's timbre characteristics to understand the speaker's basic information (such as gender, approximate age, etc.), thereby understanding different application scenarios (such as parent-child education, children's learning), thus providing a data foundation for business operations such as predicting the speaker's future speech and providing feedback to the speaker.

[0068] For example, such as Figure 3 As shown, the composite speech model includes a Long Short-Term Memory (LSTM) network.

[0069] In this example, raw acoustic features can be extracted from the current raw audio signal, such as 39-dimensional MFCC (Melbourne cepstral coefficients) mean features and 80-dimensional Mel spectrum features.

[0070] The basic audio features output by the encoding module are concatenated with the original acoustic features to form the target acoustic features. The target acoustic features are then input into the long short-term memory network to extract the timbre features of the current original audio signal.

[0071] During training, the encoding module and the timbre extraction module can be regarded as a complete timbre network, which is responsible for extracting timbre features from the speech signal.

[0072] Furthermore, the timbre network can be trained in a supervised manner, during which the parameters of the encoding module are kept unchanged while the parameters of the timbre extraction module are updated.

[0073] Step 105: Decode the basic audio features according to the current compression level to obtain the chat information.

[0074] like Figure 2 As shown, the composite speech model includes a decoding module, which is responsible for decoding the basic audio features according to the current compression level to obtain chat information.

[0075] Generally, the features lost during compression of a target speech signal are positively correlated with the compression level. That is, the higher the compression level and the greater the compression ratio, the more features the target speech signal loses during compression. Conversely, the lower the compression level and the smaller the compression ratio, the fewer features the target speech signal loses during compression.

[0076] Therefore, the decoding module can evaluate the lost features of the basic audio features according to the current compression level, thereby dynamically decoding the basic audio features and seeking a balance between resource consumption and decoding accuracy, thus improving the decoding efficiency.

[0077] During training, the encoding and decoding modules can be regarded as a complete speech recognition network, with the speech recognition module responsible for performing speech recognition on the speech signal.

[0078] Furthermore, the speech recognition network can be trained in a supervised manner, during which the parameters of the encoding module are kept unchanged while the parameters of the decoding module are updated.

[0079] In one embodiment of the present invention, each compression level can be divided into a first level range and a second level range, wherein the compression levels in the second level range are all higher than the compression levels in the first level range.

[0080] For example, if the compression levels include Level 1 compression, Level 2 compression, and Level 3 compression, then the first level ranges from Level 1 compression to Level 2 compression, and the second level ranges from Level 3 compression.

[0081] Therefore, in this embodiment, step 105 may include the following steps:

[0082] Step 1051: If the current compression level belongs to the first level range, then input the basic audio features into the decoder to decode them into low-pressure audio features.

[0083] Step 1052: Input the low-voltage audio features into the head structure to generate chat information.

[0084] In this embodiment, as Figure 4 As shown, the decoding module includes a decoder and a head structure. The decoder is responsible for capturing text temporal dependencies and aligning audio semantics from the features output by the encoding module, while the head structure is responsible for mapping the features output by the decoder into a semantically recognized text sequence.

[0085] In practical applications, the decoding module can reuse third-party pre-trained model structures, such as Conformer and Transformer, and divide them into decoder and head structures. Alternatively, a custom model structure can be defined. This embodiment does not impose any restrictions on this.

[0086] Taking Transformer as an example, the head structure consists of the linear layer at the end and the activation function Softmax, while the rest is the decoder.

[0087] If the current compression level is in the first level range, the basic audio features are lost in small amounts and within an acceptable range. Even if some characters are misidentified, it will not have a negative impact on the LLM's understanding of the speaker's intention based on contextual semantics. In this case, the basic audio features can be input into the decoder to decode into low-pressure audio features, and the low-pressure audio features can be input into the header structure to generate chat information.

[0088] In this case, the decoding module has a simple link structure and low resource consumption.

[0089] Step 1053: If the current compression level belongs to the second level range, the basic audio features are input into the first compensation layer to compensate for the features lost during compression, and the first high-voltage audio features are obtained.

[0090] In this embodiment, as Figure 4 As shown, the decoding module includes a first compensation layer, a decoder, a second compensation layer, and a header structure. The first compensation layer is responsible for compensating for features lost during speech signal compression during encoding, and the second compensation layer is responsible for compensating for features lost during speech signal compression during decoding.

[0091] If the current compression level is in the second level range, many basic audio features are lost, and the number of misidentified texts increases, which may negatively affect the LLM's understanding of the speaker's intention based on contextual semantics. In this case, a decoding module with a longer link structure can be called to perform decoding to improve decoding accuracy.

[0092] In this embodiment, basic audio features can be input into the first compensation layer, which compensates for features lost during compression in the encoding process to obtain the first high-voltage audio features.

[0093] In the specific implementation, the first compensation layer includes an attention layer and a fully connected layer (FC).

[0094] The standard audio features of uncompressed sample audio signals are searched in the preset template library. That is, the standard audio features are the average of the features extracted by the encoder from multiple users' uncompressed sample audio signals, which belong to clean audio features.

[0095] In the first compensation layer, the basic audio features and standard audio features are input into the attention layer. The basic audio features are used as the query matrix Q and the standard audio features are used as the key matrix K and value matrix V. The attention mechanism is applied to generate standard attention weights. The attention mechanism focuses on the group (i.e., general) key acoustic frames (such as vowel frames and consonant frames) and weakens the noise frames, thus compensating for the feature loss caused by compression in the encoding stage.

[0096] The basic audio features are multiplied element-wise with the standard attention weights to obtain the first candidate audio features. The first candidate audio features are then input into the fully connected layer and mapped to the first high-pressure audio features to align the dimensions.

[0097] Step 1054: Input the first high-voltage audio feature into the decoder and decode it into the second high-voltage audio feature.

[0098] In this embodiment, the first high-voltage audio feature can be input into the decoder, and the decoder decodes the first high-voltage audio feature into a second high-voltage audio feature.

[0099] Step 1055: Input the second high-voltage audio feature into the second compensation layer to compensate for the features lost during compression, and obtain the third high-voltage audio feature.

[0100] In this embodiment, the second high-voltage audio feature can be input into the second compensation layer, which compensates for the features lost during compression during the decoding process to obtain the third high-voltage audio feature.

[0101] In its implementation, the second compensation layer includes an attention layer and a fully connected layer (FC).

[0102] Within the context of the conversation, the lowest-compression audio feature in history is found based on timbre characteristics and used as the individual audio feature. That is, the similarity between the timbre characteristics of the current original speech signal and the timbre characteristics of the historical original speech signals in the conversation is calculated. When the similarity is greater than or equal to a preset confidence threshold, it is determined that the current original speech signal and the historical original speech signal belong to the same speaker. For the same speaker, the historical original speech signal with the lowest compression level (i.e., the lowest compression ratio) is selected. If there are multiple historical original speech signals with the lowest compression level, the most recent historical original speech signal is used as the standard, and the low-compression audio feature corresponding to the historical original speech signal is set as the individual audio feature.

[0103] In the second compensation layer, the second high-pressure audio features and individual audio features are input into the attention layer. The second high-pressure audio features are used as the query matrix Q and the individual audio features are used as the key matrix K and value matrix V. An attention mechanism is applied to generate object attention weights. The semantics of the encoded features are weakened. The attention mechanism is used to strengthen the focus of the decoding on the key semantic frames of the individual during the encoding. The feature loss caused by compression is compensated in the decoding stage.

[0104] The second high-pressure audio feature is multiplied element-wise with the object attention weight to obtain the second candidate audio feature. The second candidate audio feature is then input into the fully connected layer and mapped to the third high-pressure audio feature to align the dimensions.

[0105] Step 1056: Input the third high-voltage audio feature into the head structure to generate chat information.

[0106] In this embodiment, the third high-voltage audio feature is input into the head structure, and the head structure generates chat information based on the third high-voltage audio feature.

[0107] In this embodiment, the link structure of the decoding module under low compression is the same as that of the decoding module under high compression. The link of the decoding module under high compression is an enhanced structure of the link of the decoding module under low compression, which realizes structural reuse, reduces development, training and deployment costs, maintains speech recognition consistency, avoids fluctuations in recognition accuracy when switching scenes, has strong scalability, and is easy to maintain and optimize in the later stage.

[0108] In addition, during high compression, the decoding module first compensates for compression loss by encoding based on the characteristics of the group, and then compensates for compression loss by decoding based on the characteristics of the individual, thus achieving layered compensation. Group compensation strengthens the basic recognition accuracy, while individual compensation optimizes personalized adaptation. The two work together to cover both general and specific compensation needs, especially adapting to the needs of multiple users, multiple accents, and multiple timbres, taking into account both generalization and personalization, and can effectively improve the accuracy of features.

[0109] Step 106: Generate a new compression level based on network status, chat information, and voice characteristics.

[0110] In this embodiment, the new compression level for the next stage can be predicted from two dimensions. One dimension is the network status between the terminal device and the cloud server, and the other dimension is the conversational scenario of the terminal device, which is represented by the dialogue process of the speaker constructed from chat information and voice features.

[0111] In one embodiment of the present invention, step 106 may include the following steps:

[0112] Step 1061: Use the current network state to generate network features that represent the future network state.

[0113] In this embodiment, the current network state (time series data) can be input into a preset time series prediction model (such as RNN (Recurrent Neural Network) and its derivative models) to predict the future network state. Features are extracted from the intermediate layers of the time series prediction model to obtain network features representing the future network state.

[0114] Step 1062: Concatenate the timbre features with the text features of the chat messages to form the original dialogue features.

[0115] In this embodiment, each chat message can be input into a preset language model (such as BERT (Bidirectional Encoder Representations from Transformers)) to extract text features.

[0116] For the same target audio signal, its corresponding timbre features can be concatenated with the text features of the chat message to form the original dialogue features.

[0117] Step 1063: Generate target dialogue features representing future chat information based on the original dialogue features.

[0118] In this embodiment, the original dialogue features are the contextual representations of speakers in different roles in the conversation. The future chat information of the speakers can be predicted based on the original dialogue features, thereby generating target dialogue features that represent the chat information spoken by the future speakers. The target dialogue features can characterize the size of the file (speech signal) to a certain extent.

[0119] In the specific implementation, all the original dialogue features can be concatenated into global dialogue features. K-means can be used to cluster the original dialogue features based on timbre features to obtain multiple clusters. Each cluster can represent a sound source (including the speaker, background noise, etc.).

[0120] The number of original dialogue features in each cluster is counted. The cluster with the most original dialogue features is usually the one where the main users of the terminal device are the speakers, and they are most likely to speak in the future.

[0121] Therefore, for the cluster with the largest number of original dialogue features, the original dialogue features in the cluster are concatenated into local dialogue features, and the global dialogue features and local dialogue features are concatenated into enhanced dialogue features.

[0122] The enhanced dialogue features are input into the generative model, and the enhanced dialogue features are used to generate target dialogue features that represent future chat information.

[0123] Step 1064: Fuse network features and target dialogue features into multimodal features.

[0124] In this embodiment, network features and target dialogue features can be fused across modalities into multimodal features.

[0125] For example, network features and target dialogue features are concatenated to form fused features. The fused features are then input into the attention layer. Using the fused features as the query matrix Q, key matrix K, and value matrix V, an attention mechanism is applied to generate self-attention weights. The fused features are then multiplied element-wise with the self-attention weights to obtain multimodal features.

[0126] Step 1065: Classify new compression levels based on multimodal features.

[0127] In this embodiment, multimodal features can be input into a preset classifier for binary or multi-class classification to obtain a new compression level for the next stage.

[0128] Step 107: Transmit the new compression level to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level.

[0129] In this embodiment, the cloud server transmits the new compression level to the terminal device. When the terminal device acquires the new original audio signal, it compresses the new original audio signal into a new target audio signal according to the new compression level.

[0130] In addition, the cloud server can construct a conversation scenario based on voice characteristics and chat information, call the generative model to generate business information based on the current chat information in the conversation scenario, and transmit the business information to the terminal device to execute the corresponding business operation according to the business information.

[0131] Among them, business information can be reply information, used to respond to the speaker's chat information and form a dialogue, or it can be operation instructions, used to control the operation of terminal devices, such as playing songs, adjusting the color and / or brightness of lights, etc.

[0132] In this embodiment, the network status between the terminal device and the device is monitored. Upon receiving the target audio signal transmitted by the terminal device, the target audio signal is decompressed into the original audio signal according to the current compression level. The original audio signal is encoded into basic audio features. The timbre features of the original audio signal are extracted from the basic audio features. The basic audio features are decoded according to the current compression level to obtain chat information. A new compression level is generated based on the network status, chat information, and timbre features. The new compression level is transmitted to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level. This embodiment determines the compression level of the audio signal based on adaptive factors such as network status and dynamic scene of the conversation, and dynamically adjusts speech recognition according to the compression level. It seeks a balance between the resource consumption of the terminal device in transmitting audio signals and the accuracy of speech recognition, effectively improving the efficiency of audio signal transmission, maintaining the feedback efficiency to the terminal device, and maintaining the response efficiency of the terminal device to the user.

[0133] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0134] Reference Figure 5 The diagram illustrates an audio transmission device based on hierarchical compression according to an embodiment of the present invention, which may specifically include the following modules:

[0135] The network status monitoring module 501 is used to monitor the network status between the terminal device and the terminal device.

[0136] The audio signal decompression module 502 is used to decompress the target audio signal into the original audio signal according to the current compression level when receiving the target audio signal transmitted by the terminal device.

[0137] Audio encoding module 503 is used to encode the current original audio signal into basic audio features;

[0138] The timbre feature extraction module 504 is used to extract the timbre features of the current original audio signal from the basic audio features;

[0139] The audio decoding module 505 is used to decode the basic audio features according to the current compression level to obtain chat information;

[0140] Compression level generation module 506 is used to generate a new compression level based on the network status, the chat information and the timbre characteristics;

[0141] Compression level transmission module 507 is used to transmit a new compression level to the terminal device so as to compress the new original audio signal into a new target audio signal according to the new compression level.

[0142] In one embodiment of the present invention, the audio decoding module 505 is further configured to:

[0143] If the current compression level is in the first level range, the basic audio features are input into the decoder and decoded into low-pressure audio features;

[0144] The low-voltage audio features are input into the head structure to generate chat information;

[0145] If the current compression level belongs to the second level range, the basic audio features are input into the first compensation layer to compensate for the features lost during compression, and the first high-voltage audio features are obtained.

[0146] The first high-voltage audio feature is input into the decoder and decoded into the second high-voltage audio feature;

[0147] The second high-voltage audio feature is input into the second compensation layer to compensate for the features lost during compression, thus obtaining the third high-voltage audio feature.

[0148] The third high-voltage audio feature is input into the head structure to generate chat information;

[0149] The compression levels in the second level range are all higher than those in the first level range.

[0150] In one embodiment of the present invention, the audio decoding module 505 is further configured to:

[0151] Find standard audio features in uncompressed sample speech signals;

[0152] In the first compensation layer, the basic audio features are used as the query matrix and the standard audio features are used as the key matrix and value matrix. An attention mechanism is applied to generate standard attention weights. The basic audio features are multiplied by the standard attention weights to obtain the first candidate audio features. The first candidate audio features are then mapped to the first high-voltage audio features.

[0153] In one embodiment of the present invention, the audio decoding module 505 is further configured to:

[0154] Based on the timbre characteristics, find the lowest compression level of the low-pressure audio feature in history, and use it as the individual audio feature;

[0155] In the second compensation layer, the second high-voltage audio feature is used as the query matrix and the individual audio feature is used as the key matrix and value matrix. An attention mechanism is applied to generate object attention weights. The second high-voltage audio feature is multiplied by the object attention weights to obtain the second candidate audio feature. The second candidate audio feature is then mapped to the third high-voltage audio feature.

[0156] In one embodiment of the present invention, the compression level generation module 506 is further configured to:

[0157] Use the current network state to generate network features that represent the future network state;

[0158] The timbre features are concatenated with the text features of the chat messages to form the original dialogue features;

[0159] Based on the original dialogue features, target dialogue features representing future chat information are generated;

[0160] The network features and the target dialogue features are fused into multimodal features;

[0161] A new compression level is classified based on the aforementioned multimodal features.

[0162] In one embodiment of the present invention, the compression level generation module 506 is further configured to:

[0163] All the original dialogue features are concatenated into a global dialogue feature;

[0164] Based on the timbre features, the original dialogue features are clustered to obtain multiple clusters;

[0165] For the cluster with the largest number of original dialogue features, the original dialogue features in the cluster are concatenated into local dialogue features;

[0166] The global dialogue features and the local dialogue features are concatenated to form enhanced dialogue features;

[0167] The enhanced dialogue features are used to generate target dialogue features representing future chat information.

[0168] In one embodiment of the present invention, the compression level generation module 506 is further configured to:

[0169] The network features and the target dialogue features are concatenated to form a fused feature;

[0170] Using the fused features as the query matrix, key matrix, and value matrix, an attention mechanism is applied to generate self-attention weights. The fused features are then multiplied by the self-attention weights to obtain multimodal features.

[0171] In one embodiment of the present invention, it further includes:

[0172] The business information generation module is used to generate business information based on the timbre characteristics and the chat information;

[0173] The business information transmission module is used to transmit the business information to the terminal device so as to perform business operations according to the business information.

[0174] The present invention provides an audio transmission device based on hierarchical compression. By applying the audio transmission device based on hierarchical compression, the steps in the aforementioned audio transmission method embodiments based on hierarchical compression can be implemented.

[0175] It should be noted that the module division in the various audio transmission devices based on hierarchical compression provided in the above embodiments is illustrative and only represents a logical functional division. In actual implementation, other division methods may also be used. Furthermore, the functional modules in the various embodiments of this invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0176] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of the embodiments of the present invention can be embodied in the form of a computer program product, which is stored in a computer storage medium and includes several instructions to cause an electronic device or processor to execute all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned computer storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0177] Furthermore, the audio transmission device based on hierarchical compression and the audio transmission method based on hierarchical compression provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0178] Reference Figure 6 The diagram illustrates an electronic device according to an embodiment of the present invention. Figure 6 As shown, the electronic device in this embodiment of the invention includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the above-described embodiment of the audio transmission method based on hierarchical compression. Alternatively, when the processor executes the computer program, it implements the functions of each module in the above-described embodiment of the audio transmission device based on hierarchical compression.

[0179] For example, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which can be used to describe the execution process of the computer program in the electronic device.

[0180] The electronic device may be a desktop computer, a cloud server, or other computing device. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 6 This is merely one example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0181] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0182] The memory can be an internal storage unit of the electronic device, such as a hard drive or RAM. Alternatively, it can be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory can include both internal and external storage units. The memory is used to store the computer program and other programs and data required by the electronic device. The memory can also be used to temporarily store data that has been output or will be output.

[0183] This invention also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the audio transmission method based on hierarchical compression as described in the foregoing embodiments.

[0184] This invention also discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio transmission method based on hierarchical compression as described in the foregoing embodiments.

[0185] This invention also discloses a computer program product that, when run on a computer, causes the computer to execute the audio transmission method based on hierarchical compression described in the foregoing embodiments.

[0186] The embodiments described above are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An audio transmission method based on hierarchical compression, characterized in that, include: Monitor the network status between the device and the terminal equipment; When receiving the target audio signal transmitted by the terminal device, the target audio signal is decompressed into the original audio signal according to the current compression level; Encode the original audio signal into basic audio features; Extract the timbre features of the current original audio signal from the basic audio features; If the current compression level is in the first level range, the basic audio features are input into the decoder and decoded into low-pressure audio features; The low-voltage audio features are input into the head structure to generate chat information; If the current compression level is in the second level range, then search for the standard audio features of the uncompressed sample speech signal; wherein, the compression levels in the second level range are all higher than the compression levels in the first level range; In the first compensation layer, the basic audio features are used as the query matrix and the standard audio features are used as the key matrix and value matrix. An attention mechanism is applied to generate standard attention weights. The basic audio features are multiplied by the standard attention weights to obtain the first candidate audio features. The first candidate audio features are then mapped to the first high-voltage audio features. The first high-voltage audio feature is input into the decoder and decoded into the second high-voltage audio feature; Based on the timbre characteristics, find the lowest compression level of the low-pressure audio feature in history, and use it as the individual audio feature; In the second compensation layer, the second high-voltage audio feature is used as the query matrix and the individual audio feature is used as the key matrix and value matrix. An attention mechanism is applied to generate object attention weights. The second high-voltage audio feature is multiplied by the object attention weights to obtain the second candidate audio feature. The second candidate audio feature is then mapped to the third high-voltage audio feature. The third high-voltage audio feature is input into the head structure to generate chat information; the timbre feature and the chat information are used to construct a conversation scenario; A new compression level is generated based on the network status, the chat information, and the timbre characteristics; The new compression level is transmitted to the terminal device to compress the new original audio signal into a new target audio signal according to the new compression level.

2. The method according to claim 1, characterized in that, The process of generating a new compression level based on the network status, the chat information, and the timbre features includes: Use the current network state to generate network features that represent the future network state; The timbre features are concatenated with the text features of the chat messages to form the original dialogue features; Based on the original dialogue features, target dialogue features representing future chat information are generated; The network features and the target dialogue features are fused into multimodal features; A new compression level is classified based on the aforementioned multimodal features.

3. The method according to claim 2, characterized in that, The step of generating target dialogue features representing future chat information based on the original dialogue features includes: All the original dialogue features are concatenated into a global dialogue feature; Based on the timbre features, the original dialogue features are clustered to obtain multiple clusters; For the cluster with the largest number of original dialogue features, the original dialogue features in the cluster are concatenated into local dialogue features; The global dialogue features and the local dialogue features are concatenated to form enhanced dialogue features; The enhanced dialogue features are used to generate target dialogue features representing future chat information.

4. The method according to claim 2, characterized in that, The process of fusing the network features and the target dialogue features into multimodal features includes: The network features and the target dialogue features are concatenated to form a fused feature; Using the fused features as the query matrix, key matrix, and value matrix, an attention mechanism is applied to generate self-attention weights. The fused features are then multiplied by the self-attention weights to obtain multimodal features.

5. The method according to any one of claims 1-4, characterized in that, Also includes: Business information is generated based on the timbre characteristics and the chat information; The service information is transmitted to the terminal device to perform service operations according to the service information.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio transmission method based on hierarchical compression as described in any one of claims 1-5.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the audio transmission method based on hierarchical compression as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Data transmission method and device, terminal, storage medium and system

    CN111314335A

  • Cloud edge cooperative control method

    CN119561945A