Audio encoding and decoding method and system
Through time domain and frequency domain analysis combined with neural network models for dynamic parameter calculation, the problem that existing Bluetooth headset audio processing solutions cannot be dynamically adjusted is solved, and the dual optimization of audio processing accuracy and processing time is realized, providing a personalized auditory experience and reducing system power consumption.
Patent Information
- Application Number
- CN202510305914.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Bluetooth headset audio processing solutions cannot be dynamically adjusted to adapt to the listening habits and audio characteristics of different users, making it difficult for the sound quality experience to meet personalized needs. At the same time, the calculation burden is too heavy when processing high-quality audio streams, which is prone to audio data distortion and delay.
Through dual analysis of time domain and frequency domain, combined with user listening habits, an audio feature matrix is generated, and dynamic parameter calculation is performed through neural network models to generate an adaptive coding matrix. Based on the hierarchical processing architecture, the audio input signals are divided layer by layer, and encoded by heterogeneous computing units to realize efficient compression and quality control of audio data.
It realizes dual optimization of audio processing accuracy and processing time, provides a personalized auditory experience, ensures high-quality audio output and real-time processing capabilities, and significantly reduces the system's memory footprint and overall power consumption, and optimizes data transmission efficiency.
Smart Images

Figure CN120089147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio encoding and decoding, and particularly to an audio encoding and decoding method and system. Background Art
[0002] With the popularization of Bluetooth audio devices, users' requirements for audio quality and auditory experience are constantly increasing. However, the existing audio processing solutions for Bluetooth headsets face multiple key challenges. Traditional fixed-parameter audio encoding and decoding schemes cannot dynamically adjust according to different users' listening habits and audio characteristics, resulting in the difficulty of meeting personalized needs in terms of sound quality experience. At the same time, as a portable device, the processing power and battery life of Bluetooth headsets are strictly limited, and they often face the problem of excessive computational burden when processing high-quality audio streams.
[0003] Current audio processing technologies need to consider multiple constraints such as transmission bandwidth, computational power, and environmental noise while achieving high-quality sound reproduction. Existing single encoding and decoding schemes are prone to audio data distortion and delay during Bluetooth transmission, especially in complex environments, it is difficult to ensure stable sound output quality. Summary of the Invention
[0004] The present invention provides an audio encoding and decoding method and system for achieving dual optimization of audio processing accuracy and processing time.
[0005] In a first aspect, the present invention provides an audio encoding and decoding method, and the audio encoding and decoding method includes: Performing time-domain analysis and frequency-domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; Performing feature extraction and fusion processing on the audio feature matrix and user audio operation data to generate a target feature vector including frequency band energy distribution and user preference information; Inputting the target feature vector into a preset neural network model for dynamic parameter calculation to generate an adaptive encoding matrix; According to the adaptive encoding matrix, hierarchically partitioning the audio input signal to obtain basic sound quality layer data and enhanced sound quality layer data; Performing encoding processing on the basic sound quality layer data and the enhanced sound quality layer data respectively through a heterogeneous computing unit, performing lossless encoding of the basic sound quality layer data by the main processor, and performing lossy encoding of the enhanced sound quality layer data by the digital signal processor to generate a target encoded data stream.
[0006] In a second aspect, the present invention provides an audio encoding and decoding system, and the audio encoding and decoding system includes: An analysis module for performing time-domain analysis and frequency-domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; A fusion module, configured to perform feature extraction and fusion processing on the audio feature matrix and user audio operation data, and generate a target feature vector including band energy distribution and user preference information; A calculation module, configured to input the target feature vector into a preset neural network model to perform dynamic parameter calculation, and generate an adaptive coding matrix; A partitioning module, configured to perform hierarchical partitioning on the audio input signal according to the adaptive coding matrix, and obtain basic audio quality layer data and enhanced audio quality layer data; An encoding module, configured to perform encoding processing on the basic audio quality layer data and the enhanced audio quality layer data respectively through a heterogeneous computing unit, perform lossless encoding of the basic audio quality layer data by a main processor, and perform lossy encoding of the enhanced audio quality layer data by a digital signal processor, and generate a target encoded data stream.
[0007] A third aspect of the present invention provides an audio codec device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to enable the audio codec device to execute the above audio codec method.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is enabled to execute the above audio codec method.
[0009] In the technical solution provided by the present invention, through dual analysis in the time domain and frequency domain combined with the user's listening habits, accurate extraction of audio features is achieved, and a neural network model is used for dynamic parameter calculation, so that the coding parameters can be adaptively adjusted. Based on a hierarchical processing architecture, efficient compression and quality control of audio data are realized. This method makes full use of the respective advantages of the main processor and the digital signal processor through the collaborative work of the heterogeneous computing unit, and adopts a task priority scheduling mechanism to achieve efficient allocation of processing resources, and can dynamically adjust the processing strategy to control system power consumption. In addition, the present invention improves the robustness of the decoding process, enhances the system's ability to process data errors, and realizes double optimization of audio processing accuracy and processing time. In terms of user experience, this method provides a personalized auditory experience through adaptive parameter adjustment, ensuring high-quality audio output and real-time processing capabilities. At the same time, through hierarchical coding technology and dynamic parameter adaptive methods, the system memory occupancy and overall power consumption are significantly reduced, the data transmission efficiency is optimized, and the conservation and utilization of system resources are realized.
[0010] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures particularly pointed out in the specification, claims, and drawings.
[0011] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Brief Description of the Drawings
[0012] Figure 1 It is a schematic diagram of an embodiment of the audio encoding and decoding method in an embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of the audio encoding and decoding system in an embodiment of the present invention; Figure 3 It is a schematic diagram of an embodiment of the audio encoding and decoding device in an embodiment of the present invention. Detailed Embodiments
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0014] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0015] For ease of understanding of this embodiment, first, a detailed introduction is given to an audio encoding and decoding method disclosed in the embodiments of the present invention. As Figure 1 shown, the method includes the following steps: 101. Perform time-domain analysis and frequency-domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; It can be understood that the execution subject of the present invention can be an audio encoding and decoding system, or a terminal or a server. Specifically, no limitation is made here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.
[0016] Specifically, the audio input signal is framed according to a preset time window length, dividing the continuous audio signal into multiple small segments, and each small segment is regarded as an independent time frame. During the framing process, the length of the time window is optimally selected according to the application scenario and the characteristics of the audio signal, choosing a parameter setting that can balance the frequency-domain resolution and the time-domain resolution, such as a window length of 20 milliseconds or 25 milliseconds. For each time frame, a Hamming window weighting process is applied to reduce the spectral leakage problem introduced during the signal framing process. The Hamming window is a weighting function that effectively reduces the discontinuity generated in the frequency-domain analysis after framing by smoothing the signals at both ends of the time frame. The fast Fourier transform is performed on the weighted time frame sequence to transform the signal from the time domain to the frequency domain, extracting the characteristics of the signal in the frequency dimension, that is, obtaining the frequency-domain coefficient sequence of the signal through the fast Fourier transform, and these coefficients reflect the intensity distribution of different frequency components in the signal. Nonlinear frequency band division is performed on the frequency-domain coefficient sequence. Based on the human ear's auditory characteristics, the frequency domain is divided into multiple critical bands, and these critical bands are divided according to the variation law of the human ear's frequency resolution ability. The human ear has a higher frequency resolution ability for the low-frequency part and a lower resolution ability for the high-frequency part. Therefore, the critical bands are narrower in the low-frequency region and gradually widen in the high-frequency region. This division process is achieved by using tools such as a Mel filter bank. After the division is completed, the frequency-domain coefficient sequence is mapped to these critical bands to form a critical band sequence. The energy of each band in the critical band sequence is calculated, and the energy value is calculated by squaring and summing, reflecting the intensity of the signal within that band. The calculated band energy values are arranged in a matrix according to the indices of the time frames and the indices of the bands, forming an initial matrix of band energy. This matrix records the energy distribution of each critical band at different time frames in the form of a two-dimensional structure. To adapt to the changes in different signals and environments and improve the robustness of the audio features, the initial matrix of band energy is normalized. The normalization process eliminates the absolute differences between different signal intensities, making the subsequent feature extraction more universal, and obtaining a normalized matrix of band energy. The normalized matrix of band energy is convolved with a preset auditory weighting coefficient. The convolution process applies a weighting factor to different bands, making it closer to the actual situation of the human ear's perception of sound. The auditory weighting coefficient is based on a specific auditory model, such as an equal-loudness curve or a Bark frequency scale, and these models can effectively improve the matching degree of the audio features to the human ear's perception. The final audio feature matrix is obtained after the convolution operation.
[0017] 102. Feature extraction and fusion processing are performed on the audio feature matrix and the user's audio operation data to generate a target feature vector containing the band energy distribution and user preference information; Specifically, perform a temporal continuity analysis on the band energy values in the audio feature matrix to mine the dynamic change laws of the audio signal. By observing the change trend of the band energy values in the time dimension, generate a band energy change sequence, which reflects the distribution characteristics of the energy in the audio signal over time. At the same time, collect data on the user's volume adjustment operations, equalizer setting parameters, and ambient noise adaptation parameters to generate a user operation data sequence. The user's volume adjustment operations reflect their volume requirements in different usage scenarios, the equalizer setting parameters reflect the user's preferences for enhancing or attenuating specific frequencies, and the ambient noise adaptation parameters show the noise characteristics of the user's environment and the device's response strategy to noise. By continuously collecting and recording these data, construct a user operation data sequence in the form of a time series. Conduct a statistical analysis of the user operation data sequence to extract the user's common volume level and preferred frequency response curve. The common volume level is obtained by calculating the mean or median of the volume adjustment operations, while the preferred frequency response curve is generated by weighted statistics on the frequency distribution of the equalizer setting parameters. These statistical results form the core part of the user preference feature data, which can reflect the user's personalized needs in volume and frequency adjustment. Align the band energy change sequence with the user preference feature data to ensure their consistency in the time and feature dimensions. Use interpolation methods or dynamic time warping techniques to synchronize the data to generate an aligned feature sequence. Standardize the aligned feature sequence to eliminate the dimensional differences between different feature dimensions, improve the comparability between features, and obtain standardized feature data. Perform feature selection on the standardized feature data, and screen out key feature dimensions by calculating the information gain value to obtain a reduced-dimensional feature subset. The information gain value measures the importance of a certain feature for the target variable, and the feature selection process can remove redundant or irrelevant feature dimensions, reduce the data dimension, and improve the efficiency of the model. Perform a non-linear transformation on the reduced-dimensional feature subset through deep neural networks, kernel function transformations, or other mapping methods to further mine complex feature relationships on the basis of the reduced-dimensional feature subset and generate a target feature vector with higher expression ability.
[0018] 103. Input the target feature vector into a preset neural network model for dynamic parameter calculation to generate an adaptive coding matrix; Specifically, the target feature vector is decomposed into a frequency band feature sub-vector and a user feature sub-vector. The target feature vector is a high-dimensional feature representation that comprehensively represents the frequency band energy distribution characteristics of the audio signal and the user preference information. By logically partitioning the target feature vector, the frequency band features and user features are processed independently to extract the key information of both with higher pertinence. The frequency band feature sub-vector is input into the feature extraction network in the neural network model. The feature extraction network consists of three convolutional layers and two fully connected layers. Each convolutional layer captures the local receptive field of the frequency band features through convolutional operations and extracts their local correlations in the time and frequency dimensions. After each convolutional operation, the ReLU activation function is applied to introduce non-linear expressive ability, and the data distribution is standardized through the batch normalization layer to ensure the stability and fast convergence of the network during training. After three convolutional operations, the frequency band feature sub-vector is mapped into a high-dimensional frequency band feature mapping matrix. At the same time, the user feature sub-vector in the target feature vector is input into the user modeling network in the neural network model. The user modeling network consists of four fully connected layers. The first and second fully connected layers use the ReLU activation function, and these layers perform preliminary non-linear expansion and feature capture on the user features through non-linear transformations. The third fully connected layer uses the Sigmoid activation function to refine the non-linear relationship of the user features and limit the feature range to [0, 1], which is suitable for capturing probabilistic or weight features. The fourth fully connected layer uses the Tanh activation function to enhance the expressive ability of the features and adjust the feature range to [-1, 1] to ensure that the output can contain negatively correlated information. After four fully connected operations, the user feature sub-vector is mapped into a user feature mapping matrix. Attention feature fusion is performed on the frequency band feature mapping matrix and the user feature mapping matrix. The attention mechanism selectively strengthens the key information by calculating the importance weights of each feature element, while suppressing irrelevant or redundant information. The similarity weight matrix between the frequency band feature mapping matrix and the user feature mapping matrix is calculated, and then each element in the frequency band feature mapping matrix is weighted and summed to generate a fused feature tensor. The fused feature tensor is input into the parameter generation network in the neural network model. The parameter generation network adopts a multi-layer perceptron structure and processes the fused feature tensor through a series of fully connected layers to output an encoded parameter vector. The transposed convolution operation is performed on the encoded parameter vector. The transposed convolution operation remaps the low-dimensional encoded parameter vector into a high-dimensional adaptive encoding matrix through an inverse convolution process.
[0019] 104. According to the adaptive encoding matrix, the audio input signal is hierarchically partitioned to obtain the basic sound quality layer data and the enhanced sound quality layer data; Specifically, based on the bit allocation parameters in the adaptive coding matrix, the audio input signal is divided into frequency bands. The adaptive coding matrix contains bit allocation parameters dynamically calculated by a neural network, and these parameters determine the priority of each frequency band and the amount of resources allocated. Through wavelet transform or filter bank technology, the audio signal is divided into multiple groups of frequency band signals, and each frequency band signal represents the audio characteristics within a specific frequency range. After the division is completed, the energy of each frequency band signal in the multi-frequency band signal group is calculated and the spectral density is analyzed. The energy calculation evaluates the intensity of the signal by summing the squares of the signal, while the spectral density analysis quantifies the spectral structure of the signal through the distribution characteristics in the frequency domain. These analysis results generate frequency band characteristic parameters. Based on the frequency band characteristic parameters and the human ear auditory masking threshold, the perceptual importance of the multi-frequency band signal group is evaluated. The auditory masking model is based on the perceptual characteristics of the human ear, where strong signals mask weak signals, especially when the frequencies are close. By comparing the frequency band characteristic parameters with the masking threshold, the contribution of each frequency band to the overall sound quality is calculated, and this contribution is quantified as a frequency band weight sequence. The higher the weight value of a frequency band, the stronger its perceptual importance to the audio signal. The maximum entropy threshold segmentation is performed on the frequency band weight sequence to achieve hierarchical division of the frequency bands. The maximum entropy threshold segmentation is a segmentation method based on information theory. By finding a threshold that maximizes the distribution entropy value of the frequency band weights after segmentation, it ensures that the segmentation result can fully reflect the importance differences of the frequency bands. After the segmentation is completed, the frequency bands with weight values higher than the threshold are included in the critical frequency band set, and the frequency bands with weight values lower than the threshold are included in the enhancement frequency band set, forming the frequency band hierarchical result. According to the frequency band hierarchical result, the multi-frequency band signal group is reorganized. For the signals in the critical frequency band set, they are reconstructed into time-domain signals through inverse wavelet transform to obtain the basic sound quality layer data. The basic sound quality layer data retains the main content and the perceptual key parts of the audio signal, ensuring that even under bandwidth constraints, the transmitted audio signal can still meet the basic requirements of hearing. At the same time, for the signals in the enhancement frequency band set, they are organized into a hierarchical structure through wavelet coefficient rearrangement and coefficient matrix transformation to obtain the enhanced sound quality layer data. The wavelet coefficient rearrangement sorts and adjusts the enhancement frequency band signals according to their importance, enabling dynamic adjustment of the data allocation and storage efficiency at different resolutions.
[0020] 105. The basic sound quality layer data and the enhanced sound quality layer data are respectively encoded by heterogeneous computing units. The main processor performs lossless encoding of the basic sound quality layer data, and the digital signal processor performs lossy encoding of the enhanced sound quality layer data to generate the target encoded data stream.
[0021] Specifically, the basic sound quality layer data is input into the lossless coding module of the main processor. The goal of the lossless coding module is to compress the basic sound quality layer data with the highest fidelity to ensure that its core audio information is fully restored without distortion after transmission and decoding. The main processor performs integer transform and entropy coding operations on the basic sound quality layer data. Integer transform is a linear transform that can reduce data redundancy. By converting time domain or frequency domain data into integer form, the complexity of storage and processing is reduced. Entropy coding uses the probability of occurrence of different symbols in the source to efficiently compress data through methods such as Huffman coding or arithmetic coding to generate a basic layer lossless code stream. The basic layer lossless code stream is grouped. Through a grouping operation based on audio frame boundary information, the basic layer lossless code stream is divided into multiple coding units to form a basic layer coding unit sequence. The purpose of the grouping process is to adapt to the framing requirements of Bluetooth transmission on the basis of ensuring the continuity of the code stream. At the same time, the enhanced sound quality layer data is input into the quantization coding module of the digital signal processor, and the data is nonlinearly quantized according to the dynamic quantization parameter. The setting of dynamic quantization parameters is dynamically adjusted based on signal characteristics and user needs. The nonlinear quantization process effectively reduces the amount of data by mapping the signal value to a limited number of quantization levels in segments, while retaining the perceived sound quality details as much as possible. The quantized data is organized into a quantization coefficient matrix. The quantization coefficient matrix is entropy coded and run-length coded. Entropy coding further compresses the remaining redundant information in the quantized data, while run-length coding uses the characteristics of continuous repeated values in the quantized data to achieve a higher compression ratio by recording the number of repeated values. Combining the results of entropy coding and run-length coding, the enhanced layer coded data is generated according to the data redundancy and quantization accuracy. A multi-layer coded data packet is established based on the base layer coding unit sequence and the enhanced layer coded data to form an audio data representation with a clear hierarchical structure. In the process of constructing the multi-layer coded data packet, data check and synchronization mark insertion operations are performed. By adding checksums and synchronization information to the header of the data packet, the reliability of the data packet and the synchronization of the transmission are improved, thereby reducing the bit error rate and data loss risk in Bluetooth transmission, and obtaining a coded data stream with checksum. In order to adapt to the processing load and power status of the Bluetooth chip, the transmission of the coded data stream with checksum is optimized. By adding layer identification and priority information to the data packet, the transmission order and strategy of the base layer and enhancement layer data can be flexibly adjusted to obtain the target coded data stream. As the core part of the audio, the base layer data is set to a higher priority to ensure that users can still receive complete basic sound quality information even in the case of insufficient bandwidth or transmission interruption. The enhancement layer data is transmitted as an additional part when bandwidth permits, thereby providing additional enhanced details for the audio experience.
[0022] Perform packet parsing on the target encoded data stream to isolate the multi-layer decoded data contained therein. The packet parsing is based on the structure of the multi-layer encoded packets. By analyzing the layer identification and synchronization information in the packet header, the independent bitstreams of the base layer and the enhancement layer are extracted. Perform state modeling on the multi-layer decoded data. The decoding state modeling module maps key factors such as signal-to-noise ratio, frequency distortion degree, and quantization error distribution to a belief state vector by analyzing the current decoding state space. The signal-to-noise ratio reflects the quality of packet transmission. The frequency distortion degree measures the distortion of the signal in the frequency domain during the decoding process. The quantization error distribution reveals the possible deviation of the quantization coefficients in the enhancement layer during reconstruction. These state parameters are combined in the form of a probability model and mapped to a belief vector representing the current decoding environment and state, that is, the initial decoding belief state. Perform Markov chain prediction on the initial decoding belief state. The Markov chain model predicts the state distribution at the next moment through time series analysis of the decoding state and corrects the initial belief state based on historical state information. This prediction process makes full use of the dynamic change characteristics of the decoding state, so that the generated corrected belief state can more accurately reflect the changes in the current decoding environment. After the correction is completed, the corrected belief state is mapped to an optimized decoding parameter set, which contains the core parameters that need to be dynamically adjusted during the decoding process, such as the probability distribution of entropy decoding, the non-linear coefficient of quantization inverse mapping, and the weight adjustment of the frequency-domain signal. Perform entropy decoding on the base layer bitstream in the multi-layer decoded data according to the optimized decoding parameter set. Through reverse Huffman decoding and run-length decoding, the compressed data is restored to an uncompressed frequency-domain signal representation. After the entropy decoding is completed, perform an integer inverse transform on the data to restore the integer transform operation performed during the encoding process, and convert the frequency-domain signal into a time-domain base audio signal to obtain the base layer decoded data. At the same time, for the enhancement layer bitstream, the decoding process mainly focuses on inverse quantization and the conversion from the frequency domain to the time domain. According to the optimized decoding parameter set, use a non-linear quantization inverse mapping function to perform inverse quantization on the quantization coefficients of the enhancement layer and restore them to a frequency-domain representation closer to the original signal. The inverse mapping function combines the non-linear characteristics during the quantization process and the adjustment parameters of the current decoding state to minimize the impact of quantization error on the reconstructed signal. After the inverse quantization is completed, use wavelet inverse transform to reconstruct the frequency-domain signal into a time-domain enhancement signal to generate the enhancement layer decoded data. Perform time-domain superposition synthesis on the base layer decoded data and the enhancement layer decoded data to generate an audio output signal. By performing sample-by-sample superposition of the two layers of signals, the detailed information of the enhancement layer is incorporated into the base layer signal to form a more complete and delicate audio signal. During the superposition process, for problems such as phase error or amplitude mismatch, perform dynamic correction according to the adjustment coefficients in the optimized decoding parameter set to ensure the consistency of the synthesized signal in time and amplitude.
[0023] In the embodiments of the present invention, through dual analysis in the time domain and frequency domain combined with the user's listening habits, the accurate extraction of audio features is achieved, and a neural network model is used to calculate dynamic parameters, enabling the coding parameters to be adaptively adjusted. Based on a hierarchical processing architecture, efficient compression and quality control of audio data are realized. This method, through the collaborative work of heterogeneous computing units, makes full use of the respective advantages of the main processor and the digital signal processor, adopts a task priority scheduling mechanism to achieve efficient allocation of processing resources, and can dynamically adjust the processing strategy to control system power consumption. In addition, the present invention introduces a modified belief Markov decision process to enhance the robustness of the decoding process, and enhances the system's ability to handle data errors through a multi-agent actor-attention-critic algorithm, achieving a dual optimization of audio processing accuracy and processing time. In terms of user experience, this method provides a personalized auditory experience through adaptive parameter adjustment, ensuring high-quality audio output and real-time processing capabilities. At the same time, through hierarchical coding technology and dynamic parameter adaptation methods, the system's memory occupancy and overall power consumption are significantly reduced, the data transmission efficiency is optimized, and the conservation and utilization of system resources are realized.
[0024] In a specific embodiment, the process of executing step 101 may specifically include the following steps: Frame the audio input signal according to a preset time window length to obtain multiple time frames, and perform Hamming window weighting on each time frame to obtain a weighted time frame sequence; Perform a fast Fourier transform on the weighted time frame sequence to obtain a frequency domain coefficient sequence, and perform non-linear frequency band division on the frequency domain coefficient sequence. Based on the human ear auditory characteristics, divide the frequency band into multiple critical frequency bands to obtain a critical frequency band sequence; Calculate the energy of each frequency band in the critical frequency band sequence to obtain the frequency band energy value, and arrange the frequency band energy values in a matrix according to the time frame and frequency band index to obtain the initial frequency band energy matrix; Perform normalization processing on the initial frequency band energy matrix to obtain a normalized frequency band energy matrix, and perform convolution operation on the normalized frequency band energy matrix with a preset auditory weighting coefficient to obtain an audio feature matrix.
[0025] Specifically, frame the audio input signal according to a preset time window length. Assume the audio signal is , where represents the discrete time index, and the entire signal is divided into multiple overlapping time frames. The length of each frame is , and the frame shift is . Through framing, the audio signal is decomposed into multiple frames , defined as: ; where Indicates the index of the frame. Assume that sampling points, and the frame shift is , then the overlapping part of adjacent frames is sampling points. Overlapping can reduce the spectral distortion introduced during the frame segmentation process. After frame segmentation, each time frame is weighted with a Hamming window. The Hamming window has the following formula: ; The weighted time frame signal is: ; Through this step, the discontinuity at the frame boundary is reduced, thereby reducing spectral leakage. The fast Fourier transform is performed on the weighted time frame sequence to convert the time-domain signal to the frequency domain. The mathematical definition of the fast Fourier transform is: ; where represents the frequency-domain coefficient of the th frame, is the frequency index, is the imaginary unit. Through the fast Fourier transform, the spectral information of the signal is obtained. For example, for a sine signal , its frequency-domain coefficient will present a peak at the corresponding frequency . Nonlinear frequency band division is performed on the frequency-domain coefficient sequence . Based on the human auditory characteristics, the frequency is divided into multiple critical bands. The division of critical bands uses the Mel frequency scale, and its formula is: ; After converting the linear frequency to the Mel frequency, the frequency band is divided by designing a filter bank. Assume that 20 Mel filters are used, and the frequency range of each filter is determined by the center frequency and bandwidth. Then the output of each filter is: ; where is the energy of the th frame in the th frequency band, is the frequency response of the th filter. After calculating the energy of each frequency band in the critical band sequence, the frequency band energy value matrix is obtained, where represents the time frame index, represents the frequency band index. These energy values reflect the distribution characteristics of the audio signal in the time-frequency plane. To make the frequency band energy matrix adapt to different signals and environmental conditions, it is normalized. The normalization uses the following formula: ; Normalized band energy matrix The value range is restricted between [0, 1], thereby improving the stability and adaptability of the features. The normalized band energy matrix and the preset auditory weighting coefficients are subjected to a convolution operation to enhance the perceptual relevance of the features, obtaining an audio feature matrix. The convolution formula is: ; where represents the convolved audio feature matrix, is the weighting coefficient, which is used to simulate the sensitivity of the human ear to different frequencies.
[0026] In a specific embodiment, the process of executing step 102 may specifically include the following steps: Perform a temporal continuity analysis on the band energy values in the audio feature matrix to obtain a band energy change sequence, and collect data on the user's volume adjustment operations, equalizer setting parameters, and environmental noise adaptation parameters to obtain a user operation data sequence; According to the user operation data sequence, perform data statistics on the user's common volume level and preferred frequency response curve to obtain user preference feature data; Align the band energy change sequence and the user preference feature data to obtain an aligned feature sequence, and perform data normalization processing on the aligned feature sequence to obtain normalized feature data; Perform feature selection on the normalized feature data, screen the key feature dimensions by calculating the information gain value to obtain a dimensionality-reduced feature subset, and perform a non-linear transformation on the dimensionality-reduced feature subset to obtain a target feature vector.
[0027] Specifically, assume that the audio feature matrix is represented as , where is the time frame index, is the band index, and each element reflects the energy value of a certain time frame in a certain band. To capture the characteristics of the audio signal changing over time, a temporal continuity analysis is performed on it to calculate the change amount between adjacent time frames. The change amount is expressed by the first-order difference formula as: ; The band energy change sequence generated by this process , and represent the user's volume adjustment, the gain setting of the equalizer for each frequency band, and the noise level of the current environment respectively. Combining these data, a user operation data sequence is constructed : ; This step reflects how user operations affect the processing of audio signals. For example, when the user increases the gain of a certain frequency band, the corresponding value will increase, thus affecting subsequent analysis. After obtaining the user operation data sequence, the user's common volume level and preferred frequency response curve are extracted through statistical analysis. The user's common volume level is expressed as: ; where represents the total number of time frames. The preferred frequency response curve of the equalizer is calculated by taking the average of the gains for each frequency band: ; For example, if the user tends to frequently increase the gain in certain frequency bands, then the values of these frequency bands will be higher, reflecting the user's preference characteristics. Align the frequency band energy change sequence and the user preference characteristic data Perform an alignment operation. Synchronize the features in the time and frequency dimensions to generate an aligned feature sequence . The alignment formula is defined as: ; In this formula, is a scaling factor used to introduce the influence of the user's equalizer gain preference into the dynamic change of the frequency band energy. Normalize the aligned feature sequence to eliminate the influence of different feature dimensions. The normalization formula is: ; The result of normalization is to adjust the feature values to a fixed range, such as [0,1], thereby improving the stability of subsequent analysis. By calculating the information gain value screen the key feature dimensions to generate a reduced-dimensional feature subset. Information gain measures the importance of a certain feature to the target variable , and its formula is: ; where represents the entropy of the target variable , represents the entropy given the feature The conditional entropy of the target variable at that time. By retaining the feature with the largest information gain value, the data dimension is reduced and the expressive ability of the feature is improved. A non-linear transformation is performed on the reduced-dimensional feature subset to generate the target feature vector. The non-linear transformation is completed through an activation function, such as the Tanh function: ; This transformation maps the features to the range of [-1, 1], while enhancing the model's ability to capture non-linear feature relationships.
[0028] In a specific embodiment, the process of executing step 103 may specifically include the following steps: Input the band feature sub-vector in the target feature vector into the feature extraction network in the neural network model. The feature extraction network includes three convolutional layers and two fully connected layers. Each convolutional layer uses the ReLU activation function and the batch normalization layer. The local correlation of the band features is extracted through convolutional operations to obtain the band feature mapping matrix; Input the user feature sub-vector in the target feature vector into the user modeling network in the neural network model. The user modeling network includes four fully connected layers. The first fully connected layer and the second fully connected layer use the ReLU activation function, the third fully connected layer uses the Sigmoid activation function, and the fourth fully connected layer uses the Tanh activation function. The user preference features are captured through non-linear transformation to obtain the user feature mapping matrix; Perform attention feature fusion on the band feature mapping matrix and the user feature mapping matrix to obtain the fused feature tensor; Input the fused feature tensor into the parameter generation network in the neural network model. The parameter generation network adopts a multi-layer perceptron structure to obtain the encoded parameter vector, and perform a transposed convolution operation on the encoded parameter vector to obtain the adaptive encoding matrix.
[0029] Specifically, decompose the target feature vector to obtain the band feature sub-vector and the user feature sub-vector . Among them, represents the part related to the frequency band energy distribution of the audio signal, with a dimension of , while represents the part related to the user preference characteristics, with a dimension of . Input the band feature sub-vector into the feature extraction network, which includes three convolutional layers and two fully connected layers. The local correlation of the band features is extracted through convolution to generate a high-order feature representation. In the first convolution, apply a one-dimensional convolution operation to , with a convolution kernel size of , a stride of , and the number of filters of 。The mathematical expression of convolution is as follows: ; Where, and are the weights and biases of the first layer of convolution, BN represents batch normalization, * represents the convolution operation, and ReLU is the activation function. The first layer of convolution extracts the local patterns of the band features and stabilizes the training process through batch normalization. The second and third layers of convolution use similar formulas in sequence: ; ; After three layers of convolution, the obtained feature map matrix captures the local correlations of the band features. The features are further compressed and refined through two fully connected layers, and the formula is: ; ; Where, and are the weight matrices of the fully connected layers, and are the bias vectors. The generated band feature map matrix provides a high-dimensional description of the band characteristics. At the same time, the user feature sub-vector is input into the user modeling network, which contains four fully connected layers. The first and second fully connected layers use the ReLU activation function to capture the linear and non-linear relationships of user preferences: ; ; The third fully connected layer uses the Sigmoid activation function to extract the probabilistic features of user preferences: ; where Sigmoid . The fourth fully connected layer uses the Tanh activation function to generate a high-order representation of user features: ; The generated user feature map matrix is a non-linear representation of user preference features. The band feature map matrix and the user feature map matrix are fused through the attention mechanism to generate the fused feature tensor . The attention mechanism captures the correlation between the two features by calculating the weight matrix : ; Among them, Softmax is a normalization operation that maps the values in the weight matrix to the range of [0, 1]. The calculation formula for the fused feature tensor is as follows: ; The fused feature tensor combines the mutual information of the frequency band and user features. The fused feature tensor is input into a parameter generation network, which adopts a multi-layer perceptron structure. Through multi-layer non-linear transformation, an encoded parameter vector is generated : ; The multi-layer perceptron structure includes several fully connected layers, and the activation function of each layer selects ReLU or Tanh. For the encoded parameter vector transpose convolution operation is performed to generate an adaptive encoding matrix . The formula for transpose convolution is: ; Among them, is the transpose convolution kernel. The adaptive encoding matrix is a high-dimensional parameter matrix used to dynamically adjust the parameters of the encoding process.
[0030] In a specific embodiment, the process of executing step 104 may specifically include the following steps: According to the bit allocation parameters in the adaptive encoding matrix, the audio input signal is divided into frequency bands to obtain a multi-frequency band signal group, and energy calculation and spectral density analysis are performed on each frequency band signal in the multi-frequency band signal group to obtain frequency band characteristic parameters; According to the frequency band characteristic parameters and the human ear auditory masking threshold, the multi-frequency band signal group is evaluated for perceptual importance to obtain a frequency band weight sequence; Perform maximum entropy threshold segmentation on the frequency band weight sequence, divide the frequency bands with weight values higher than the threshold into the key frequency band set, and divide the frequency bands with weight values lower than the threshold into the enhanced frequency band set to obtain a frequency band stratification result; According to the frequency band stratification result, the multi-frequency band signal group is reorganized, and the signals in the key frequency band set are reconstructed into a time-domain signal through inverse wavelet transform to obtain the basic sound quality layer data; Perform wavelet coefficient rearrangement on the signals in the enhanced frequency band set, and organize the frequency band signals into a hierarchical structure through coefficient matrix transformation to obtain the enhanced sound quality layer data.
[0031] Specifically, the adaptive encoding matrix contains dynamically calculated bit allocation parameters , where is the frequency index, Specifies the amount of bit resources allocated for each frequency band. These parameters directly determine the way the audio input signal is divided in different frequency ranges. Assume the audio input signal is , where represents the discrete-time index, and the frequency band division is achieved through wavelet transform or filter bank. Taking wavelet transform as an example, by applying the discrete wavelet transform to , a multi-band signal group is obtained, and the formula is: ; where is the wavelet basis function, represents the frequency band, and represents the time. The wavelet transform decomposes the signal into multiple frequency bands, and each frequency band contains the energy information of a specific frequency range. For the multi-band signal group , energy calculation and spectral density analysis are performed to extract the characteristic parameters of each frequency band. The formula for energy calculation is: ; where represents the total energy of the frequency band , and the energy accumulation over the entire time dimension is obtained by summation. The spectral density analysis calculates the distribution characteristics of the signal in the frequency domain, and the formula is: ; where represents the average power spectral density of the frequency band , and is the total number of time frames. These characteristic parameters reflect the frequency characteristics of the signal. According to the frequency band characteristic parameters and the masking threshold in the human ear auditory model, the perceptual importance of each frequency band is evaluated. The formula for evaluating the perceptual importance is: ; where is the threshold calculated according to the human ear masking effect, indicating the lowest signal energy that can be perceived by the human ear in this frequency band. When is close to or lower than , the weight of this frequency band will be lower; conversely, when is significantly higher than , will be higher. The perceptual importance weight sequence provides a quantitative description of the contribution of each frequency band to the sound quality. To achieve frequency band stratification, the weight sequence is divided into critical frequency bands and enhanced frequency bands by the maximum entropy threshold segmentation method. The goal of the maximum entropy threshold segmentation is to find a threshold , such that the entropy of the segmented subsets reaches the maximum. The formula is expressed as: ; Among them, and are respectively the frequency band sets higher and lower than , and are the corresponding probability distributions. By optimizing , the optimal threshold is determined. The frequency bands higher than are classified into the key frequency band set , and the frequency bands lower than are classified into the enhanced frequency band set . After obtaining the hierarchical result, data reorganization is performed on the multi-band signal group. For the key frequency band set , its signal is reconstructed into a time-domain signal through inverse wavelet transform, and the formula is: ; The reconstructed signal is the basic audio quality layer data, retaining the main information of the audio signal to ensure the basic audio quality experience. For the enhanced frequency band set , its signal is subjected to wavelet coefficient rearrangement and hierarchical structure organization. The wavelet coefficient rearrangement sorts according to the importance of the enhanced frequency band signal, adjusting the coefficient order to improve storage and transmission efficiency. Through the coefficient matrix transformation , these frequency band signals are organized into a hierarchical structure: ; Among them, is the enhanced layer coefficient matrix, used to represent the hierarchical relationship of the frequency band signals. The organized data is the enhanced audio quality layer data, containing the detailed information of the signal, used to enhance the overall audio experience.
[0032] In a specific embodiment, the process of executing step 105 may specifically include the following steps: Input the basic audio quality layer data into the lossless encoding module of the main processor, perform integer transformation and entropy encoding operations on the basic audio quality layer data, and obtain the basic layer lossless bitstream; Perform data grouping on the basic layer lossless bitstream, and divide the data into multiple coding units based on the audio frame boundary information to obtain the basic layer coding unit sequence; Input the enhanced audio quality layer data into the quantization encoding module of the digital signal processor, perform non-linear quantization on the enhanced audio quality layer data according to the dynamic quantization parameters, and obtain the quantization coefficient matrix; Perform entropy coding and run-length coding on the quantized coefficient matrix, and generate enhanced layer coded data according to the data redundancy and quantization accuracy; Based on the base layer coding unit sequence and the enhanced layer coded data, establish a multi-layer coded data packet, and perform data verification and synchronization marker insertion operations on the multi-layer coded data packet. Add the check code and synchronization information to the packet header to obtain a coded data stream with check; According to the processing load and power status of the Bluetooth chip, optimize the transmission of the coded data stream with check, and add layer identification and priority information to the data packet to obtain the target coded data stream.
[0033] Specifically, input the base audio quality layer data into the lossless coding module of the main processor for processing. Assume that the base audio quality layer data is , where represents the discrete time index. The first step of lossless coding is integer transformation, which converts the floating-point represented audio data into integer representation to reduce quantization error and improve compression efficiency. Integer transformation is achieved through a linear mapping, and the formula is: ; where is the quantization factor, represents the floor operation. The data after integer transformation retains the core information of the audio signal and is also convenient for subsequent entropy coding. Perform entropy coding on the data after integer transformation. Entropy coding is a lossless compression technology based on data distribution, which achieves efficient compression by assigning shorter coding lengths to frequently occurring values. Entropy coding methods include Huffman coding and arithmetic coding, and their basic formula is: ; where is the number of bits after coding, is the occurrence probability of the integer data . The data stream after entropy coding is the base layer lossless code stream , which has a high compression ratio and retains the complete information of the base audio quality layer. Perform data grouping on the base layer lossless code stream to adapt to the framing requirements of Bluetooth transmission. Based on the audio frame boundary information, divide into multiple coding units, and the length of each unit matches the length of the Bluetooth transmission frame. The formula for data grouping is: ; where represents the th coding unit. After the grouping operation, obtain the base layer coding unit sequence , ensure that the data is restored efficiently and error - free during transmission. At the same time, input the enhanced audio quality layer data into the quantization and encoding module of the DSP. To compress the enhanced layer data, perform non - linear quantization on it according to the dynamic quantization parameters. The formula for non - linear quantization is: ; where is the non - linear quantization factor, which controls the balance between quantization accuracy and range. The quantized data is organized into a quantization coefficient matrix , where represents the frequency band index, represents the time - frame index. The quantized coefficient matrix is entropy - encoded and run - length - encoded to achieve compression. Entropy encoding is carried out in a similar way to the base layer, while run - length encoding specifically compresses the redundancy of consecutive repeated values. The formula for run - length encoding is ; where is the repeated quantization value, is the number of its repetitions, is the total number of runs. After encoding, the enhanced layer encoded data is generated. Combine the base layer encoded unit sequence and the enhanced layer encoded data to establish a multi - layer encoded data packet. Add independent header information to the data of each layer to mark the hierarchical structure and data boundaries. The formula for the multi - layer encoded data packet is: ; where, represents the concatenation operation of the data packet, and are the header information of the base layer and the enhanced layer respectively. To ensure the integrity and synchronization of the data packet, add a checksum and a synchronization mark to the header. The checksum is generated according to the parity of the data or CRC check, while the synchronization mark is used to identify the starting position of the data packet. The data packet after adding the checksum and synchronization information is represented as: ; The final data packet is the encoded data stream with check . According to the processing load and power status of the Bluetooth chip, optimize the transmission of the encoded data stream with check. Add a layer identifier and priority information to the data packet to adjust the transmission order of the base layer and enhanced layer data. The formula for setting the priority is: ; The data in the base layer has a higher priority to ensure that the core audio quality information can be transmitted completely in case of limited bandwidth or transmission interruption; the data in the enhancement layer has a lower priority and provides additional audio quality improvement when conditions permit.
[0034] In a specific embodiment, the execution of the audio encoding and decoding method further includes the following steps: Perform packet parsing on the target encoded data stream to obtain multi-layer decoded data; Input the multi-layer decoded data into the state modeling module to model the decoding state space, where the decoding state space includes signal-to-noise ratio, frequency distortion degree, and quantization error distribution, and map the probability distribution of the current decoding state to a belief state vector to obtain the initial decoding belief state; Perform Markov chain prediction on the initial decoding belief state to obtain the corrected decoding belief state, and perform parameter mapping on the corrected decoding belief state to obtain an optimized decoding parameter set; Perform entropy decoding on the base layer code stream in the multi-layer decoded data according to the optimized decoding parameter set, convert the compressed data of Huffman coding and run-length coding into time-domain audio samples, and perform inverse integer transformation to restore the basic audio signal to obtain the base layer decoded data; Perform inverse quantization processing on the enhancement layer code stream in the multi-layer decoded data according to the optimized decoding parameter set, restore the quantization coefficient to the frequency-domain signal representation based on the non-linear quantization inverse mapping function, and reconstruct the time-domain enhancement signal through wavelet inverse transformation to obtain the enhancement layer decoded data; Perform time-domain superposition synthesis on the base layer decoded data and the enhancement layer decoded data to obtain the audio output signal.
[0035] Specifically, input the target encoded data stream into the packet parsing module. By identifying its packet header information and layer identifier, separate the multi-layer decoded data into the base layer code stream and the enhancement layer code stream . The basic formula for packet parsing is: ; Among them, the Parse function parses according to the synchronization mark and the check code in the packet header, verifies the integrity of the packet, and decouples the data streams of different layers. For example, if the layer identifier in the packet header indicates that the current packet belongs to the base layer, the data will be allocated to after parsing; if the identifier is the enhancement layer, it corresponds to . The parsed multi-layer decoded data enters the state modeling module for establishing a decoding state space model. The decoding state space includes signal-to-noise ratio (SNR), frequency distortion degree (FDR), and quantization error distribution (QED). Assuming the decoding state is s, its definition is: ; The signal-to-noise ratio is calculated as: ; Wherein, is the original signal, is the noise signal. The frequency distortion is described by calculating the frequency-domain error: ; The quantization error distribution is generated by statistically enhancing the error values of the layer quantization data. The probability distribution of the decoding state is established on the statistical model of historical data and mapped to the belief state vector through the Bayesian formula : ; Wherein, is the currently observed decoding state parameter. After obtaining the initial decoding belief state , the modified state is predicted using a Markov chain. The transition equation of the Markov chain is: ; Wherein, is the state transition matrix, representing the transition probability from the current state to the next state. The modified belief state more accurately reflects the dynamic changes occurring during the decoding process. The modified belief state is converted into an optimized decoding parameter set through the parameter mapping function : : ; The parameter set includes entropy coding parameters, quantization reflection mapping parameters, and wavelet transform parameters for decoding. According to the optimized decoding parameter set , the base layer bitstream is entropy decoded. Entropy decoding includes Huffman decoding and run-length decoding, and is used to restore the compressed data to an integer form: ; The signal in integer form is restored to the base audio signal through an inverse integer transform : ; Wherein, is the quantization factor used during the original encoding. For the enhancement layer bitstream , inverse quantization processing is performed through an optimized quantization reflection mapping function to restore the quantization coefficients to a frequency-domain signal : ; Wherein, They are non-linear quantization parameters. The frequency-domain signal after inverse quantization is reconstructed into a time-domain signal through inverse wavelet transform : ; where is the wavelet basis function. After obtaining the base layer decoded data and the enhancement layer decoded data the two are superimposed in the time domain to synthesize the final audio output signal : ; wherein is the gain control factor of the enhancement layer, which is used to adjust the contribution of the enhancement layer to the overall audio.
[0036] The audio encoding and decoding method in the embodiments of the present invention is described above. Next, the audio encoding and decoding system in the embodiments of the present invention will be described. Please refer to Figure 2 One embodiment of the audio encoding and decoding system in the embodiments of the present invention includes: An analysis module 201, configured to perform time-domain analysis and frequency-domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; A fusion module 202, configured to perform feature extraction and fusion processing on the audio feature matrix and the user audio operation data to generate a target feature vector including frequency band energy distribution and user preference information; A calculation module 203, configured to input the target feature vector into a preset neural network model for dynamic parameter calculation to generate an adaptive encoding matrix; A partitioning module 204, configured to perform hierarchical partitioning on the audio input signal according to the adaptive encoding matrix to obtain base sound quality layer data and enhanced sound quality layer data; An encoding module 205, configured to perform encoding processing on the base sound quality layer data and the enhanced sound quality layer data respectively through a heterogeneous computing unit, perform lossless encoding of the base sound quality layer data by a main processor, and perform lossy encoding of the enhanced sound quality layer data by a digital signal processor to generate a target encoded data stream.
[0037] Through the collaborative cooperation of the above-mentioned various components, through dual analysis in the time domain and frequency domain combined with the user's listening habits, the accurate extraction of audio features is achieved, and a neural network model is used for dynamic parameter calculation, enabling the encoding parameters to be adaptively adjusted. Based on a hierarchical processing architecture, efficient compression and quality control of audio data are realized. This method, through the collaborative work of heterogeneous computing units, makes full use of the respective advantages of the main processor and the digital signal processor, and adopts a task priority scheduling mechanism to achieve efficient allocation of processing resources, and can dynamically adjust the processing strategy to control system power consumption. In addition, the present invention introduces a modified belief Markov decision process to enhance the robustness of the decoding process, and enhances the system's ability to handle data errors through a multi-agent actor-attention-critic algorithm, achieving a dual optimization of audio processing accuracy and processing time. In terms of user experience, this method provides a personalized auditory experience through adaptive parameter adjustment, ensuring high-quality audio output and real-time processing capabilities. At the same time, through hierarchical coding technology and dynamic parameter adaptation methods, the system's memory occupancy and overall power consumption are significantly reduced, the data transmission efficiency is optimized, and the conservation and utilization of system resources are realized.
[0038] Above Figure 2 The audio codec system in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the audio codec device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0039] Figure 3 It is a schematic structural diagram of an audio codec device provided by an embodiment of the present invention. The audio codec device 300 may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device ends). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the audio codec device 300. Further, the processor 310 may be set to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the audio codec device 300 to implement the steps of the above audio codec method.
[0040] The audio codec device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The structure of the illustrated audio codec device does not constitute a limitation on the audio codec device provided by the present invention, and may include more or fewer components than those shown, or combine certain components, or have different component arrangements.
[0041] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the audio codec method.
[0042] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0043] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0044] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An audio coding and decoding method, characterized in that: The method comprises: Perform time domain analysis and frequency domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; Performing feature extraction and fusion processing on the audio feature matrix and the user audio operation data to generate a target feature vector including frequency band energy distribution and user preference information; Inputting the target feature vector into a preset neural network model to perform dynamic parameter calculation to generate an adaptive coding matrix; According to the adaptive coding matrix, the audio input signal is divided into layers to obtain basic sound quality layer data and enhanced sound quality layer data; The basic sound quality layer data and the enhanced sound quality layer data are encoded and processed separately by heterogeneous computing units, the main processor performs lossless encoding of the basic sound quality layer data, and the digital signal processor performs lossy encoding of the enhanced sound quality layer data to generate a target encoded data stream.
2. The audio encoding and decoding method according to claim 1, characterized in that: The step of performing time domain analysis and frequency domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix includes: Performing frame processing on the audio input signal according to a preset time window length to obtain a plurality of time frames, and performing Hamming window weighting processing on each of the time frames to obtain a weighted time frame sequence; Performing a fast Fourier transform on the weighted time frame sequence to obtain a frequency domain coefficient sequence, and performing nonlinear frequency band division on the frequency domain coefficient sequence to divide the frequency band into a plurality of critical frequency bands based on the auditory characteristics of the human ear to obtain a critical frequency band sequence; Performing energy calculation on each frequency band in the critical frequency band sequence to obtain frequency band energy values, and arranging the frequency band energy values in a matrix according to time frames and frequency band indices to obtain a frequency band energy initial matrix; The initial frequency band energy matrix is normalized to obtain a normalized frequency band energy matrix, and the normalized frequency band energy matrix is convolved with a preset auditory weighting coefficient to obtain an audio feature matrix.
3. The audio encoding and decoding method according to claim 2, characterized in that: The step of extracting and fusing the audio feature matrix and the user audio operation data to generate a target feature vector including frequency band energy distribution and user preference information includes: Performing a time series continuity analysis on the frequency band energy values in the audio feature matrix to obtain a frequency band energy change sequence, and collecting data on the user's volume adjustment operation, equalizer setting parameters, and environmental noise adaptation parameters to obtain a user operation data sequence; According to the user operation data sequence, data statistics are collected on the user's commonly used volume levels and preferred frequency response curves to obtain user preference feature data; Performing data alignment on the frequency band energy change sequence and the user preference feature data to obtain an aligned feature sequence, and performing data standardization processing on the aligned feature sequence to obtain standardized feature data; Feature selection is performed on the standardized feature data, key feature dimensions are screened by calculating information gain values to obtain a reduced-dimensional feature subset, and a nonlinear transformation is performed on the reduced-dimensional feature subset to obtain a target feature vector.
4. The audio encoding and decoding method according to claim 3, characterized in that: The step of inputting the target feature vector into a preset neural network model to perform dynamic parameter calculation to generate an adaptive coding matrix includes: Inputting the frequency band feature subvector in the target feature vector into a feature extraction network in a neural network model, wherein the feature extraction network comprises three convolutional layers and two fully connected layers, each convolutional layer uses a ReLU activation function and a batch normalization layer, extracts the local correlation of the frequency band features through a convolution operation, and obtains a frequency band feature mapping matrix; Inputting the user feature subvector in the target feature vector into the user modeling network in the neural network model, the user modeling network comprises four fully connected layers, wherein the first fully connected layer and the second fully connected layer use the ReLU activation function, the third fully connected layer uses the Sigmoid activation function, and the fourth fully connected layer uses the Tanh activation function, and the user preference features are captured through nonlinear transformation to obtain a user feature mapping matrix; Performing attention feature fusion on the frequency band feature mapping matrix and the user feature mapping matrix to obtain a fused feature tensor; The fused feature tensor is input into a parameter generation network in the neural network model, the parameter generation network adopts a multi-layer perceptron structure, a coding parameter vector is obtained, and a transposed convolution operation is performed on the coding parameter vector to obtain an adaptive coding matrix.
5. The audio encoding and decoding method according to claim 1, characterized in that: The step of dividing the audio input signal into layers according to the adaptive coding matrix to obtain basic sound quality layer data and enhanced sound quality layer data includes: Dividing the audio input signal into frequency bands according to the bit allocation parameters in the adaptive coding matrix to obtain a multi-band signal group, and performing energy calculation and spectral density analysis on each frequency band signal in the multi-band signal group to obtain frequency band characteristic parameters; Performing a perceptual importance evaluation on the multi-band signal group according to the frequency band characteristic parameters and the human ear auditory masking threshold to obtain a frequency band weight sequence; Performing maximum entropy threshold segmentation on the frequency band weight sequence, classifying frequency bands with weight values higher than the threshold into a key frequency band set, and classifying frequency bands with weight values lower than the threshold into an enhanced frequency band set, to obtain a frequency band stratification result; Reorganize the multi-band signal group according to the frequency band stratification result, reconstruct the signal in the key frequency band set into a time domain signal by inverse wavelet transform, and obtain basic sound quality layer data; The wavelet coefficients of the signals in the enhanced frequency band set are rearranged, and the frequency band signals are organized into a hierarchical structure through coefficient matrix transformation to obtain enhanced sound quality layer data.
6. The audio encoding and decoding method according to claim 1, characterized in that: The encoding process of the basic sound quality layer data and the enhanced sound quality layer data is performed by the heterogeneous computing unit, the main processor performs lossless encoding of the basic sound quality layer data, and the digital signal processor performs lossy encoding of the enhanced sound quality layer data to generate a target encoded data stream, including: Inputting the basic sound quality layer data into a lossless coding module of a main processor, performing integer transformation and entropy coding operations on the basic sound quality layer data to obtain a basic layer lossless code stream; Grouping the base layer lossless code stream, dividing the data into multiple coding units based on audio frame boundary information, and obtaining a base layer coding unit sequence; Inputting the enhanced sound quality layer data into a quantization coding module of a digital signal processor, performing nonlinear quantization on the enhanced sound quality layer data according to a dynamic quantization parameter, and obtaining a quantization coefficient matrix; Performing entropy coding and run-length coding on the quantization coefficient matrix, and generating enhanced layer coding data according to data redundancy and quantization accuracy; Establishing a multi-layer coded data packet based on the base layer coding unit sequence and the enhanced layer coded data, performing data check and synchronization mark insertion operations on the multi-layer coded data packet, adding a check code and synchronization information to a data packet header, and obtaining a coded data stream with check; The transmission of the coded data stream with verification is optimized according to the processing load and power status of the Bluetooth chip, and the layer identification and priority information are added to the data packet to obtain the target coded data stream.
7. The audio encoding and decoding method according to claim 6, characterized in that: The audio encoding and decoding method also includes: Parsing the target coded data stream for data packets to obtain multi-layer decoded data; Input the multi-layer decoding data into the state modeling module, model the decoding state space, the decoding state space includes signal-to-noise ratio, frequency distortion and quantization error distribution, and map the probability distribution of the current decoding state into a belief state vector to obtain the decoding initial belief state; Performing Markov chain prediction on the initial decoding belief state to obtain a revised decoding belief state, and performing parameter mapping on the revised decoding belief state to obtain an optimized decoding parameter set; Performing entropy decoding on the base layer bitstream in the multi-layer decoded data according to the optimized decoding parameter set, converting the compressed data of Huffman coding and run-length coding into time-domain audio samples, and performing integer inverse transformation to restore the basic audio signal, thereby obtaining base layer decoded data; Dequantizing the enhanced layer bitstream in the multi-layer decoded data according to the optimized decoding parameter set, restoring the quantized coefficients to frequency domain signal representation based on a nonlinear quantization inverse mapping function, and reconstructing the time domain enhanced signal through inverse wavelet transform to obtain enhanced layer decoded data; The base layer decoded data and the enhanced layer decoded data are combined in time domain to obtain an audio output signal.
8. An audio codec system, characterized in that: Used to perform the audio coding and decoding method according to any one of claims 1 to 7, the system comprises: An analysis module is used to perform time domain analysis and frequency domain analysis on the audio input signal of the Bluetooth headset to generate an audio feature matrix; A fusion module, used for performing feature extraction and fusion processing on the audio feature matrix and the user audio operation data to generate a target feature vector containing frequency band energy distribution and user preference information; A calculation module, used for inputting the target feature vector into a preset neural network model to perform dynamic parameter calculation and generate an adaptive coding matrix; A division module, used for dividing the audio input signal into layers according to the adaptive coding matrix to obtain basic sound quality layer data and enhanced sound quality layer data; The encoding module is used to encode the basic sound quality layer data and the enhanced sound quality layer data respectively through heterogeneous computing units, the main processor performs lossless encoding of the basic sound quality layer data, and the digital signal processor performs lossy encoding of the enhanced sound quality layer data to generate a target encoded data stream.
9. An audio codec device, characterized in that: The audio codec device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the audio codec device to perform the audio codec method according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the audio encoding and decoding method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Digitalized coding method and system for medium-wave transmitter
CN120567370A
A digital coding method and system for medium wave transmitter
CN120567370B
Fast coding and decoding method and system for real-time audio
CN120766691A
Speech reconstruction method and system based on entropy coding residual quantization and spectrum repair
CN121096348A