Low-power communication method, device, SoC chip and storage medium for wireless audio devices

By sampling and frequency domain analysis of heterogeneous audio data of wireless audio equipment, extracting frequency domain features, performing speech enhancement and device grouping encoding, the problem of low energy utilization efficiency in the prior art is solved, and high-efficiency and low-power audio data transmission is achieved.

CN119649835BActive Publication Date: 2025-07-25SHENZHEN HESHENGCHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510179933.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-25
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The existing low-power communication methods for wireless audio equipment ignore the multi-dimensional characteristics of audio data and the heterogeneous characteristics between devices, resulting in low energy utilization efficiency during transmission and difficult to ensure the transmission quality of audio data.

Method used

By sampling the heterogeneous audio data of multiple target audio devices, obtaining frequency domain feature data, performing multi-device acoustic feature extraction and multi-frequency subband speech enhancement, heterogeneous segmentation grouping and differentiated encoding, then performing transmission unit division and collaborative signal modulation, and finally performing collaborative resource allocation and data stream transmission to achieve low-power transmission.

Benefits of technology

It effectively reduces data redundancy during transmission, fully considers the characteristics of different devices, which not only ensures audio quality, but also reduces transmission power consumption, and realizes efficient and low-power transmission of heterogeneous audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649835B_ABST
    Figure CN119649835B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of wireless communication technologies, and discloses a low-power communication method, apparatus, SoC chip and storage medium for wireless audio devices. The method includes: sampling the acoustic features of heterogeneous audio data of multiple target audio devices, calculating the voice spectra corresponding to the device characteristics, and extracting the acoustic features of multiple devices, and performing voice enhancement in multiple frequency subbands, heterogeneous segmentation and grouping of multiple target audio devices, check value calculation and differential coding on the extracted acoustic feature sequences to obtain multi-source coded voice segments; dividing the multi-source coded voice segments into transmission units and performing cooperative signal modulation, cooperative resource allocation and data stream transmission for each transmission unit to obtain a transmitted voice data stream; obtaining the transmitted voice data stream for low-power transmission and performing audio restoration on the obtained data to obtain a target audio data stream transmitted by multiple target audio devices. This application reduces the communication power consumption of multiple target audio devices and ensures the audio transmission quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular, to a low-power communication method, apparatus, SoC chip, and storage medium for wireless audio devices. Background Art

[0002] In the field of wireless audio communication, low-power data transmission between audio devices is one of the key issues currently under research. With the popularization of intelligent audio devices, the demand for audio data transmission between multiple devices is increasing day by day, while the problem of device energy consumption is becoming increasingly prominent. Currently, the industry generally adopts the method of compressing and encoding audio data before transmission to reduce energy consumption, or uses adaptive power control technology to adjust the transmission power to reduce transmission power consumption. However, these methods have many problems when dealing with the transmission of multi-device heterogeneous audio data, such as insufficient data feature extraction, serious signal distortion, and high transmission energy consumption, and it is difficult to meet the actual application requirements.

[0003] Currently, sub-band coding technology is used to perform frequency division processing on audio data, or a cooperative transmission scheme is adopted for resource scheduling in order to reduce the power consumption of audio data transmission. However, these methods still face challenges in aspects such as insufficient utilization of data features, poor cooperation between devices, and low resource allocation efficiency. Especially when dealing with the data transmission of multiple heterogeneous audio devices, these methods often ignore some unique features of audio data, such as device characteristic differences, signal spectrum characteristics, data priorities, etc., and these features have a significant impact on transmission energy consumption and audio quality. That is, the existing low-power communication methods for wireless audio devices ignore the multi-dimensional features of audio data and the heterogeneous characteristics between devices, resulting in low energy utilization efficiency during the transmission process and it is difficult to ensure the transmission quality of audio data. Summary of the Invention

[0004] The main objective of the present invention is to solve the problem that the existing low-power communication methods for wireless audio devices ignore the multi-dimensional features of audio data and the heterogeneous characteristics between devices, resulting in low energy utilization efficiency during the transmission process and it is difficult to ensure the transmission quality of audio data.

[0005] The first aspect of the present invention provides a low-power communication method for wireless audio devices. The low-power communication method for wireless audio devices includes: sampling the acoustic features of heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points, and calculating the speech spectra corresponding to the device characteristics for each of the initial acoustic sampling points to obtain the frequency-domain feature data corresponding to each of the target audio devices; extracting the multi-device acoustic features from the frequency-domain feature data to obtain multiple paths of acoustic feature sequences, and performing speech enhancement on multiple frequency subbands for each of the acoustic feature sequences to obtain multiple paths of enhanced speech streams; performing heterogeneous segmentation and grouping of multiple target audio devices on each of the enhanced speech streams to obtain a speech stream grouping result, and calculating a check value and performing differential coding on the speech stream grouping result to obtain multi-source coded speech segments; dividing the multi-source coded speech segments into transmission units and performing cooperative signal modulation on each transmission unit to obtain a modulated audio stream, and performing cooperative resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream; obtaining the transmitted speech data stream for low-power transmission, and performing audio restoration on the transmitted speech data stream to obtain the target audio data stream transmitted by multiple target audio devices.

[0006] Optionally, in the first implementation manner of the first aspect of the present invention, the sampling the acoustic features of heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points includes: performing downsampling on the heterogeneous audio data of multiple target audio devices based on a preset device sampling rate to obtain raw audio data, and performing frame division on the raw audio data based on a preset fixed sampling interval to obtain initial audio data frames; performing sound pressure detection on the initial audio data frames to obtain audio energy distribution data, and performing threshold detection of speech activity on the audio energy distribution data to obtain speech segment data; performing continuity detection and timing marking on the speech segment data to obtain audio frame sequence data, and performing marking and sampling value quantization on the audio frame sequence data based on the device identification code corresponding to each of the target audio devices to obtain multiple paths of initial acoustic sampling points.

[0007] Optionally, in the second implementation manner of the first aspect of the present invention, the calculating the frequency-domain characteristic data corresponding to each target audio device by performing voice spectrum calculation of corresponding device characteristics on each of the initial acoustic sampling points includes: performing mean and variance statistics on the initial acoustic sampling points to obtain an audio statistical sequence, and calculating a deviation threshold of the audio statistical sequence based on a preset confidence interval and audio statistical parameters; extracting the sampling points with sampling point amplitudes greater than the deviation threshold from the initial acoustic sampling points to obtain effective audio sampling points, and calculating frequency-domain transformation values of each of the effective audio sampling points based on the corresponding frequency response coefficients of the target audio device to obtain an initial spectrum sequence; calculating the energy distribution of each of the initial spectrum sequences in a preset spectrum frequency band, and filtering out frequency components greater than a preset device response interval to obtain a filtered spectrum sequence, and integrating the filtered spectrum sequence with the corresponding audio sampling time and the device codes corresponding to each of the target audio devices to obtain the frequency-domain characteristic data corresponding to each of the target audio devices.

[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the extracting multi-device acoustic characteristics from the frequency-domain characteristic data to obtain multiple paths of acoustic characteristic sequences includes: respectively calculating the energy distribution entropy values of the frequency-domain characteristic data within a preset multiple time window to obtain a frequency-domain entropy value sequence, and calculating the spectral flatness of the frequency-domain characteristic data to obtain a frequency-domain flatness sequence, and extracting the modulation frequency components of the frequency-domain characteristic data to obtain a frequency-domain modulation characteristic sequence, and performing parameter integration on the frequency-domain entropy value sequence, the frequency-domain flatness sequence and the frequency-domain modulation characteristic sequence to construct an audio characteristic parameter table; performing a linear transformation on the characteristic parameter table to obtain a reference audio characteristic set, and performing a characteristic projection transformation on the frequency-domain audio characteristic data to obtain an acoustic characteristic sequence; performing frequency band division on the acoustic characteristic sequence to obtain a sub-band characteristic sequence, and performing energy normalization on the sub-band characteristic sequence to obtain multiple paths of acoustic characteristic sequences.

[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, the performing voice enhancement of multiple frequency sub-bands on each of the acoustic characteristic sequences to obtain multiple paths of enhanced voice streams includes: estimating the sub-band noise power spectrum of the sub-band characteristic sequences in the acoustic characteristic sequences to obtain a sub-band noise spectrum, and calculating the signal power spectrum of the sub-band characteristic sequences to obtain a sub-band signal spectrum; based on the sub-band noise spectrum, performing spectral subtraction noise reduction on each sub-band characteristic in the sub-band signal spectrum to obtain a noise-reduced sub-band signal; performing adaptive gain and phase compensation on the noise-reduced sub-band signal to obtain an enhanced sub-band signal, and performing adaptive sub-band synthesis on each of the enhanced sub-band signals to obtain multiple paths of enhanced voice streams.

[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, the heterogeneous segmentation and grouping of each of the enhanced speech streams for multiple target audio devices, obtaining a speech stream grouping result, and calculating a check value and performing differential encoding on the speech stream grouping result to obtain a multi-source encoded speech segment, includes: calculating the short-time energy and zero-crossing rate of each of the enhanced speech streams to obtain an energy sequence and a zero-crossing rate sequence, and performing feature combination on the energy sequence and the zero-crossing rate sequence to obtain device feature parameters; calculating the cross-correlation coefficients between the device feature parameters, generating a device correlation table, and calculating the audio feature distances between the target audio devices based on the device correlation table to generate a device distance matrix; constructing a hierarchical structure of the device distance matrix to obtain a device grouping structure, and performing index assignment grouping on the enhanced speech streams corresponding to the target audio devices in the device grouping structure to obtain a speech stream grouping result; dividing the speech stream grouping result into blocks to obtain a data block sequence, and averaging the data block sequences corresponding to different target audio devices in the same group to obtain a common audio feature; calculating the signal-to-noise ratio and weight normalization of the common audio feature to obtain audio device weights, and performing weighted combination on the common audio feature based on the audio device weights to obtain a fusion feature vector; performing adaptive quantization and coding verification on the fusion feature vector to obtain a multi-source encoded speech segment.

[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, the division of the multi-source encoded speech segment into transmission units and the cooperative signal modulation of each transmission unit to obtain a modulated audio stream, and the cooperative resource allocation and data stream transmission of the modulated audio stream to obtain a transmitted speech data stream, includes: calculating the information amount and transmission level division of the multi-source encoded speech segment to obtain an audio transmission priority, and performing dynamic bit allocation and configuration of multiple cooperative transmission resource parameters on the multi-source encoded speech segment based on the audio transmission priority to generate a modulated audio stream; calculating the modulation power consumption value of a preset modulation link when modulating the modulated audio stream, and performing segmented threshold division and resource amount configuration calculation on the modulation power consumption value to obtain an initial resource allocation value; integrating and allocating multiple modulation allocation parameters for the initial resource allocation value to generate resource allocation parameters, and performing multi-device timing division power level division on the resource allocation parameters to obtain power configuration parameters; calculating the traffic parameter of the power configuration parameters to obtain a transmission data traffic parameter, and performing data stream transmission on the modulated audio stream based on the transmission data traffic parameter and the resource allocation parameters to obtain a transmitted speech data stream.

[0012] In a second aspect of the present invention, a low-power communication device for a wireless audio device is provided. The low-power communication device for a wireless audio device includes: a spectrum calculation module, configured to sample acoustic features of heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points, and perform voice spectrum calculation of corresponding device characteristics on each of the initial acoustic sampling points to obtain frequency-domain feature data corresponding to each of the target audio devices; a voice enhancement module, configured to extract multi-device acoustic features from the frequency-domain feature data to obtain multiple paths of acoustic feature sequences, and perform voice enhancement on multiple frequency sub-bands of each of the acoustic feature sequences to obtain multiple paths of enhanced voice streams; a segmentation and grouping module, configured to perform heterogeneous segmentation and grouping of multiple target audio devices on each of the enhanced voice streams to obtain a voice stream grouping result, and perform check value calculation and differential coding on the voice stream grouping result to obtain multi-source coded voice segments; a signal modulation module, configured to perform transmission unit division and cooperative signal modulation of each transmission unit on the multi-source coded voice segments to obtain a modulated audio stream, and perform cooperative resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted voice data stream; an audio recovery module, configured to obtain the transmitted voice data stream transmitted with low power, and perform audio recovery on the transmitted voice data stream to obtain a target audio data stream transmitted by multiple target audio devices.

[0013] In a third aspect of the present invention, an SoC chip is provided, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the SoC chip to execute each step of the above-mentioned low-power communication method for a wireless audio device.

[0014] In a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute each step of the above-mentioned low-power communication method for a wireless audio device.

[0015] The above-mentioned low-power communication method, device, SoC chip and storage medium for wireless audio devices. In the embodiments of the present invention, by sampling the acoustic features of heterogeneous audio data of multiple target audio devices, multiple initial acoustic sampling points are obtained, and the voice spectra corresponding to the device characteristics are calculated for each initial acoustic sampling point to obtain the frequency-domain feature data corresponding to each target audio device; the multi-device acoustic features are extracted from the frequency-domain feature data to obtain multiple acoustic feature sequences, and the voice enhancement of multiple frequency sub-bands is performed on each acoustic feature sequence to obtain multiple enhanced voice streams; the heterogeneous segmentation and grouping of multiple target audio devices are performed on each enhanced voice stream to obtain the voice stream grouping result, and the check value calculation and differential coding are performed on the voice stream grouping result to obtain multi-source coded voice segments; the transmission unit division and the cooperative signal modulation of each transmission unit are performed on the multi-source coded voice segments to obtain the modulated audio stream, and the cooperative resource allocation and data stream transmission are performed on the modulated audio stream to obtain the transmitted voice data stream; the transmitted voice data stream for low-power transmission is obtained, and the audio restoration is performed on the transmitted voice data stream to obtain the target audio data stream transmitted by multiple target audio devices. Compared with the prior art, the present application samples and analyzes the spectra of heterogeneous data of multiple audio devices to obtain the frequency-domain features of each device, then extracts the acoustic features of these frequency-domain features, and performs voice enhancement through sub-band processing; then performs device grouping and coding processing on the enhanced voice stream; then divides the coded voice segments into transmission units and performs signal modulation, and performs data transmission according to the resource allocation strategy; finally, performs audio signal restoration on the received data. Through hierarchical data processing and transmission control, the problem of cooperative transmission of multi-device heterogeneous audio data is solved. Especially in feature extraction and data fusion, the characteristic differences of different devices are fully considered, the data redundancy in the transmission process is effectively reduced, and the hierarchical modulation and resource allocation strategy are adopted, which not only ensures the audio quality but also reduces the transmission power consumption, thus realizing the efficient low-power transmission of heterogeneous audio data as a whole.

[0016] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the specification, the claims and the drawings.

[0017] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically described in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the first embodiment of the low-power communication method for wireless audio devices in the embodiments of the present invention;

[0019] Figure 2Schematic diagram of an embodiment of the low-power communication device of the wireless audio device in the embodiment of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0022] For ease of understanding of this embodiment, the specific process of the embodiment of the present invention will be described below. Please refer to Figure 1 , the first embodiment of the low-power communication method for wireless audio devices in the embodiments of the present invention includes:

[0023] 101. Sample the acoustic features of the heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points, and perform voice spectrum calculations of corresponding device characteristics on each initial acoustic sampling point to obtain the frequency-domain feature data corresponding to each target audio device;

[0024] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems.

[0025] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0026] In this embodiment, the target audio device here refers to various audio devices for audio data transmission, such as wireless headphones, smart speakers, microphone arrays, portable speakers, etc.; the heterogeneous audio data here refers to audio data from different types of audio devices with different characteristics (such as parameters like sampling rate, bit depth, number of channels, etc.). For example, the audio data collected by headphones may have a sampling rate of 48 kHz and 16-bit quantization; the data collected by a microphone array may have multiple channels, a sampling rate of 16 kHz, and 24-bit quantization; the audio data output by a speaker may have a sampling rate of 44.1 kHz and 32-bit quantization; the device characteristics here refer to the inherent characteristic differences of different audio devices in terms of hardware performance and audio acquisition and processing. Based on a preset device sampling rate, downsample the heterogeneous audio data of multiple target audio devices to obtain original audio data, and based on a preset fixed sampling interval, frame the original audio data to obtain initial audio data frames; perform sound pressure detection on the initial audio data frames to obtain audio energy distribution data, and perform threshold detection of voice activity on the audio energy distribution data to obtain voice segment data; perform continuity detection and timing marking on the voice segment data to obtain an audio frame sequence data, and based on the device identification codes corresponding to each target audio device, mark and sample value quantization on the audio frame sequence data to obtain multiple initial acoustic sampling points; perform mean and variance statistics on the initial acoustic sampling points to obtain an audio statistical sequence, and based on a preset confidence interval and audio statistical parameters, calculate the deviation threshold of the audio statistical sequence; extract the sampling points in the initial acoustic sampling points whose sampling point amplitudes are greater than the deviation threshold to obtain effective audio sampling points, and based on the corresponding frequency response coefficients of the target audio devices, calculate the frequency domain transformation values of each effective audio sampling point to obtain an initial frequency spectrum sequence; calculate the energy distribution of each initial frequency spectrum sequence in a preset frequency spectrum band, and filter out the frequency components greater than the preset device response interval to obtain a filtered frequency spectrum sequence, and integrate the filtered frequency spectrum sequence with the corresponding audio sampling time and the device codes corresponding to each target audio device to obtain the frequency domain feature data corresponding to each target audio device.

[0027] In practical applications, first, set the preset sampling rate parameters according to the sampling capabilities of different audio devices. For example, high-performance devices can be set to 48 kHz, and low-performance devices can be set to 16 kHz. Then, downsample the heterogeneous audio data collected by each device according to the lowest sampling rate standard, so as to unify the audio data of different devices to the same sampling rate standard and generate the original audio data. Furthermore, based on the preset fixed sampling interval (such as 20 ms), perform frame segmentation on these original audio data respectively, cut the continuous audio data into equal-length data frames to form the initial audio data frames. Then, by calculating the sound pressure value of each frame of data (that is, converting the audio signal amplitude into the sound pressure level in decibels), obtain the audio energy distribution data reflecting the audio energy distribution characteristics of each frame, and compare this data with the preset voice activity detection threshold value (usually set in the range of 20 - 30 dB) to screen out the effective voice segments exceeding the threshold value and generate the voice segment data. Furthermore, for the obtained voice segment data, by detecting the energy continuity and time relationship between adjacent frames, judge the integrity and timing relationship of the voice segments, and at the same time add timestamp marks to each voice segment to generate the audio frame sequence data with timing information. Finally, according to the unique device identification codes (such as device ID or MAC address) pre-allocated to each audio device, add device identification information to the audio frame sequence data corresponding to each audio device, and perform quantization processing on the sampling values according to the quantization accuracy of the device (such as 16-bit quantization), and finally obtain the multiplexed initial acoustic sampling points with device identification and timing information. Thus, the orderly organization of multiplexed audio data is realized, and the audio data of multiple devices is efficiently processed and transmitted.

[0028] Secondly, by calculating the mean (reflecting the DC bias level of the signal) and variance (reflecting the degree of signal fluctuation) of each sampling sequence in the initial acoustic sampling points, an audio statistical sequence describing the overall statistical characteristics of the audio signal is formed. Then, based on a pre-set confidence interval (usually choosing a confidence level of 95% or 99%) and the obtained statistical parameters, the deviation threshold that can distinguish normal signals and outliers is calculated using statistical theory. Next, each sampling value in the original initial acoustic sampling points is compared with the calculated deviation threshold, and the sampling points with amplitudes greater than the threshold are extracted (where these sampling points contain valid audio information), thus obtaining valid audio sampling points. Then, according to the frequency response characteristics of different audio devices (including the gain response curve of the device at different frequencies), the corresponding frequency response coefficients are determined for each audio device, and these coefficients are applied to the frequency-domain conversion process of the valid audio sampling points. The sampling points in the time domain are converted into a spectral sequence in the frequency domain through methods such as the fast Fourier transform to obtain the initial spectral sequence. Furthermore, after obtaining the initial spectral sequence, the energy distribution in each pre-set frequency band (such as the low-frequency band 20 - 200 Hz, the middle-frequency band 200 - 2000 Hz, and the high-frequency band 2000 - 20000 Hz) is calculated. At the same time, according to the pre-set effective frequency response range of each device (i.e., the frequency interval in which the device can work normally), the frequency components outside this range are filtered out to obtain the filtered spectral sequence. Finally, these filtered spectral sequences are integrated with the time information recorded during data acquisition and the unique identification codes of each device to form a complete frequency-domain feature data packet. This not only realizes the validity judgment and frequency-domain conversion of audio data, but also makes the processing result more in line with the working characteristics of the actual device by considering the frequency response characteristics of the device. At the same time, through the screening and integration of frequency components, the redundancy of data transmission is reduced and the processing efficiency is improved.

[0029] 102. Extract the multi-device acoustic features from the frequency-domain feature data to obtain multiple acoustic feature sequences, and perform speech enhancement on each acoustic feature sequence in multiple frequency sub-bands to obtain multiple enhanced speech streams;

[0030] In this embodiment, the energy distribution entropy values of the frequency-domain feature data within a preset multi-time window are calculated respectively to obtain a frequency-domain entropy value sequence, and the spectral flatness of the frequency-domain feature data is calculated to obtain a frequency-domain flatness sequence, and the modulation frequency components of the frequency-domain feature data are extracted to obtain a frequency-domain modulation feature sequence. Then, the frequency-domain entropy value sequence, the frequency-domain flatness sequence, and the frequency-domain modulation feature sequence are parameter-integrated to construct an audio feature parameter table; a linear transformation is performed on the feature parameter table to obtain a reference audio feature set, and a feature projection transformation is performed on the frequency-domain audio feature data to obtain an acoustic feature sequence; the acoustic feature sequence is divided into frequency bands to obtain a sub-band feature sequence, and the sub-band feature sequence is energy-normalized to obtain a multi-channel acoustic feature sequence; the sub-band noise power spectrum of the sub-band feature sequence in the acoustic feature sequence is estimated to obtain a sub-band noise spectrum, and the signal power spectrum of the sub-band feature sequence is calculated to obtain a sub-band signal spectrum; based on the sub-band noise spectrum, spectral subtraction noise reduction is performed on each sub-band feature in the sub-band signal spectrum to obtain a noise-reduced sub-band signal; an adaptive gain and phase compensation are performed on the noise-reduced sub-band signal to obtain an enhanced sub-band signal, and an adaptive sub-band synthesis is performed on each enhanced sub-band signal to obtain a multi-channel enhanced speech stream.

[0031] In practical applications, first, calculate the energy distribution entropy value of the frequency-domain feature data within multiple preset time windows (such as analysis windows of different scales like 10 ms, 20 ms, 30 ms, etc.). Here, this entropy value reflects the uncertainty and complexity of the signal in the frequency domain, so as to calculate the frequency-domain entropy value sequence that describes the time-varying characteristics of the signal. At the same time, conduct spectral flatness analysis on the frequency-domain feature data, calculate the uniformity of the energy distribution in each frequency band, obtain the frequency-domain flatness sequence that characterizes the spectral shape characteristics of the signal, and extract the modulation characteristics of the signal from the frequency-domain feature data through methods such as band-pass filtering (where these modulation characteristics reflect the periodic change characteristics of the signal, forming a frequency-domain modulation feature sequence). Then, integrate these three different-dimensional feature sequences (entropy value sequence, flatness sequence, and modulation feature sequence) according to the preset feature weights and combination rules to construct an audio feature parameter table containing multi-dimensional feature information. Next, apply linear transformation processing (such as principal component analysis or linear discriminant analysis) to this feature parameter table, transform the original feature space into a new feature space to obtain a more representative reference audio feature set, and use this feature set as a reference to perform feature projection transformation on the original frequency-domain feature data, map the high-dimensional features to a low-dimensional feature space to obtain a more compact acoustic feature sequence. Furthermore, according to the frequency characteristics of the audio signal, divide the entire frequency band into multiple sub-bands according to the preset frequency boundaries (such as dividing by critical bands or octaves), process the features within each sub-band respectively to obtain sub-band feature sequences. Finally, perform energy normalization processing on the feature sequences of each sub-band to eliminate the energy differences between different sub-bands, making the features more comparable, thereby obtaining a standardized multi-channel acoustic feature sequence. It realizes considering the multi-dimensional characteristics of the audio signal and improves the expression ability of the features through feature transformation and normalization processing, providing a feature basis for speech enhancement and data fusion.

[0032] Secondly, estimate the noise power spectrum of the feature sequences in each sub-band. By analyzing the non-speech segments or low-energy segments of the signal, estimate the background noise level in each sub-band to obtain the sub-band noise spectrum reflecting the noise characteristics. At the same time, perform power spectral density analysis on the sub-band feature sequences to calculate the energy distribution of the signal at each frequency point to obtain the sub-band signal spectrum characterizing the useful signal characteristics. Then, based on the sub-band noise spectrum, perform spectral subtraction on the signal spectrum and the noise spectrum in each sub-band of the sub-band signal spectrum, that is, subtract the estimated noise spectrum from the signal spectrum, and this process needs to consider the balance of over-subtraction and under-subtraction. Usually, soft thresholding is used for spectral subtraction to obtain the sub-band signal after preliminary noise reduction. Then, according to the energy distribution and signal-to-noise ratio of each sub-band signal, calculate the corresponding gain compensation coefficient to adjust the amplitude of the denoised signal. At the same time, based on the phase characteristics of the signal, perform phase compensation processing to eliminate the phase distortion that may be introduced during the spectral subtraction process to obtain the enhanced sub-band signal improved in both amplitude and phase. Finally, according to the importance degree and energy distribution of each sub-band signal, determine the appropriate sub-band synthesis weight, and use the weighted overlap and add method to perform weighted synthesis on all enhanced sub-band signals to reconstruct the complete speech signal. The weighted overlap and add formula is:

[0033] ;

[0034] where: is the weight coefficient of the k-th sub-band, is the synthesis window function, is the sub-band signal, R is the frame shift, n is the time index or frame index, k is the index of the sub-band or channel, M is the length of the synthesis window function, r is the sample index of the window function. For example: set 8 sub-bands, 50% overlap between adjacent frames, and use an improved triangular window for synthesis. Each sub-band is assigned a weight according to its signal-to-noise ratio, so as to obtain a multi-channel enhanced speech stream with improved quality. Through sub-band processing, fine-grained suppression of noise is achieved, and the naturalness of the signal is ensured through gain and phase compensation. The final sub-band synthesis guarantees the integrity of the enhanced speech, thereby maintaining the intelligibility and naturalness of the speech signal while reducing noise interference.

[0035] 103. Perform heterogeneous segmentation and grouping of multi-target audio devices on each enhanced speech stream to obtain the speech stream grouping result, and calculate the check value and differential coding on the speech stream grouping result to obtain the multi-source coded speech segment;

[0036] In this embodiment, short-time energy calculation and zero-crossing rate calculation are performed on each enhanced speech stream to obtain an energy sequence and a zero-crossing rate sequence, and feature combination is performed on the energy sequence and the zero-crossing rate sequence to obtain device feature parameters; the cross-correlation coefficients between the device feature parameters are calculated to generate a device correlation table, and based on the device correlation table, the audio feature distances between the target audio devices are calculated to generate a device distance matrix; a hierarchical structure of the device distance matrix is constructed to obtain a device grouping structure, and index assignment grouping is performed on the enhanced speech streams corresponding to the target audio devices in the device grouping structure to obtain a speech stream grouping result; the speech stream grouping result is segmented to obtain a data block sequence, and the mean values of the data block sequences corresponding to different target audio devices in the same group are calculated to obtain common audio features; signal-to-noise ratio calculation and weight normalization are performed on the common audio features to obtain audio device weights, and based on the audio device weights, weighted combination is performed on the common audio features to obtain a fusion feature vector; adaptive quantization and coding verification are performed on the fusion feature vector to obtain a multi-source coded speech segment.

[0037] In practical applications, first, an energy sequence reflecting the signal intensity change is obtained by calculating the signal energy value within each short-time frame, and at the same time, a zero-crossing rate sequence reflecting the signal frequency characteristics is obtained by calculating the number of times the signal crosses the zero level within each analysis frame. Then, these two sequences are weighted and fused according to a preset feature combination rule. The weighted fusion formula for the device feature parameters is as follows:

[0038] ;

[0039] Among them, represents the energy of the nth frame, represents the zero-crossing rate, is the weighting coefficient, is the pitch feature, is the adjustment coefficient, N is the total number of frames, d is the identifier of the device, t is the index of the time period, and F(d) is the value of the characteristic parameter of device d within the time period t. Thus, the device characteristic parameter that can characterize the audio characteristics of each device is calculated. Then, the cross-correlation coefficients between different device characteristic parameters are calculated. These correlation coefficients reflect the similarity degree between the signals collected by different devices, thereby forming a device correlation table that describes the correlation between devices. Based on this correlation table, metric methods such as Euclidean distance or Mahalanobis distance are used to calculate the characteristic distances between devices, constructing a device distance matrix that characterizes the degree of difference between devices. Then, methods such as hierarchical clustering are applied to the distance matrix to construct a hierarchical structure, grouping devices with similar characteristic distances into one group to form a device grouping structure, and assigning a unique index number to each group. The enhanced speech streams corresponding to each device are grouped and marked to obtain the speech stream grouping result. Then, the grouped speech streams are segmented according to the preset data block size to obtain a data block sequence that is convenient for processing. The mean value is calculated for the data block sequences of different devices within the same group, and the common audio characteristics that reflect the common characteristics of the devices in this group are extracted. Then, the signal-to-noise ratio is calculated for these common characteristics to evaluate the quality of the signals collected by each device, and the signal-to-noise ratio values are converted into weights after normalization processing, which are used to represent the importance of different devices in feature fusion. Then, based on these weights, the common audio characteristics are weighted and combined to fuse the characteristics of different devices within the same group. The weighted combination formula is:

[0040] ;

[0041] where, is the time-domain characteristic of device i, is the frequency response function, is the variance, is the gain function, is the fusion weight, N is the total number of devices, t is the time index, and i is the index of the device. Thus, the fusion feature vector that can represent the overall characteristics of the devices in this group is calculated. Finally, the fusion feature vector is subjected to fixed-point quantization processing and a check code is added for encoding to obtain the final multi-source coded speech segment. It realizes the effective organization of multi-device heterogeneous data, ensures the reasonable utilization of the characteristics of each device, and prepares for subsequent low-power transmission through quantization coding. It not only maintains the characteristics of the signals of each device but also realizes the effective fusion of data.

[0042] 104. Divide the multi-source coded speech segment into transmission units and perform cooperative signal modulation on each transmission unit to obtain a modulated audio stream, and perform cooperative resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream;

[0043] In this embodiment, the information amount of the multi-source encoded speech segments is calculated and the transmission levels are divided to obtain the audio transmission priority. Based on the audio transmission priority, dynamic bit allocation and the configuration of various cooperative transmission resource parameters are performed on the multi-source encoded speech segments to generate a modulated audio stream. The modulation power consumption value of a preset modulation link when modulating the modulated audio stream is calculated, and the modulation power consumption value is subjected to segmented threshold division and resource amount configuration calculation to obtain an initial resource allocation value. The initial resource allocation value is integrally allocated with various modulation allocation parameters to generate a resource allocation parameter, and the power level of the multi-device timing is divided for the resource allocation parameter to obtain a power configuration parameter. The flow parameter calculation is performed on the power configuration parameter to obtain a transmission data flow parameter, and based on the transmission data flow parameter and the resource allocation parameter, the modulated audio stream is transmitted as a data stream to obtain a transmitted speech data stream.

[0044] In practical applications, the information entropy value of each encoded segment is calculated to evaluate the amount of information, and then these encoded segments are divided into different transmission levels according to the distribution of the information amount to obtain the audio transmission priority representing the importance degree of the data. Furthermore, based on this priority, bit allocation is performed on the encoded speech segments, and more bit resources are allocated to the high-priority data. At the same time, transmission resource parameters including modulation mode, coding mode, and transmission power are configured, and a modulated audio stream suitable for transmission is generated through the combined configuration of these parameters. Then, the power consumption of each modulation link (including signal modulation, power amplification, channel coding, etc.) during the generation of the modulated audio stream is calculated to obtain the modulation power consumption value reflecting the energy consumption situation, and these power consumption values are divided into different levels according to the preset power consumption threshold. Combining the available resource status, an initial resource allocation scheme is calculated to obtain the initial resource allocation value. Then, this initial allocation value is integrated with multiple modulation parameters (such as modulation order, code rate, time slot length, etc.) to generate a complete resource allocation parameter, and the transmission timing of multiple devices is planned according to these parameters. At the same time, the transmission power is divided into different levels to obtain a detailed power configuration parameter. Then, the data flow size of each transmission time slot is calculated based on the power configuration parameter to obtain the transmission data flow parameter for controlling the transmission rate. Finally, these transmission parameters are applied to the actual transmission process of the modulated audio stream, and data transmission is performed according to the allocated time slot, power, and bit resources to obtain the final transmitted speech data stream. Through multi-level parameter configuration and resource allocation, fine control of the transmission process is achieved, which not only ensures the transmission quality of important data but also realizes the rational utilization of energy and low-power communication.

[0045] 105. Obtain the transmitted speech data stream for low-power transmission, and perform audio restoration on the transmitted speech data stream to obtain the target audio data stream transmitted by multiple target audio devices.

[0046] In this embodiment, the transmitted voice data stream is classified according to priority to obtain classified voice data, and the classified voice data is divided into cache regions to obtain a data block position table; the access times of the cached data in the data block position table are counted to obtain a data access record, and the importance of the data access record is calculated to obtain a replacement order table; the occupancy status of the cached data in the data block position table is detected to obtain a cache occupancy value, and the cache occupancy value is compared with a threshold to obtain cache screening data; the cache screening data is reorganized according to the replacement order table to obtain reorganized cache data, and the reorganized cache data is signal-reconstructed to obtain a target audio data stream.

[0047] In practical applications, the target communication device receives the transmitted voice data stream transmitted with low power consumption, and then classifies the data stream according to the priority identification information carried in the data packet, divides data with different importance levels into different priority categories to obtain classified voice data, and then allocates corresponding cache regions for data with different priorities according to the available cache resources and the capacity requirements of each level of data, establishes the mapping relationship between the data blocks and the storage locations to form a data block position table. Then, the access frequency of the data blocks stored in each cache region is counted, the number of times each data block is read and used is recorded to generate a data access record reflecting the data usage situation, and the importance value of each data block is calculated based on these access records (where this value comprehensively considers the priority and access frequency of the data and is used to generate an order table for guiding the replacement of cached data). Then, the status of each cache region recorded in the data block position table is detected, the data occupancy of each region is counted to obtain a cache occupancy value reflecting the cache usage status, and these occupancy values are compared with a preset cache capacity threshold to identify the cache regions that need data replacement or merging to obtain cache screening data. Then, the screened cache data is reorganized according to the previously generated replacement order table, the data with high importance is retained in the fast access area, while the data with lower importance is moved to the standby area or the occupied space is released to obtain optimized reorganized cache data. Finally, signal reconstruction processing is performed on these reorganized data, including splicing the segmented data, restoring the timing relationship of the signal, and restoring the spectral characteristics of the audio signal, etc., to finally obtain a complete target audio data stream. Thus, through cache management and data reorganization, not only is the fast access of important data ensured, but also the efficient utilization of storage resources is achieved, realizing a stable and reliable communication data recovery function.

[0048] In the embodiments of the present invention, by sampling and performing spectrum analysis on the heterogeneous data of multiple audio devices, the frequency-domain characteristics of each device are obtained. Then, the acoustic characteristics of these frequency-domain characteristics are extracted, and speech enhancement is performed through sub-band processing. Next, the enhanced speech stream is subjected to device grouping and encoding processing. Then, the encoded speech segments are divided into transmission units and signal modulation is performed, and data transmission is carried out according to the resource allocation strategy. Finally, the received data is subjected to audio signal recovery. Through hierarchical data processing and transmission control, the problem of collaborative transmission of multi-device heterogeneous audio data is solved. Especially in feature extraction and data fusion, the characteristic differences of different devices are fully considered, effectively reducing data redundancy during the transmission process. And a hierarchical modulation and resource allocation strategy is adopted, which not only ensures the audio quality but also reduces the transmission power consumption, thus achieving efficient and low-power transmission of heterogeneous audio data as a whole.

[0049] The low-power communication method of the wireless audio device in the embodiments of the present invention has been described above. Next, the low-power communication device of the wireless audio device in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the low-power communication device of the wireless audio device in the embodiments of the present invention includes:

[0050] A spectrum calculation module 201, configured to sample the acoustic characteristics of the heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points, and perform speech spectrum calculation of the corresponding device characteristics on each of the initial acoustic sampling points to obtain the frequency-domain characteristic data corresponding to each of the target audio devices;

[0051] A speech enhancement module 202, configured to extract the multi-device acoustic characteristics from the frequency-domain characteristic data to obtain multiple paths of acoustic characteristic sequences, and perform speech enhancement on multiple frequency sub-bands of each of the acoustic characteristic sequences to obtain multiple paths of enhanced speech streams;

[0052] A segmentation and grouping module 203, configured to perform heterogeneous segmentation and grouping of multiple target audio devices on each of the enhanced speech streams to obtain a speech stream grouping result, and perform check value calculation and differential encoding on the speech stream grouping result to obtain multi-source encoded speech segments;

[0053] A signal modulation module 204, configured to divide the multi-source encoded speech segments into transmission units and perform collaborative signal modulation on each transmission unit to obtain a modulated audio stream, and perform collaborative resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream;

[0054] An audio recovery module 205, configured to obtain the transmitted speech data stream of low-power transmission, and perform audio recovery on the transmitted speech data stream to obtain the target audio data stream transmitted by multiple target audio devices.

[0055] In the embodiments of the present invention, by sampling and performing spectral analysis on the heterogeneous data of multiple audio devices, the frequency-domain characteristics of each device are obtained. Then, the acoustic characteristics of these frequency-domain characteristics are extracted, and speech enhancement is performed through sub-band processing. Next, the enhanced speech stream is subjected to device grouping and encoding processing. Then, the encoded speech segments are divided into transmission units and signal modulation is performed, and data transmission is carried out according to the resource allocation strategy. Finally, the received data is subjected to audio signal restoration. Through hierarchical data processing and transmission control, the problem of collaborative transmission of multi-device heterogeneous audio data is solved. Especially in feature extraction and data fusion, the characteristic differences of different devices are fully considered, effectively reducing data redundancy during the transmission process. And by adopting hierarchical modulation and resource allocation strategies, both the audio quality is guaranteed and the transmission power consumption is reduced. Thus, the efficient and low-power transmission of heterogeneous audio data is achieved as a whole.

[0056] Above Figure 2 The low-power device of the wireless audio device in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the Soc chip in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0057] The Soc chip may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing applications or data (for example, one or more mass storage device ends). Among them, the memory and the storage medium may be transient storage or persistent storage. The program stored in the storage medium may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the Soc chip. Further, the processor may be configured to communicate with the storage medium and execute a series of instruction operations in the storage medium on the Soc chip to implement the steps of the above-mentioned low-power method for wireless audio devices.

[0058] The Soc chip may also include one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that the structure of the Soc chip does not constitute a limitation on the Soc chip provided by the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0059] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for the low-power wireless audio device.

[0060] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0061] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other various media that can store program codes.

[0062] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A low-power communication method for a wireless audio device, characterized in that, The low-power communication method of the wireless audio device includes: Sampling the acoustic features of the heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points, and calculating the voice spectra corresponding to the device characteristics for each of the initial acoustic sampling points to obtain the frequency-domain feature data corresponding to each of the target audio devices; Extracting the multi-device acoustic features from the frequency-domain feature data to obtain multiple paths of acoustic feature sequences, and performing voice enhancement on multiple frequency sub-bands for each of the acoustic feature sequences to obtain multiple paths of enhanced voice streams; Performing heterogeneous segmentation and grouping of multiple target audio devices on each of the enhanced voice streams to obtain a voice stream grouping result, and partitioning the voice stream grouping result to obtain a data block sequence, and calculating the mean value of the corresponding data block sequences of different target audio devices within the same group to obtain the common audio features of the common features between devices in the same group; calculating the signal-to-noise ratio and weight normalization for the common audio features to obtain audio device weights, and based on the audio device weights and a preset weighted combination formula, performing weighted fusion on the corresponding data block sequences of different target audio devices within the same group to obtain a fused feature vector; performing adaptive quantization and coding verification on the fused feature vector to obtain multi-source coded voice segments; Performing transmission unit division and cooperative signal modulation on each of the transmission units for the multi-source coded voice segments to obtain a modulated audio stream, and performing cooperative resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted voice data stream; Obtaining the transmitted voice data stream for low-power transmission, and performing audio restoration on the transmitted voice data stream to obtain the target audio data stream transmitted by multiple target audio devices.

2. The low-power communication method of the wireless audio device according to claim 1, wherein The sampling of the acoustic features of the heterogeneous audio data of multiple target audio devices to obtain multiple paths of initial acoustic sampling points includes: Performing downsampling on the heterogeneous audio data of multiple target audio devices based on a preset device sampling rate to obtain raw audio data, and performing frame division on the raw audio data based on a preset fixed sampling interval to obtain initial audio data frames; Performing sound pressure detection on the initial audio data frames to obtain audio energy distribution data, and performing threshold detection of voice activity on the audio energy distribution data to obtain voice segment data; Performing continuity detection and timing marking on the voice segment data to obtain audio frame sequence data, and performing marking and sampling value quantization on the audio frame sequence data based on the device identification codes corresponding to each of the target audio devices to obtain multiple paths of initial acoustic sampling points.

3. The low-power communication method for a wireless audio device according to claim 1, wherein The calculation of the voice spectra corresponding to the device characteristics for each of the initial acoustic sampling points to obtain the frequency-domain feature data corresponding to each of the target audio devices includes: Performing mean and variance statistics on the initial acoustic sampling points to obtain an audio statistical sequence, and calculating the deviation threshold for the audio statistical sequence based on a preset confidence interval and audio statistical parameters; Extract the sampling points in the initial acoustic sampling points whose sampling point amplitudes are greater than the deviation threshold to obtain valid audio sampling points, and calculate the frequency-domain transformation values of each of the valid audio sampling points based on the corresponding frequency response coefficients of the target audio device to obtain an initial frequency spectrum sequence; Calculate the energy distribution of each of the initial frequency spectrum sequences in a preset frequency spectrum band, and filter out the frequency components greater than a preset device response interval to obtain a filtered frequency spectrum sequence, and integrate the filtered frequency spectrum sequence with the corresponding audio sampling time and the device codes corresponding to each of the target audio devices to obtain the frequency-domain feature data corresponding to each of the target audio devices.

4. The low-power communication method of the wireless audio device according to claim 1, wherein The extraction of multi-device acoustic features from the frequency-domain feature data to obtain multiple acoustic feature sequences includes: Calculate the energy distribution entropy values of the frequency-domain feature data in a preset multi-time window respectively to obtain a frequency-domain entropy value sequence, calculate the spectral flatness of the frequency-domain feature data to obtain a frequency-domain flatness sequence, and extract the modulation frequency components of the frequency-domain feature data to obtain a frequency-domain modulation feature sequence, and perform parameter integration on the frequency-domain entropy value sequence, the frequency-domain flatness sequence, and the frequency-domain modulation feature sequence to construct an audio feature parameter table; Perform a linear transformation on the feature parameter table to obtain a reference audio feature set, and perform a feature projection transformation on the frequency-domain feature data to obtain an acoustic feature sequence; Perform frequency band division on the acoustic feature sequence to obtain a sub-band feature sequence, and perform energy normalization on the sub-band feature sequence to obtain multiple acoustic feature sequences.

5. The low-power communication method of the wireless audio device according to claim 4, characterized in that The multi-frequency sub-band voice enhancement of each of the acoustic feature sequences to obtain multiple enhanced voice streams includes: Estimate the sub-band noise power spectrum of the sub-band feature sequence in the acoustic feature sequence to obtain a sub-band noise spectrum, and calculate the signal power spectrum of the sub-band feature sequence to obtain a sub-band signal spectrum; Based on the sub-band noise spectrum, perform spectral subtraction noise reduction on each sub-band feature in the sub-band signal spectrum to obtain a noise-reduced sub-band signal; Perform adaptive gain and phase compensation on the noise-reduced sub-band signal to obtain an enhanced sub-band signal, and perform adaptive sub-band synthesis on each of the enhanced sub-band signals to obtain multiple enhanced voice streams.

6. The low-power communication method of the wireless audio device according to claim 1, wherein, The heterogeneous segmentation and grouping of each of the enhanced voice streams for multiple target audio devices to obtain a voice stream grouping result includes: Calculate the short-time energy and zero-crossing rate of each of the enhanced voice streams to obtain an energy sequence and a zero-crossing rate sequence, and perform feature combination on the energy sequence and the zero-crossing rate sequence to obtain device feature parameters; Calculate the cross-correlation coefficients between each of the device feature parameters to generate a device correlation table, and based on the device correlation table, calculate the audio feature distances between each of the target audio devices to generate a device distance matrix; Construct a hierarchical structure of the device distance matrix to obtain a device grouping structure, and perform index assignment grouping on the enhanced voice streams corresponding to each of the target audio devices in the device grouping structure to obtain a voice stream grouping result.

7. The low-power communication method of the wireless audio device according to claim 1, characterized in that, Performing transmission unit division on the multi-source encoded speech segments and co-signal modulation of each transmission unit to obtain a modulated audio stream, and performing co-resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream, including: Calculating the information amount and performing transmission level division on the multi-source encoded speech segments to obtain the audio transmission priority, and based on the audio transmission priority, performing dynamic bit allocation and configuration of various co-transmission resource parameters on the multi-source encoded speech segments to generate a modulated audio stream; Calculating the modulation power consumption value of a preset modulation link when modulating the modulated audio stream, and performing segmented threshold division and resource amount configuration calculation on the modulation power consumption value to obtain an initial resource allocation value; Integrating and allocating various modulation allocation parameters to the initial resource allocation value to generate a resource allocation parameter, and performing multi-device timing division and power level division on the resource allocation parameter to obtain a power configuration parameter; Calculating the traffic parameter of the power configuration parameter to obtain a transmitted data traffic parameter, and based on the transmitted data traffic parameter and the resource allocation parameter, performing data stream transmission on the modulated audio stream to obtain a transmitted speech data stream.

8. A low-power communication device for a wireless audio device, characterized in that, The low-power communication device of the wireless audio device includes: A spectrum calculation module, configured to sample the acoustic features of the heterogeneous audio data of multiple target audio devices to obtain multiple initial acoustic sampling points, and perform speech spectrum calculation of the corresponding device characteristics on each of the initial acoustic sampling points to obtain the frequency domain feature data corresponding to each of the target audio devices; A speech enhancement module, configured to extract the multi-device acoustic features from the frequency domain feature data to obtain multiple acoustic feature sequences, and perform speech enhancement of multiple frequency sub-bands on each of the acoustic feature sequences to obtain multiple enhanced speech streams; A segmented grouping module, configured to perform heterogeneous segmented grouping of multiple target audio devices on each of the enhanced speech streams to obtain a speech stream grouping result, and perform chunking on the speech stream grouping result to obtain a data block sequence, and calculate the mean value of the corresponding data block sequences of different target audio devices in the same group to obtain a common audio feature of the common features between devices in the same group; calculating the signal-to-noise ratio and weight normalization of the common audio feature to obtain an audio device weight, and based on the audio device weight and a preset weighted combination formula, performing weighted fusion on the corresponding data block sequences of different target audio devices in the same group to obtain a fused feature vector; performing adaptive quantization and coding verification on the fused feature vector to obtain a multi-source encoded speech segment; A signal modulation module, configured to perform transmission unit division on the multi-source encoded speech segment and co-signal modulation of each transmission unit to obtain a modulated audio stream, and perform co-resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream; An audio recovery module, configured to obtain the transmitted speech data stream transmitted with low power consumption, and perform audio recovery on the transmitted speech data stream to obtain the target audio data stream transmitted by multiple target audio devices.

9. A SoC chip, characterized in that, The SoC chip includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor invokes the instructions in the memory to cause the SoC chip to perform each step of the low-power communication method for a wireless audio device according to any one of claims 1-7.

10. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, each step of the low-power communication method for a wireless audio device according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Data transmission cross-protocol transparent coexistence method based on precoding

    CN119363172A

  • Wireless audio data transmission and reading method and audio playing equipment

    CN119400188A