Low-power communication method and apparatus for wireless audio device, soc chip, and storage medium

WO2026175268A1PCT designated stage Publication Date: 2026-08-27SHENZHEN HESHENGCHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/078454
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-19
Filing Date
2026-02-11
Publication Date
2026-08-27

Smart Images

  • Figure CN2026078454_27082026_PF_FP_ABST
    Figure CN2026078454_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A low-power communication method and apparatus for a wireless audio device, an SoC chip, and a storage medium. The low-power communication method for a wireless audio device comprises: sampling acoustic features of heterogeneous audio data from multiple target audio devices, calculating speech spectra corresponding to device characteristics, extracting multi-device acoustic features, and performing multi-frequency sub-band speech enhancement, heterogeneous segmentation and grouping of the multiple target audio devices, check value calculation, and differentiated encoding on each extracted acoustic feature sequence to obtain a multi-source encoded speech segment; dividing the multi-source encoded speech segment into transmission units, and performing cooperative signal modulation, cooperative resource allocation, and data stream transmission on each transmission unit to obtain a transmitted speech data stream; and acquiring the transmitted speech data stream transmitted with low power, and performing audio recovery on the transmitted speech data stream to obtain a target audio data stream transmitted by the multiple target audio devices, thereby reducing the communication power consumption of the multiple target audio devices and ensuring the audio transmission quality.
Need to check novelty before this filing date? Find Prior Art

Description

Low-power Communication Method, Device, SoC Chip and Storage Medium for Wireless Audio Devices Technical Field The present invention relates to the field of wireless communication technology, and particularly to a low-power communication method, device, SoC chip and storage medium for wireless audio devices. Background Art In the field of wireless audio communication, low-power data transmission between audio devices is one of the key issues currently under research. With the popularization of smart audio devices, the demand for audio data transmission between multiple devices is increasing day by day, while the problem of device energy consumption is becoming increasingly prominent. At present, the industry generally adopts the method of compressing and encoding audio data before transmission to reduce energy consumption, or adopts adaptive power control technology to adjust the transmission power to reduce transmission power consumption. However, these methods have many problems when dealing with the transmission of multi-device heterogeneous audio data, such as insufficient data feature extraction, serious signal distortion, high transmission energy consumption, etc., and it is difficult to meet the actual application requirements. Currently, sub-band coding technology is used to perform frequency division processing on audio data, or a cooperative transmission scheme is adopted for resource scheduling in order to reduce the power consumption of audio data transmission. However, these methods still face challenges in aspects such as insufficient utilization of data features, poor cooperation between devices, and low resource allocation efficiency. Especially when dealing with the data transmission of multiple heterogeneous audio devices, these methods often ignore some unique features of audio data, such as device characteristic differences, signal spectrum characteristics, data priorities, etc., and these features have a significant impact on transmission energy consumption and audio quality. That is, the existing low-power communication methods for wireless audio devices ignore the multi-dimensional features of audio data and the heterogeneous characteristics between devices, resulting in low energy utilization efficiency during the transmission process and it is difficult to ensure the transmission quality of audio data. Summary of the Invention The main object of the present invention is to solve the problem that the existing low-power communication methods for wireless audio devices ignore the multi-dimensional features of audio data and the heterogeneous characteristics between devices, resulting in low energy utilization efficiency during the transmission process and it is difficult to ensure the transmission quality of audio data. This invention provides a low-power communication method for wireless audio devices. The method includes: sampling acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points; calculating the speech spectrum of each initial acoustic sampling point based on the corresponding device characteristics to obtain frequency domain feature data corresponding to each target audio device; extracting multi-device acoustic features from the frequency domain feature data to obtain multiple acoustic feature sequences; and performing multi-frequency sub-band speech enhancement on each acoustic feature sequence to obtain multiple enhanced speech streams; and further... The enhanced speech stream is heterogeneously segmented and grouped by multiple target audio devices to obtain speech stream grouping results. Checksum calculation and differential coding are performed on the speech stream grouping results to obtain multi-source coded speech segments. Transmission unit division and cooperative signal modulation of each transmission unit are performed on the multi-source coded speech segments to obtain a modulated audio stream. Cooperative resource allocation and data stream transmission are performed on the modulated audio stream to obtain a transmitted speech data stream. A low-power transmission transmitted speech data stream is acquired, and audio recovery is performed on the transmitted speech data stream to obtain the target audio data stream transmitted by the multiple target audio devices. In this embodiment of the invention, acoustic features are sampled from heterogeneous audio data of multiple target audio devices to obtain multiple initial acoustic sampling points. Speech spectrum calculations are performed on each initial acoustic sampling point to obtain frequency domain feature data corresponding to each target audio device. Multi-device acoustic features are extracted from the frequency domain feature data to obtain multiple acoustic feature sequences. Multi-frequency sub-band speech enhancement is performed on each acoustic feature sequence to obtain multiple enhanced speech streams. Heterogeneous segmentation and grouping of each enhanced speech stream from multiple target audio devices is performed to obtain speech stream grouping results. Checksum calculation and differential coding are performed on the speech stream grouping results to obtain multi-source coded speech segments. Transmission unit division and cooperative signal modulation are performed on the multi-source coded speech segments to obtain modulated audio streams. Cooperative resource allocation and data stream transmission are performed on the modulated audio streams to obtain transmitted speech data streams. Low-power transmission of the transmitted speech data stream is obtained, and audio recovery is performed on the transmitted speech data streams to obtain the target audio data streams transmitted by the multiple target audio devices. Compared to existing technologies, this application samples and performs spectral analysis on heterogeneous data from multiple audio devices to obtain the frequency domain characteristics of each device. Then, it extracts the acoustic features of these frequency domain characteristics and performs speech enhancement through sub-band processing. Next, it groups and encodes the enhanced speech stream by device. Then, it divides the encoded speech segments into transmission units and modulates the signals, transmitting the data according to a resource allocation strategy. Finally, it recovers the audio signal from the received data. Through hierarchical data processing and transmission control, it solves the problem of collaborative transmission of heterogeneous audio data from multiple devices. Especially in feature extraction and data fusion, it fully considers the differences in characteristics between different devices, effectively reducing data redundancy during transmission. Furthermore, by adopting a hierarchical modulation and resource allocation strategy, it ensures audio quality while reducing transmission power consumption, thus achieving efficient and low-power transmission of heterogeneous audio data overall. Attached Figure Description Figure 1 is a schematic diagram of the first embodiment of the low-power communication method for wireless audio devices in this invention; Figure 2 is a schematic diagram of an embodiment of a low-power communication device for wireless audio equipment according to the present invention. Detailed Implementation To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. To facilitate understanding of this embodiment, the specific process of this embodiment is described below. Please refer to Figure 1. The first embodiment of the low-power communication method for wireless audio devices in this embodiment includes: 101. Sample the acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points, and calculate the speech spectrum of each initial acoustic sampling point according to the corresponding device characteristics to obtain the frequency domain feature data corresponding to each target audio device. The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning. In this embodiment, the target audio device refers to various audio devices that transmit audio data, such as wireless headphones, smart speakers, microphone arrays, and portable speakers. The heterogeneous audio data refers to audio data from different types of audio devices with different characteristics (such as sampling rate, bit depth, and number of channels). For example, audio data collected by headphones may have a sampling rate of 48kHz and 16-bit quantization; data collected by a microphone array may have multiple channels and a sampling rate of 16kHz and 24-bit quantization; audio data output by a speaker may have a sampling rate of 44.1kHz and 32-bit quantization. Device characteristics refer to the inherent differences in hardware performance and audio acquisition and processing between different audio devices. Based on a preset device sampling rate, the heterogeneous audio data from multiple target audio devices is downsampled to obtain the original audio data. Based on a preset fixed sampling interval, the original audio data is framed to obtain initial audio data frames. Sound pressure level detection is performed on the initial audio data frames to obtain audio energy distribution data, and the audio energy distribution is further analyzed. The system performs threshold detection on the speech activity data to obtain speech segment data. Continuity detection and temporal labeling are then performed on the speech segment data to obtain audio frame sequence data. Based on the device identification code corresponding to each target audio device, the audio frame sequence data is labeled and sampled values ​​are quantized to obtain multiple initial acoustic sampling points. Mean and variance statistics are performed on the initial acoustic sampling points to obtain an audio statistical sequence. Based on a preset confidence interval and audio statistical parameters, a deviation threshold is calculated for the audio statistical sequence. Sampling points with amplitudes greater than the deviation threshold are extracted from the initial acoustic sampling points to obtain valid audio sampling points. Based on the frequency response coefficients of the target audio devices, the frequency domain transform values ​​of each valid audio sampling point are calculated to obtain an initial spectrum sequence. The energy distribution of each initial spectrum sequence in a preset frequency band is calculated, and frequency components greater than a preset device response interval are filtered to obtain a filtered spectrum sequence. The filtered spectrum sequence is then integrated with the corresponding audio sampling time and the device code corresponding to each target audio device to obtain frequency domain feature data corresponding to each target audio device. In practical applications, a preset sampling rate parameter is first set according to the sampling capabilities of different audio devices. For example, high-performance devices can be set to 48kHz, and low-performance devices to 16kHz. Then, the heterogeneous audio data collected by each device is downsampled according to the lowest sampling rate standard, thereby unifying the audio data from different devices to the same sampling rate standard and generating original audio data. Then, based on a preset fixed sampling interval (e.g., 20ms), these original audio data are segmented into frames, dividing the continuous audio data into data frames of equal length to form initial audio data frames. Then, by calculating the sound pressure level of each frame (i.e., converting the audio signal amplitude into sound pressure level in decibels), audio energy distribution data reflecting the energy distribution characteristics of each frame is obtained, and this data is then processed. The audio data is compared with a pre-set speech activity detection threshold (typically set in the 20-30 dB range) to filter out valid speech segments exceeding the threshold, generating speech segment data. Then, for the obtained speech segment data, the integrity and temporal sequence of the speech segments are determined by detecting the energy continuity and temporal relationship between adjacent frames. Simultaneously, a timestamp is added to each speech segment, generating audio frame sequence data with temporal information. Finally, based on the unique device identification code (such as device ID or MAC address) pre-assigned to each audio device, device identification information is added to the audio frame sequence data corresponding to each audio device, and the sampled values ​​are quantized according to the device's quantization precision (such as 16-bit quantization), ultimately obtaining multiple initial acoustic sampling points with device identification and temporal information. This achieves the orderly organization of multiple audio data streams and efficiently processes and transmits audio data from multiple devices. Secondly, by calculating the mean (reflecting the DC bias level of the signal) and variance (reflecting the degree of signal fluctuation) of each sampling sequence in the initial acoustic sampling points, an audio statistical sequence describing the overall statistical characteristics of the audio signal is formed. Then, based on a pre-set confidence interval (usually a 95% or 99% confidence level) and the obtained statistical parameters, a deviation threshold that can distinguish normal signals from outliers is calculated using statistical theory. Next, each sample value in the original initial acoustic sampling points is compared with the calculated deviation threshold, and sampling points with amplitudes greater than the threshold (where these sampling points contain valid audio information) are extracted, thus obtaining valid audio sampling points. Then, based on the frequency response characteristics of different audio devices (including the gain response curves of the device at different frequencies), a specific frequency response is determined for each audio device. The corresponding frequency response coefficients are used in the frequency domain conversion of effective audio sampling points. Methods such as Fast Fourier Transform are employed to convert the time-domain sampling points into a frequency-domain spectral sequence, yielding an initial spectral sequence. Following this initial spectral sequence, the energy distribution within pre-defined frequency bands (e.g., low-frequency band 20-200Hz, mid-frequency band 200-2000Hz, high-frequency band 2000-20000Hz) is calculated. Simultaneously, based on the preset effective frequency response range of each device (i.e., the frequency range in which the device can operate normally), frequency components exceeding this range are filtered out, resulting in a filtered spectral sequence. Finally, these filtered spectral sequences are integrated with the time information recorded during data acquisition and the unique identifier of each device to form a complete frequency domain feature data packet. This approach not only achieves effective audio data verification and frequency domain conversion but also, by considering the frequency response characteristics of the devices, ensures that the processing results better reflect the actual operating characteristics of the devices. Furthermore, the filtering and integration of frequency components reduces data transmission redundancy and improves processing efficiency. 102. Extract the acoustic features of multiple devices from the frequency domain feature data to obtain multiple acoustic feature sequences, and perform multi-frequency sub-band speech enhancement on each acoustic feature sequence to obtain multiple enhanced speech streams; In this embodiment, the energy distribution entropy values ​​of the frequency domain feature data within a preset multi-time window are calculated to obtain a frequency domain entropy sequence. The spectral flatness of the frequency domain feature data is also calculated to obtain a frequency domain flatness sequence. The modulation frequency components of the frequency domain feature data are extracted to obtain a frequency domain modulation feature sequence. The frequency domain entropy sequence, frequency domain flatness sequence, and frequency domain modulation feature sequence are then parameterized to construct an audio feature parameter table. A linear transformation is performed on the feature parameter table to obtain a reference audio feature set. Finally, a feature projection transformation is performed on the frequency domain audio feature data to obtain an acoustic feature sequence. The sequence is divided into frequency bands to obtain sub-band feature sequences, and the energy of the sub-band feature sequences is normalized to obtain multiple acoustic feature sequences. Sub-band noise power spectrum is estimated for the sub-band feature sequences in the acoustic feature sequences to obtain sub-band noise spectrum, and signal power spectrum is calculated for the sub-band feature sequences to obtain sub-band signal spectrum. Based on the sub-band noise spectrum, spectral subtraction is performed on each sub-band feature in the sub-band signal spectrum to obtain denoised sub-band signals. Adaptive gain and phase compensation are performed on the denoised sub-band signals to obtain enhanced sub-band signals, and adaptive sub-band synthesis is performed on each enhanced sub-band signal to obtain multiple enhanced speech streams. In practical applications, the energy distribution entropy value of the frequency domain feature data is first calculated within multiple preset time windows (such as analysis windows of different scales like 10ms, 20ms, and 30ms). This entropy value reflects the uncertainty and complexity of the signal in the frequency domain, thus obtaining a frequency domain entropy value sequence describing the time-varying characteristics of the signal. Simultaneously, spectral flatness analysis is performed on the frequency domain feature data to calculate the uniformity of energy distribution in each frequency band, resulting in a frequency domain flatness sequence characterizing the signal's spectral shape. Furthermore, modulation features of the signal are extracted from the frequency domain feature data using methods such as bandpass filtering. These modulation features reflect the periodic variation characteristics of the signal, forming a frequency domain modulation feature sequence. Finally, these three feature sequences of different dimensions (entropy value sequence, flatness sequence, and modulation feature sequence) are integrated according to preset feature weights and combination rules. An audio feature parameter table containing multidimensional feature information is constructed. Then, a linear transformation (such as principal component analysis or linear discriminant analysis) is applied to this parameter table to transform the original feature space into a new feature space, resulting in a more representative benchmark audio feature set. Using this feature set as a benchmark, feature projection transformation is performed on the original frequency domain feature data, mapping high-dimensional features to a low-dimensional feature space, resulting in a more compact acoustic feature sequence. Furthermore, based on the frequency characteristics of the audio signal, the entire frequency band is divided into multiple sub-bands according to preset frequency boundaries (such as critical bands or octaves). The features within each sub-band are processed separately to obtain sub-band feature sequences. Finally, energy normalization is performed on the feature sequences of each sub-band to eliminate energy differences between different sub-bands, making the features more comparable, thus obtaining a standardized multi-channel acoustic feature sequence. This approach considers the multidimensional characteristics of audio signals and improves the expressive power of features through feature transformation and normalization, providing a feature foundation for speech enhancement and data fusion. Secondly, noise power spectrum estimation is performed on the characteristic sequences within each sub-band. By analyzing the non-speech or low-energy segments of the signal, the background noise level in each sub-band is estimated, resulting in a sub-band noise spectrum reflecting noise characteristics. Simultaneously, power spectral density analysis is performed on the sub-band characteristic sequences to calculate the energy distribution of the signal at each frequency point, obtaining the sub-band signal spectrum characterizing the useful signal characteristics. Then, based on the sub-band noise spectrum, spectral subtraction is performed on the signal spectrum and noise spectrum within each sub-band, i.e., the estimated noise spectrum is subtracted from the signal spectrum. This process needs to consider the balance between over-subtraction and under-subtraction, typically using a soft thresholding method. Spectral subtraction yields the initially denoised subband signals. Then, based on the energy distribution and signal-to-noise ratio of each subband signal, corresponding gain compensation coefficients are calculated to adjust the amplitude of the denoised signal. Simultaneously, based on the signal's phase characteristics, phase compensation processing is performed to eliminate phase distortion that may be introduced during spectral subtraction, resulting in enhanced subband signals with improved amplitude and phase. Finally, based on the importance and energy distribution of each subband signal, appropriate subband synthesis weights are determined. A weighted overlapping addition method is used to weightedly synthesize all enhanced subband signals to reconstruct the complete speech signal. The weighted overlapping addition formula is as follows:

[0034] Wherein: g k h represents the weighting coefficient of the k-th sub-band. k (r) is the synthesis window function, s k (n) represents the subband signal, R is the frame shift, n is the time index or frame index, k is the index of the subband or channel, M is the length of the synthesis window function, and r is the sample index of the window function. For example, with 8 subbands and 50% overlap between adjacent frames, an improved triangular window is used for synthesis. Each subband is weighted according to its signal-to-noise ratio, resulting in a multi-channel enhanced speech stream with improved quality. Subband processing achieves refined noise suppression, and gain and phase compensation ensure the naturalness of the signal. The final subband synthesis guarantees the integrity of the enhanced speech, thus maintaining the intelligibility and naturalness of the speech signal while reducing noise interference. 103. Perform heterogeneous segmentation and grouping of each enhanced speech stream using multi-target audio devices to obtain speech stream grouping results. Then, calculate the check value and perform differential coding on the speech stream grouping results to obtain multi-source coded speech segments. In this embodiment, short-time energy and zero-crossing rate calculations are performed on each enhanced speech stream to obtain energy sequences and zero-crossing rate sequences. These sequences are then combined to obtain device feature parameters. The cross-correlation coefficients between these device feature parameters are calculated to generate a device correlation table. Based on this table, the audio feature distances between target audio devices are calculated to generate a device distance matrix. A hierarchical structure of the device distance matrix is ​​constructed to obtain a device grouping structure. The enhanced speech streams corresponding to each target audio device in the device grouping structure are indexed and grouped to obtain speech stream grouping results. The speech stream grouping results are then divided into blocks to obtain data block sequences. The mean of the data block sequences corresponding to different target audio devices within the same group is used to obtain common audio features. The signal-to-noise ratio (SNR) of the common audio features is calculated and weights are normalized to obtain audio device weights. Based on these weights, the common audio features are weighted and combined to obtain a fused feature vector. The fused feature vector is then subjected to adaptive quantization and encoding verification to obtain a multi-source coded speech segment. In practical applications, firstly, an energy sequence reflecting signal intensity changes is obtained by calculating the signal energy value within each short frame. Simultaneously, a zero-crossing rate sequence reflecting the signal frequency characteristics is obtained by calculating the number of times the signal crosses the zero level within each analysis frame. Then, these two sequences are weighted and fused according to a preset feature combination rule. The weighted fusion formula for the device characteristic parameters is as follows: Among them, E n Z represents the energy of the nth frame. n W represents the zero-crossing rate. n P is the weighting coefficient. nLet α, β, and γ be the pitch features, α, β, and γ be the adjustment coefficients, N be the total number of frames, d be the device identifier, t be the index of the time period, and F(d) be the feature parameter value of device d within the time period t. This allows us to calculate device feature parameters that characterize the audio characteristics of each device. Next, we calculate the cross-correlation coefficients between different device feature parameters. These correlation coefficients reflect the similarity between the signals acquired by different devices, thus forming a device correlation table describing the inter-device relationship. Based on this correlation table, we use metrics such as Euclidean distance or Mahalanobis distance to calculate the feature distance between each device, constructing a device distance matrix that characterizes the degree of difference between devices. Then, we apply hierarchical clustering and other methods to this distance matrix to construct a hierarchical structure, grouping devices with similar feature distances together. The system groups devices into a structure and assigns a unique index number to each group. The enhanced speech streams corresponding to each device are then grouped and labeled to obtain the speech stream grouping results. The grouped speech streams are then segmented according to a preset data block size to obtain easily processed data block sequences. The mean of the data block sequences from different devices within the same group is calculated to extract common audio features reflecting the shared characteristics of the devices in that group. The signal-to-noise ratio (SNR) of these common features is then calculated to evaluate the quality of the signals acquired by each device. The SNR values ​​are then normalized and converted into weights to represent the importance of different devices in feature fusion. Finally, based on these weights, the common audio features are weighted and combined to fuse the features from different devices within the same group. The weighted combination formula is as follows: Among them, X i (t) represents the time-domain characteristics of device i, H i (f) is the frequency response function, σ i G represents the variance. i (f) is the gain function, λ i For the fusion weights, N is the total number of devices, t is the time index, and i is the device index. This allows for the calculation of a fusion feature vector representing the overall characteristics of the group of devices. Finally, the fusion feature vector undergoes fixed-point quantization and is encoded with a checksum to obtain the final multi-source coded speech segment. This approach effectively organizes heterogeneous data from multiple devices, ensuring the rational utilization of each device's features. Furthermore, quantization encoding prepares the data for subsequent low-power transmission, preserving the characteristics of each device's signal while achieving effective data fusion. 104. Divide the multi-source coded speech segment into transmission units and perform coordinated signal modulation on each transmission unit to obtain a modulated audio stream. Then, perform coordinated resource allocation and data stream transmission on the modulated audio stream to obtain a transmitted speech data stream. In this embodiment, information content calculation and transmission level classification are performed on multi-source coded speech segments to obtain audio transmission priority. Based on the audio transmission priority, dynamic bit allocation and configuration of various cooperative transmission resource parameters are performed on the multi-source coded speech segments to generate a modulated audio stream. The modulation power consumption value of the preset modulation stage when modulating the modulated audio stream is calculated, and the modulation power consumption value is segmented into thresholds and resource allocation calculations to obtain an initial resource allocation value. The initial resource allocation value is integrated and allocated with various modulation allocation parameters to generate resource allocation parameters. The resource allocation parameters are then divided into power levels based on multi-device timing to obtain power configuration parameters. The power configuration parameters are used to calculate traffic parameters to obtain transmission data traffic parameters. Based on the transmission data traffic parameters and resource allocation parameters, the modulated audio stream is transmitted as a data stream to obtain a transmission voice data stream. In practical applications, the information entropy value of each coded segment is calculated to assess its information content. Then, based on the information content distribution, these coded segments are divided into different transmission levels to obtain audio transmission priorities that characterize the importance of the data. Bit allocation is then performed on the coded speech segments based on this priority, with higher-priority data receiving more bit resources. Simultaneously, transmission resource parameters, including modulation scheme, coding scheme, and transmit power, are configured. These parameters are combined to generate a modulated audio stream suitable for transmission. Power consumption is then calculated for each modulation stage (including signal modulation, power amplification, and channel coding) during the generation of the modulated audio stream, yielding a modulation power consumption value reflecting energy consumption. Finally, based on a preset power consumption threshold, [the following is a separate, unrelated step:] […]. These power consumption values ​​are categorized into different levels, and an initial resource allocation scheme is calculated based on available resources to obtain initial resource allocation values. These initial allocation values ​​are then integrated with multiple modulation parameters (such as modulation order, bit rate, and time slot length) to generate complete resource allocation parameters. Based on these parameters, the transmission timing of multiple devices is planned. Simultaneously, transmit power is categorized into different levels to obtain detailed power configuration parameters. Then, based on these power configuration parameters, the data flow rate of each transmission time slot is calculated to obtain transmission data flow parameters used to control the transmission rate. Finally, these transmission parameters are applied to the actual transmission process of the modulated audio stream. Data transmission is performed according to the allocated time slots, power, and bit resources to obtain the final transmitted voice data stream. Through multi-level parameter configuration and resource allocation, fine-grained control of the transmission process is achieved, ensuring the transmission quality of important data while realizing rational energy utilization and low-power communication. 105. Obtain the low-power transmission voice data stream and perform audio recovery on the transmission voice data stream to obtain the target audio data stream transmitted by multiple target audio devices. In this embodiment, the transmitted voice data stream is prioritized to obtain hierarchical voice data, and the hierarchical voice data is divided into buffer regions to obtain a data block location table. The access frequency of the buffered data in the data block location table is counted to obtain data access records, and the importance of the data access records is calculated to obtain a replacement order table. The occupancy status of the buffered data in the data block location table is detected to obtain the buffer occupancy value, and the buffer occupancy value is compared with a threshold to obtain buffer filtered data. The buffer filtered data is reassembled according to the replacement order table to obtain buffer reassembled data, and the buffer reassembled data is used for signal reconstruction to obtain the target audio data stream. In practical applications, the target communication device receives low-power transmission voice data streams and then classifies the data streams according to the priority identification information carried in the data packets, dividing data of different importance into different priority categories to obtain hierarchical voice data. Then, based on available cache resources and the capacity requirements of each level of data, corresponding cache areas are allocated for data of different priorities, establishing a mapping relationship between data blocks and storage locations to form a data block location table. Next, the access frequency of data blocks stored in each cache area is statistically analyzed, recording the number of times each data block is read and used, generating data access records reflecting data usage. Based on these access records, the importance value of each data block is calculated (wherein, this value comprehensively considers the data priority and access frequency, and is used to generate guidance for cache data). The system first generates a replacement order list, then checks the status of each cache region recorded in the data block location table, calculates the data occupancy of each region, obtains the cache occupancy value reflecting the cache usage status, and compares these occupancy values ​​with a preset cache capacity threshold to identify cache regions that need data replacement or merging, thus obtaining cache-filtered data. Next, according to the previously generated replacement order list, the filtered cache data is reorganized, retaining high-importance data in the fast access region, while moving less important data to the spare region or releasing its occupied space, resulting in optimized cache-reorganized data. Finally, signal reconstruction processing is performed on these reorganized data, including reassembling segmented data, restoring the timing relationship of the signal, and restoring the spectral characteristics of the audio signal, ultimately obtaining the complete target audio data stream. Therefore, through cache management and data reorganization, not only is fast access to important data ensured, but efficient utilization of storage resources is also achieved, realizing stable and reliable communication data recovery functionality. In this embodiment of the invention, heterogeneous data from multiple audio devices are sampled and spectral analyzed to obtain the frequency domain characteristics of each device. Then, the acoustic features of these frequency domain characteristics are extracted, and speech enhancement is performed through sub-band processing. Next, the enhanced speech stream is grouped and encoded by device. The encoded speech segments are then divided into transmission units and modulated, and data is transmitted according to a resource allocation strategy. Finally, the received data is used to recover the audio signal. Through hierarchical data processing and transmission control, the problem of collaborative transmission of heterogeneous audio data from multiple devices is solved. Particularly in feature extraction and data fusion, the differences in characteristics between different devices are fully considered, effectively reducing data redundancy during transmission. The hierarchical modulation and resource allocation strategy ensures audio quality while reducing transmission power consumption, thus achieving efficient and low-power transmission of heterogeneous audio data overall. The low-power communication method for wireless audio devices in the embodiments of the present invention has been described above. The low-power communication device for wireless audio devices in the embodiments of the present invention is described below. Referring to Figure 2, one embodiment of the low-power communication device for wireless audio devices in the embodiments of the present invention includes: The spectrum calculation module 201 is used to sample the acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points, and to perform speech spectrum calculation on each initial acoustic sampling point according to the corresponding device characteristics to obtain frequency domain feature data corresponding to each target audio device. The speech enhancement module 202 is used to extract multi-device acoustic features from the frequency domain feature data to obtain multiple acoustic feature sequences, and to perform multi-frequency sub-band speech enhancement on each of the acoustic feature sequences to obtain multiple enhanced speech streams. The segmentation and grouping module 203 is used to perform heterogeneous segmentation and grouping of each of the enhanced speech streams using multi-target audio devices to obtain speech stream grouping results, and to perform check value calculation and differential coding on the speech stream grouping results to obtain multi-source coded speech segments. The signal modulation module 204 is used to divide the multi-source coded speech segment into transmission units and perform coordinated signal modulation of each transmission unit to obtain a modulated audio stream, and to perform coordinated resource allocation and data stream transmission of the modulated audio stream to obtain a transmitted speech data stream. The audio recovery module 205 is used to acquire the low-power transmission voice data stream and perform audio recovery on the transmission voice data stream to obtain the target audio data stream transmitted by the multi-target audio device. In this embodiment of the invention, heterogeneous data from multiple audio devices are sampled and spectral analyzed to obtain the frequency domain characteristics of each device. Then, the acoustic features of these frequency domain characteristics are extracted, and speech enhancement is performed through sub-band processing. Next, the enhanced speech stream is grouped and encoded by device. The encoded speech segments are then divided into transmission units and modulated, and data is transmitted according to a resource allocation strategy. Finally, the received data is used to recover the audio signal. Through hierarchical data processing and transmission control, the problem of collaborative transmission of heterogeneous audio data from multiple devices is solved. Particularly in feature extraction and data fusion, the differences in characteristics between different devices are fully considered, effectively reducing data redundancy during transmission. The hierarchical modulation and resource allocation strategy ensures audio quality while reducing transmission power consumption, thus achieving efficient and low-power transmission of heterogeneous audio data overall. The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A low-power communication method for wireless audio devices, characterized in that, The low-power communication method for wireless audio devices includes: sampling the acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points, and performing speech spectrum calculations on each initial acoustic sampling point based on the corresponding device characteristics to obtain frequency domain feature data corresponding to each target audio device. The frequency domain feature data is subjected to multi-device acoustic feature extraction to obtain a multi-channel acoustic feature sequence, and the acoustic feature sequence is subjected to multi-frequency sub-band speech enhancement to obtain a multi-channel enhanced speech stream. The enhanced speech streams are heterogeneously segmented and grouped by multiple target audio devices to obtain speech stream grouping results. These grouping results are then divided into blocks to obtain data block sequences. The mean of the corresponding data block sequences from different target audio devices within the same group is calculated to obtain common audio features shared by the devices in the same group. The signal-to-noise ratio (SNR) of these common audio features is calculated and weighted and normalized to obtain audio device weights. Based on these audio device weights and a preset weighted combination formula, the corresponding data block sequences from different target audio devices within the same group are weighted and fused to obtain a fused feature vector. The fused feature vector is then subjected to adaptive quantization and encoding verification to obtain a multi-source coded speech segment. The multi-source coded speech segment is divided into transmission units and the signals of each transmission unit are coordinated and modulated to obtain a modulated audio stream. The modulated audio stream is then coordinated and resource-allocated and transmitted to obtain a transmitted speech data stream. The low-power transmission voice data stream is acquired, and audio recovery is performed on the transmission voice data stream to obtain the target audio data stream transmitted by the multi-target audio device.

2. The low-power communication method for wireless audio devices according to claim 1, characterized in that, The process of sampling the acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points includes: Based on a preset device sampling rate, heterogeneous audio data from multiple target audio devices are downsampled to obtain original audio data. Based on a preset fixed sampling interval, the original audio data is divided into frames to obtain initial audio data frames. Sound pressure level detection is performed on the initial audio data frame to obtain audio energy distribution data, and speech activity threshold detection is performed on the audio energy distribution data to obtain speech segment data; The speech segment data is subjected to continuity detection and temporal marking to obtain audio frame sequence data. Based on the device identification code corresponding to each target audio device, the audio frame sequence data is marked and sampled values ​​are quantized to obtain multiple initial acoustic sampling points.

3. The low-power communication method for wireless audio devices according to claim 1, characterized in that, The step of performing speech spectrum calculations on each of the initial acoustic sampling points according to the corresponding device characteristics to obtain frequency domain feature data corresponding to each of the target audio devices includes: The initial acoustic sampling points are subjected to mean and variance statistics to obtain an audio statistical sequence. Based on the preset confidence interval and audio statistical parameters, the deviation threshold of the audio statistical sequence is calculated. Extract the sampling points whose amplitude is greater than the deviation threshold from the initial acoustic sampling points to obtain effective audio sampling points, and calculate the frequency domain transformation value of each effective audio sampling point based on the corresponding frequency response coefficient of the target audio device to obtain the initial spectrum sequence; The energy distribution of each initial spectral sequence in a preset spectral band is calculated, and frequency components larger than the preset device response range are filtered out to obtain a filtered spectral sequence. The filtered spectral sequence is then integrated with the corresponding audio sampling time and the device code corresponding to each target audio device to obtain the frequency domain feature data corresponding to each target audio device.

4. The low-power communication method for wireless audio devices according to claim 1, characterized in that, The extraction of multi-device acoustic features from the frequency domain feature data to obtain a multi-channel acoustic feature sequence includes: The energy distribution entropy values ​​of the frequency domain feature data within a preset multi-time window are calculated to obtain a frequency domain entropy value sequence. The spectral flatness of the frequency domain feature data is calculated to obtain a frequency domain flatness sequence. The modulation frequency components of the frequency domain feature data are extracted to obtain a frequency domain modulation feature sequence. The frequency domain entropy value sequence, frequency domain flatness sequence, and frequency domain modulation feature sequence are then integrated with parameters to construct an audio feature parameter table. A linear transformation is performed on the feature parameter table to obtain a reference audio feature set, and a feature projection transformation is performed on the frequency domain feature data to obtain an acoustic feature sequence. The acoustic feature sequence is divided into frequency bands to obtain sub-band feature sequences, and the energy of the sub-band feature sequences is normalized to obtain multi-channel acoustic feature sequences.

5. The low-power communication method for wireless audio devices according to claim 4, characterized in that, The step of performing multi-frequency sub-band speech enhancement on each of the acoustic feature sequences to obtain a multi-stream enhanced speech stream includes: The sub-band noise power spectrum of the sub-band feature sequence in the acoustic feature sequence is estimated to obtain the sub-band noise spectrum, and the signal power spectrum of the sub-band feature sequence is calculated to obtain the sub-band signal spectrum; Based on the subband noise spectrum, spectral subtraction is performed on each subband feature in the subband signal spectrum to reduce noise and obtain the denoised subband signal. Adaptive gain and phase compensation are applied to the noise-reduced subband signal to obtain the enhanced subband signal, and adaptive subband synthesis is performed on each of the enhanced subband signals to obtain a multi-channel enhanced speech stream.

6. The low-power communication method for wireless audio devices according to claim 1, characterized in that, The process of heterogeneously segmenting and grouping the enhanced speech streams using multiple target audio devices to obtain speech stream grouping results includes: Short-time energy calculation and zero-crossing rate calculation are performed on each of the enhanced speech streams to obtain energy sequences and zero-crossing rate sequences. Then, feature combination is performed on the energy sequences and the zero-crossing rate sequences to obtain device feature parameters. Calculate the cross-correlation coefficients between the characteristic parameters of each device, generate a device correlation table, and calculate the audio feature distance between each target audio device based on the device correlation table, generating a device distance matrix; A hierarchical structure of the device distance matrix is ​​constructed to obtain a device grouping structure. The enhanced speech streams corresponding to each target audio device in the device grouping structure are indexed, assigned, and grouped to obtain speech stream grouping results.

7. The low-power communication method for wireless audio devices according to claim 1, characterized in that, The process of dividing the multi-source coded speech segment into transmission units and coordinating signal modulation of each transmission unit to obtain a modulated audio stream, and then coordinating resource allocation and data stream transmission of the modulated audio stream to obtain a transmitted speech data stream, includes: The information content of the multi-source coded speech segments is calculated and the transmission level is divided to obtain the audio transmission priority. Based on the audio transmission priority, dynamic bit allocation and configuration of various cooperative transmission resource parameters are performed on the multi-source coded speech segments to generate a modulated audio stream. The modulation power consumption value of the preset modulation stage when the modulated audio stream is modulated is calculated, and the modulation power consumption value is divided into segments with thresholds and the resource allocation is calculated to obtain the initial resource allocation value. The initial resource allocation value is integrated and allocated using multiple modulation allocation parameters to generate resource allocation parameters. The resource allocation parameters are then divided into power levels based on multi-device timing to obtain power configuration parameters. The power configuration parameters are used to calculate the flow parameters to obtain the transmission data flow parameters. Based on the transmission data flow parameters and the resource allocation parameters, the modulated audio stream is transmitted as a data stream to obtain the transmission voice data stream.

8. A low-power communication device for wireless audio equipment, characterized in that, The low-power communication device for the wireless audio equipment includes: The spectrum calculation module is used to sample the acoustic features of heterogeneous audio data from multiple target audio devices to obtain multiple initial acoustic sampling points, and to perform speech spectrum calculation on each initial acoustic sampling point according to the corresponding device characteristics to obtain the frequency domain feature data corresponding to each target audio device. The speech enhancement module is used to extract the acoustic features of multiple devices from the frequency domain feature data to obtain multiple acoustic feature sequences, and to perform speech enhancement on each of the acoustic feature sequences in multiple frequency sub-bands to obtain multiple enhanced speech streams. The segmentation and grouping module is used to perform heterogeneous segmentation and grouping of the enhanced speech streams into multiple target audio devices to obtain speech stream grouping results. The speech stream grouping results are then divided into blocks to obtain data block sequences. The mean of the corresponding data block sequences from different target audio devices within the same group is calculated to obtain common audio features shared by the devices in the same group. The signal-to-noise ratio (SNR) of these common audio features is calculated and the weights are normalized to obtain audio device weights. Based on the audio device weights and a preset weighted combination formula, the corresponding data block sequences from different target audio devices within the same group are weighted and fused to obtain a fused feature vector. The fused feature vector is then subjected to adaptive quantization and encoding verification to obtain a multi-source coded speech segment. The signal modulation module is used to divide the multi-source coded speech segment into transmission units and perform coordinated signal modulation of each transmission unit to obtain a modulated audio stream, and to perform coordinated resource allocation and data stream transmission of the modulated audio stream to obtain a transmitted speech data stream. The audio recovery module is used to acquire the transmitted voice data stream of low-power transmission and perform audio recovery on the transmitted voice data stream to obtain the target audio data stream transmitted by multiple target audio devices.

9. A SoC chip, characterized in that, The SoC chip includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the SoC chip to perform the steps of the low-power communication method for a wireless audio device as described in any one of claims 1-7.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the steps of the low-power communication method for wireless audio devices as described in any one of claims 1-7.