A method, system and storage medium for processing Bluetooth speaker data

Through multimodal data processing and ambient acoustic compensation technology, the personalization and environmental adaptability problems of traditional Bluetooth speakers are solved, personalized audio processing and environmental adaptation are achieved, and the sound quality and user experience of the speakers are improved.

CN119601023BActive Publication Date: 2025-07-18SHENZHEN AOOLIF TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411477357.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

Traditional Bluetooth speakers lack personalized music preference recognition and environmental adaptability, which makes audio data processing unable to be optimized for different users and environments, affecting the listening experience.

Method used

Through multimodal data capture, adaptive multi-resolution compression, multimodal data decoupling, dynamic spectrum reconstruction optimization, user music preference recognition and environmental acoustic feature database construction, personalized processing and environmental compensation of audio data are realized.

Benefits of technology

Provides stable, clear and rich personalized audio experience, adapts to various environments, and improves user satisfaction and speaker performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601023B_ABST
    Figure CN119601023B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of Bluetooth speakers, and particularly to a method, system and storage medium for Bluetooth speaker data processing. The method includes the following steps: capturing multi-modal data of a client to obtain original multi-modal client data; performing adaptive multi-resolution compression on the original multi-modal client data to obtain compressed multi-modal client data; transmitting the compressed multi-modal client data to a Bluetooth speaker through the client to obtain received compressed multi-modal data; decoupling the received compressed multi-modal data to obtain separated audio data; performing dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data; determining the music preference type of a user to obtain a user music preference type identifier. The present invention effectively solves the limitations of traditional Bluetooth speakers in terms of personalized services and environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Bluetooth speakers, and in particular to a method, system and storage medium for processing Bluetooth speaker data. Background Art

[0002] In traditional Bluetooth speakers, a significant technical difficulty is the lack of consideration for users' personalized music preferences, resulting in no targeted optimization of audio data processing. This means that different users will get different listening experiences when listening to the same music because the speakers cannot recognize and adapt to personal preferences. This lack of personalized optimization limits the performance of speakers in meeting personalized needs. Due to the lack of personalized optimization, speakers use a unified standard when processing audio data and cannot dynamically adjust to the preferences of different users. This results in the inability of speakers to perform targeted optimization of audio signals, such as enhancing certain frequency responses to meet the listening needs of specific music styles. In addition, the acoustic environment in which the speaker is located has a significant impact on the sound quality, but traditional speakers often ignore this. Different environments, such as indoors, outdoors, open or confined spaces, will have different effects on audio signals. If the speaker cannot adjust the audio output according to the acoustic environment in which it is located, it cannot guarantee the best listening experience in all environments. Summary of the invention

[0003] Based on this, it is necessary for the present invention to provide a method, system and storage medium for processing Bluetooth speaker data to solve at least one of the above technical problems.

[0004] To achieve the above object, a method for processing Bluetooth speaker data includes the following steps:

[0005] Step S1: Capturing multimodal data of a client to obtain original multimodal client data; performing adaptive multi-resolution compression on the original multimodal client data to obtain compressed multimodal client data;

[0006] Step S2: transmitting the compressed multimodal client data to the Bluetooth speaker through the client to obtain received compressed multimodal data; performing multimodal data decoupling on the received compressed multimodal data to obtain separated audio data; and performing dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data;

[0007] Step S3: determining the music preference type of the user to obtain a user music preference type identifier; performing phase response on the reconstructed audio data according to the user music preference type identifier to obtain phase optimized audio data;

[0008] Step S4: Construct an environmental acoustic feature database for the environment where the Bluetooth speaker is located to obtain the environmental acoustic feature database; perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustic feature database to obtain the final optimized audio data.

[0009] In the present invention, by performing adaptive multi-resolution compression on multi-modal data, the compression ratio of data transmission can be dynamically adjusted according to the actual network conditions, the size of data packets can be optimized, and data loss and distortion caused by network fluctuations can be reduced. This not only improves the robustness of data transmission but also ensures the clarity and stability of sound quality, and high audio quality can be maintained even under poor network conditions. Through the multi-modal data decoupling and dynamic spectrum reconstruction optimization technology, audio data can be more accurately separated and processed, so as to more realistically restore the audio signal. This enables the played audio signal to better retain the details and dynamic range of the original sound, providing a richer and more stereoscopic auditory experience, especially significantly improving the layering and details of music. By identifying the personalized music preferences of users and performing corresponding phase response optimization, customized audio output can be provided for different users. This means that each user can obtain the best listening experience according to their preferences, thereby improving the adaptability of the Bluetooth speaker and the satisfaction of users. By constructing an environmental acoustic feature database and performing acoustic environment compensation on the audio signal according to this data, the speaker can automatically adjust the audio output in different environments to adapt to the current acoustic conditions. This intelligent compensation mechanism ensures a consistent high-quality audio experience in various environments. Whether indoors, outdoors, in open or enclosed spaces, users can enjoy the optimized sound quality. In summary, the present invention effectively solves the limitations of traditional Bluetooth speakers in data transmission, audio processing, personalized services, and environmental adaptability, brings a more stable, clear, rich, and personalized audio experience to users, and at the same time improves the performance of Bluetooth speakers in different environments.

[0010] Preferably, the present invention also provides a system for processing Bluetooth speaker data for executing the method for processing Bluetooth speaker data as described above. The system for processing Bluetooth speaker data includes:

[0011] A data compression module for capturing multi-modal data from the client to obtain the original multi-modal client data; performing adaptive multi-resolution compression on the original multi-modal client data to obtain the compressed multi-modal client data;

[0012] An audio reconstruction module for transmitting the compressed multi-modal client data to the Bluetooth speaker through the client to obtain the received compressed multi-modal data; decoupling the multi-modal data of the received compressed multi-modal data to obtain the separated audio data; performing dynamic spectrum reconstruction optimization on the separated audio data to obtain the reconstructed audio data;

[0013] An audio phase optimization module is used to determine the type of music preference of a user, obtaining a user music preference type identifier; and perform phase response on the reconstructed audio data according to the user music preference type identifier to obtain phase-optimized audio data.

[0014] An environmental acoustics compensation module is used to construct an environmental acoustics feature database for the environment where the Bluetooth speaker is located, obtaining the environmental acoustics feature database; and perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustics feature database to obtain the final optimized audio data.

[0015] In the present invention, the data compression module can dynamically adjust the compression strategy according to the network condition and data characteristics by adopting the adaptive multi-resolution compression technology. This intelligent compression method not only reduces the bandwidth requirement for data transmission, but also improves the efficiency of data transmission, ensuring stable transmission of high-quality audio data in various network environments. The audio reconstruction module can effectively recover and enhance the audio signal by decoupling and dynamically optimizing the spectrum reconstruction of the received compressed data. This process not only reduces the distortion that occurs during data transmission, but also improves the clarity and dynamic range of the audio signal, making the played audio closer to the original sound quality and providing a richer and more delicate auditory experience. The audio phase optimization module can provide a customized audio processing solution for each user by analyzing the user's listening history and preferences. This personalized processing method enables the speaker to adjust the phase response of the audio according to the user's preferences, so that each user can obtain the audio experience that best suits their personal taste, enhancing user satisfaction and the intelligent level of the speaker. The environmental acoustics compensation module can intelligently identify the acoustic environment where the speaker is located and compensate the audio signal according to the environmental characteristics by constructing the environmental acoustics feature database. This compensation mechanism enables the speaker to automatically adjust the audio output in different environments to adapt to the current acoustic conditions, ensuring a consistent high-quality audio experience in various environments such as indoors, outdoors, open or enclosed spaces. Through the collaborative work of each module, an end-to-end intelligent processing from data capture, compression, transmission, decoupling, reconstruction, optimization to compensation is realized. This end-to-end intelligent processing not only improves the efficiency and quality of audio processing, but also enables the system to better adapt to different usage scenarios and user requirements, enhancing the flexibility of the system and the user experience.

[0016] Preferably, a computer-readable storage medium stores a method for processing Bluetooth speaker data that can be loaded and executed by a processor as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description with reference to the accompanying drawings:

[0018] Figure 1 The figure shows a schematic flowchart of the steps of a method for processing Bluetooth speaker data in an embodiment.

[0019] Figure 2 The figure shows a detailed schematic flowchart of step S23 in an embodiment.

[0020] Figure 3 The figure shows a detailed schematic flowchart of step S25 in an embodiment. Detailed implementation manners

[0021] The following clearly and completely describes the technical method of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

[0022] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and thus repeated descriptions of them will be omitted. Some of the block diagrams shown in the figures are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0023] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0024] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a method for processing Bluetooth speaker data, including the following steps:

[0025] Step S1: Perform multimodal data capture on the client to obtain original multimodal client data; perform adaptive multi-resolution compression on the original multimodal client data to obtain compressed multimodal client data;

[0026] Step S2: Transmit the compressed multimodal client data to the Bluetooth speaker through the client to obtain the received compressed multimodal data; decouple the received compressed multimodal data to obtain separated audio data; perform dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data;

[0027] Step S3: Determine the user's music preference type to obtain the user's music preference type identifier; perform phase response on the reconstructed audio data according to the user's music preference type identifier to obtain phase-optimized audio data;

[0028] Step S4: Construct an environmental acoustic feature database for the environment where the Bluetooth speaker is located to obtain the environmental acoustic feature database; perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustic feature database to obtain the final optimized audio data.

[0029] In this embodiment, first, a multifunctional data collector, such as a smartphone or a dedicated audio processing device, can be used to capture the user's multimodal data. This includes audio signals, user interaction data, and device status information. Using the device's built-in microphone and sensors, audio signals can be collected in real time, and user interaction and device status data can be obtained through the network interface. Next, an adaptive multi-resolution compression technology, such as a compression algorithm based on wavelet transform, is used to compress the original multimodal data. This algorithm can dynamically adjust the compression ratio according to the data content and network conditions to optimize the data transmission efficiency. For example, MATLAB or the PyWavelets library in Python can be used to implement this compression process. Then, the compressed multimodal data is wirelessly transmitted to the Bluetooth speaker through the Bluetooth module of the client device. At the speaker end, an audio decoder is used to receive and decouple the compressed data to separate the audio data. This involves applying specific decompression algorithms, such as Huffman decoding or LZW decoding. Next, dynamic spectrum reconstruction optimization is performed on the separated audio data. This can be achieved by using the short-time Fourier transform (STFT) to analyze the spectral content of the audio signal and applying spectral repair techniques, such as spectral subtraction or spectral filling, to enhance the spectral characteristics of the audio signal. According to the user's music preference type identifier, phase response optimization is performed on the reconstructed audio data. For example, if the user prefers music with a strong rhythm, the sense of rhythm can be enhanced by enhancing the phase consistency of the low-frequency components. Then, an environmental acoustic feature database is constructed. This can be achieved by deploying microphone arrays at different locations and using acoustic simulation software such as Raynoise for ray tracing to obtain spatial acoustic parameters. Finally, acoustic environment compensation is performed on the phase-optimized audio data according to the environmental acoustic feature database. This involves applying filters, such as equalizers or dynamic range compressors, to adjust the audio signal to adapt to a specific listening environment. Through this series of steps, the final optimized audio data can be obtained.

[0030] Preferably, step S1 includes the following steps:

[0031] Step S11: Perform multi-modal data capture on the client to obtain raw multi-modal client data;

[0032] Specifically, audio data can be captured in real time through the microphone and sensors of the client, and control instructions and device status updates can be received through the user interface. During this process, the audio signal is stored in the form of a waveform, the control instructions are stored in the form of command codes, and the device status is stored in the form of a status report. All these data will be packed together to form a raw multi-modal data set.

[0033] Step S12: Separate the modal features of the raw multi-modal client data to obtain client multi-modal feature data; wherein, the client multi-modal feature data includes audio data, control instruction data, and device status data;

[0034] Specifically, spectrum analysis technology can be used to identify audio data, a command parsing algorithm can be used to extract control instruction data, and status monitoring logic can be used to collect device status data. After separation, client multi-modal feature data is obtained, wherein the client multi-modal feature data includes audio data, control instruction data, and device status data.

[0035] Step S13: Perform run-length encoding and differential pulse coding on the control instruction data to obtain compressed control instruction data;

[0036] Specifically, run-length encoding (RLE) can be performed on the control instruction data first. This is a simple form of encoding that replaces a sequence of consecutive identical data values with a single value and a count value. For example, if there is a control instruction in the control instruction data that is ten consecutive "on" commands, RLE will encode it as "on x10". Next, differential pulse code modulation (DPCM) is applied to the data. This is a technique that encodes using the differences between signal values rather than absolute values. DPCM calculates the differences between consecutive samples and then encodes these differences to achieve data compression, ultimately obtaining compressed control instruction data.

[0037] Step S14: Perform change-detection-based incremental encoding on the device status data to obtain compressed device status data;

[0038] Specifically, the device status parameters such as volume, battery level, and connection status can be continuously monitored. When a status change is detected, the values before and after the change are recorded, and the difference between the two is calculated. For example, if the volume is adjusted from 75% to 80%, this change is recorded and an incremental code is generated indicating that the volume has increased by 5%. This change-based coding method only stores the information of the status change, rather than the absolute status values at each time point, thus significantly reducing the size of the data and finally obtaining compressed device status data.

[0039] Step S15: Perform multi-scale signal decomposition on the audio data to obtain client multi-scale audio data, and perform wavelet transform on the client multi-scale audio data to obtain feature-enhanced audio data;

[0040] Specifically, the AudioProcessor tool can be used for multi-scale signal decomposition. This tool integrates the functions of multi-scale decomposition and wavelet transform. First, the tool performs multi-scale decomposition on the audio data, which involves decomposing the signal into sub-bands of different frequency bands, and each sub-band represents a specific scale of the signal. This can be achieved through wavelet transform, where an appropriate wavelet basis function is selected to match the characteristics of the audio signal. For example, the Daubechies wavelet can be used, which is suitable for the decomposition of audio signals because it can well capture the time-frequency characteristics of the signal. During the multi-scale decomposition process, the audio signal is decomposed into a series of approximation coefficients and detail coefficients, which represent the low-frequency and high-frequency components of the signal respectively. Next, the tool performs wavelet transform on these coefficients to further extract the features of the audio signal. Wavelet transform can enhance the important features of the signal, such as sharp transients and smooth background noise. In this way, finally, feature-enhanced audio data can be obtained.

[0041] Step S16: Perform enhanced Huffman coding on the feature-enhanced audio data to obtain compressed audio data;

[0042] Specifically, enhanced Huffman coding is a lossless compression algorithm that assigns a unique code to each symbol in the data by constructing a variable-length coding tree. First, the occurrence frequency of each symbol (e.g., sampling value) in the feature-enhanced audio data can be counted. Then, a Huffman tree is constructed based on these frequencies, where the least frequent symbol is assigned the longest code, and the most frequent symbol is assigned the shortest code. For example, if a specific frequency component in the audio data is very common, it will be assigned a short code such as "01", while an uncommon component is assigned a long code such as "1110". Through coding, finally, compressed audio data is obtained.

[0043] Step S17: Perform embedded multimodal fusion on the compression control instruction data, compression device status data, and compressed audio data to obtain compressed multimodal client data.

[0044] Specifically, an embedded coding scheme such as Embedded Huffman Coding (ECC) can be used to fuse the compression control instruction data, compression device status data, and compressed audio data. ECC allows multiple data streams to be merged into a single data stream without destroying the integrity of individual data streams. For example, assuming there is an audio data stream, ECC can be used to embed control instructions and device status data in the coding tree of the audio data. In this way, there will be a flag next to each audio data block indicating whether there is control instruction or device status data associated with it. If so, this data will be appended to the corresponding audio data block. Finally, the fused data stream will be packaged into a compressed multimodal client data packet, which can be effectively transmitted to a Bluetooth speaker or other target devices.

[0045] In the present invention, by respectively performing run-length coding, differential pulse coding, and change-detection-based incremental coding on the control instruction data and device status data, the size of the data is significantly reduced while key information is retained. Such a coding method reduces redundancy during data transmission, improving the efficiency and reliability of data transmission. Through the application of multi-scale signal decomposition and wavelet transform, the audio data can be feature-enhanced before compression, which not only improves the quality of the audio data but also enables the compressed audio data to better retain the important features of the original audio, thus providing a better audio experience after decompression. Through embedded multimodal fusion, the compressed control instruction data, device status data, and audio data are effectively combined to form compressed multimodal client data. This fusion method not only reduces the data storage space but also simplifies the data transmission and processing processes, improving the overall system processing efficiency. Through targeted compression strategies, the present invention reduces the demand for storage and processing resources. This is particularly beneficial for resource-constrained devices as it allows these devices to store and process more data without sacrificing data quality or functionality due to resource limitations.

[0046] Preferably, step S2 includes the following steps:

[0047] Step S21: Transmit the compressed multimodal client data to the Bluetooth speaker through the client and receive the data to obtain the received compressed multimodal data;

[0048] Specifically, a client device can be used to package the compressed multimodal data into a format suitable for Bluetooth transmission. This involves splitting the data into small chunks and adding the necessary Bluetooth transmission protocol header information to each chunk to ensure that the data can be correctly routed and transmitted in the Bluetooth network. Then, a Bluetooth protocol stack (such as Bluetooth Classic or Bluetooth Low Energy) is used to establish a connection with the Bluetooth speaker. Once the connection is established, the data is sent to the speaker via a wireless signal. Upon reaching the Bluetooth speaker, the Bluetooth module on the speaker will receive these signals and use the corresponding Bluetooth protocol stack to parse and reconstruct the data, ultimately obtaining the complete received compressed multimodal data.

[0049] Step S22: Decouple the received compressed multimodal data to obtain separated audio data, and perform time-frequency feature extraction on the separated audio data to obtain audio time-frequency feature data;

[0050] Specifically, a MultiModalDecoder (multimodal decoder) tool can be used to process the received compressed multimodal data. First, the multimodal decoder will identify and separate the different data streams in the compressed multimodal data, such as audio data, control instruction data, and device status data. This involves parsing the tags in the embedded encoding scheme to determine which parts belong to the audio data. Once the audio data is separated, the tool will decode it to restore its original form. After decoding, an AudioFeatureExtractor (audio feature extractor) tool is used to extract the time-frequency features of the audio data. This tool uses the short-time Fourier transform (STFT) to convert the audio signal into a time-frequency representation, thereby extracting features such as frequency, amplitude, and phase. STFT achieves this by splitting the audio signal into short-time frames and performing a fast Fourier transform (FFT) on each frame. From the FFT results of each frame, key time-frequency features such as Mel frequency cepstral coefficients (MFCCs) and chroma features are extracted, ultimately obtaining the audio time-frequency feature data.

[0051] Step S23: Adjust the length of the preset initial analysis window according to the audio time-frequency feature data to obtain an adjusted window length parameter;

[0052] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of Step S23.

[0053] Step S24: Perform a short-time Fourier transform on the separated audio data according to the adjusted window length parameter to obtain an initial spectrogram;

[0054] Specifically, the separated audio data can be loaded. Then, a window function, such as a Hamming window or a Hanning window, is applied to window the audio signal. Next, according to the adjusted window length parameter calculated in step S23, the windowed audio signal is segmented into short-time frames. For example, if the adjusted window length parameter is 25 milliseconds, then a segment of the audio signal is taken every 25 milliseconds for processing. For each short-time frame, a Fast Fourier Transform (FFT), which is an efficient algorithm for calculating the spectrum of a signal, is performed. The FFT transforms each short-time frame from the time domain to the frequency domain, generating corresponding spectral data. These spectral data represent the amplitudes of different frequency components in each short-time frame. Finally, the spectral data of all short-time frames are aggregated to form a two-dimensional initial spectrogram. The horizontal axis represents time, the vertical axis represents frequency, and the color or brightness represents the amplitude at different times and frequencies. For example, assume there is an audio signal with a length of 1 second and a sampling rate of 44.1 kHz. In step S23, the calculated adjusted window length parameter is 25 milliseconds, and the corresponding number of FFT points is 1024. Then, 1024 samples are taken every 25 milliseconds for FFT, generating a series of spectral data of short-time frames, and finally forming an initial spectrogram that shows the time-frequency characteristics of this 1-second audio signal.

[0055] Step S25: Identify the key frequency bands of the initial spectrogram to obtain audio key frequency band data, and perform high spectral density analysis on the audio key frequency bands in the separated audio data according to the audio key frequency band data to obtain spectral data when the key frequency bands are enhanced;

[0056] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S25.

[0057] Step S26: Update the spectrogram of the initial spectrogram according to the spectral data when the key frequency bands are enhanced to obtain an optimized spectrogram;

[0058] Specifically, an initial spectrogram can be loaded. Meanwhile, the spectral data during key frequency band enhancement is also loaded. These data are obtained by performing key frequency band identification and enhancement processing on the initial spectrogram in step S25, highlighting the most important frequency bands in the audio signal. Next, an enhancement algorithm is applied to each key frequency band in the initial spectrogram. For example, a gain function based on an auditory psychology model is used, which can adjust the amplitude of the spectrum according to the sensitivity of the human ear to different frequencies. For each key frequency band, its amplitude is increased so that these frequency bands are more prominent in the final spectrogram. After processing all key frequency bands, the initial spectrogram is updated to generate an optimized spectrogram. For example, assume that the low-frequency and high-frequency parts of the audio signal are identified as key frequency bands and enhanced in step S25. In this step, the initial spectrogram is loaded, and the gain function and dynamic range compression are applied to these key frequency bands. After processing, a new spectrogram is generated, in which the amplitudes of the low-frequency and high-frequency parts are increased, and the dynamic range of the entire spectrum is effectively controlled.

[0059] Step S27: Perform harmonic fundamental frequency analysis on the optimized spectrogram to obtain fundamental frequency harmonic structure data, and reconstruct the fundamental frequency harmonic structure data to obtain reconstructed audio data.

[0060] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S27.

[0061] In the present invention, the compressed multi-modal data is transmitted to the Bluetooth speaker through the client, reducing the bandwidth required during wireless transmission of the data, thereby improving the transmission efficiency and ensuring fast and stable transmission of the data. Through multi-modal data decoupling, the audio data can be accurately separated, and the key time-frequency features can be extracted. This improves the accuracy of audio processing. By dynamically adjusting the length of the analysis window according to the audio time-frequency feature data, the short-time Fourier transform better conforms to the actual characteristics of the audio signal. This optimized window setting can more accurately capture the details of the audio signal and improve the accuracy of spectrum analysis. By performing key frequency band identification and high spectral density analysis on the initial spectrogram, the most important parts of the audio can be highlighted. This enhances the key frequency bands, thus significantly improving the quality and perceptual effect of the audio. Based on the spectral data during key frequency band enhancement, the initial spectrogram is updated in real time to generate an optimized spectrogram. This dynamic update mechanism enables audio processing to adapt to different audio contents and characteristics, providing more flexible and accurate processing results. By performing harmonic fundamental frequency analysis on the optimized spectrogram and reconstructing the fundamental frequency harmonic structure data, a high-quality audio signal can be restored. This reconstruction process retains the key information of the audio, making the reconstructed audio signal have better sound quality and higher fidelity.

[0062] Preferably, step S23 includes the following steps:

[0063] Step S231: Perform spectral centroid recognition on the audio time-frequency feature data to obtain a spectral centroid distribution map, and perform time-varying spectral feature recognition on the audio time-frequency feature data according to the spectral centroid distribution map to obtain an audio time-varying spectral feature matrix;

[0064] Specifically, the audio time-frequency feature data can be loaded, and the spectral centroid is calculated for the spectral data of each short time frame. The spectral centroid is the weighted average of the spectrum, where the weight is the amplitude of each frequency component. The calculation formula for the spectral centroid is: where f i is the frequency of the i-th frequency component, P i is the corresponding amplitude, and N is the total number of frequency components; after calculating the spectral centroids of all frames, a spectral centroid distribution map is generated, which shows the spectral centroid values changing with time. Then, analyze this distribution map to identify time-varying spectral features. This involves finding the local maxima and minima of the spectral centroid, or identifying rapid changes in the spectral centroid, which indicate important events or structural changes in the audio signal. Through these analyses, an audio time-varying spectral feature matrix can be constructed, which contains the spectral centroid changing with time and other related features.

[0065] Step S232: Perform main frequency period recognition on the audio time-frequency feature data according to the audio time-varying spectral feature matrix to obtain a main frequency period sequence;

[0066] Specifically, the audio time-varying spectral feature matrix can be loaded, and the spectral data of each short time frame is analyzed to identify the most prominent frequency component, i.e., the main frequency. This involves finding the frequency component with the largest amplitude in each frame, or using statistical methods to determine the most stable and prominent frequency in multiple consecutive frames. Once the main frequency is identified, track the changes of these frequencies over time to determine the main frequency period. This involves using an autoregressive (AR) model or a periodogram method to estimate the periodicity of the frequency. By analyzing the appearance and disappearance of the main frequency in consecutive frames, a main frequency period sequence can be constructed, which describes the main rhythm and melody structure in the audio signal. For example, if the audio signal is a music segment containing vocals and musical instruments, identify the fundamental frequency of the vocals and the main harmonics of the musical instruments as the main frequencies, and track the changes of these frequencies over time, thereby generating a main frequency period sequence reflecting the music rhythm and melody structure.

[0067] Step S233: Calculate the time-scale ratio of the main frequency period sequence to the preset initial analysis window length to obtain a window-signal time-scale ratio coefficient;

[0068] Specifically, the main frequency cycle sequence can be loaded. At the same time, the preset initial analysis window length is also obtained, which is the optimal value determined based on standard practices in audio analysis or previous experiments. Next, each main frequency cycle is compared with the initial analysis window length. The formula for calculating the time scale ratio is as follows: This scale factor reflects the relative size of the main frequency cycle and the analysis window length, and can be used to adjust the analysis window to better fit the characteristics of the audio signal. For example, if the main frequency cycle is much smaller than the initial analysis window length, it indicates that a smaller window is needed to more precisely capture the details of the audio signal.

[0069] Step S234: Evaluate the transient capture performance of the preset initial analysis window length according to the window-signal time scale ratio coefficient to obtain a transient capture performance index;

[0070] Specifically, the window-signal time scale ratio coefficient can be loaded and analyzed to evaluate the impact of different window lengths on transient capture. Transient capture performance evaluation involves simulating the impact of different window lengths on transient features in the audio signal, such as the percussion sounds of percussion instruments, sudden volume changes, etc. A set of predefined evaluation criteria, such as the transient response index or signal distortion metric, are used to evaluate the transient capture performance of different window lengths. For example, calculate the ratio of the peak value to the average value of the transient signal at different window lengths, or calculate the ratio of the peak value of the transient signal to the window length, to evaluate the accuracy and efficiency of transient capture. Through these evaluations, a transient capture performance index is generated, which reflects the impact of different window lengths on the ability to capture transient features. For example, if the transient capture performance index indicates that a smaller window length can better capture transient features, then a smaller window length is selected for subsequent audio analysis.

[0071] Step S235: Determine the window length adjustment based on the transient capture performance index for the preset initial analysis window length to obtain a window length adjustment factor, and perform dynamic range compression on the preset initial analysis window length according to the window length adjustment factor to obtain a preliminary adjusted window length;

[0072] Specifically, the transient capture performance index can be loaded and analyzed to determine whether the initial analysis window length needs to be adjusted. For example, if the transient capture performance index shows that a smaller window length can more effectively capture transient features, calculate a window length adjustment factor. This factor can be calculated by the following formula: Among them, the optimal window length is the length recommended based on the transient capture performance index. Next, apply the window length adjustment factor to the initial analysis window length for dynamic range compression. The process of dynamic range compression involves adjusting the amplitude of the audio signal to ensure that the dynamic range of the signal can be effectively controlled at the new window length. Use compression algorithms such as RMS (root mean square) compression or peak compression to adjust the amplitude of the signal. For example, if the adjustment factor is 0.8, multiply the initial analysis window length by 0.8 to obtain a new preliminary adjusted window length. Finally, output this preliminary adjusted window length.

[0073] Step S236: Perform binary quantization on the preliminary adjusted window length to obtain the adjusted window length parameter.

[0074] Specifically, the preliminary adjusted window length can be loaded. This window length is a floating-point number, expressed in milliseconds or seconds. Next, perform the binary quantization process. Quantization involves mapping continuous values to the closest discrete values. In this case, convert the floating-point value of the window length to binary format. This involves multiplying the floating-point number by a specific scale factor (for example, if the window length is in milliseconds, multiply it by 1000 to convert it to microseconds), then rounding the result to the closest integer, and finally converting that integer to its binary representation. For example, if the preliminary adjusted window length is 25 milliseconds, convert it to microseconds (25 milliseconds * 1000 = 25000 microseconds), then round it to the closest integer, which is 25000. Then, convert 25000 to its binary form. This binary value will be used as the adjusted window length parameter.

[0075] By performing spectral centroid recognition and time-varying spectral feature recognition on audio time-frequency feature data, the present invention can more accurately capture the essential features of audio signals. By analyzing the audio time-varying spectral feature matrix to identify the main frequency period sequence, the recognition of the main frequency period of the audio signal becomes more precise. According to the calculation of the time-scale ratio between the main frequency period sequence and the initial analysis window length, as well as the evaluation of the transient capture efficiency, the length of the analysis window can be dynamically adjusted. This makes the analysis window more conform to the actual characteristics of the audio signal, thereby improving the accuracy and effect of audio analysis. Through the evaluation of the transient capture efficiency and the dynamic range compression of the window length, the transient features in the audio signal can be captured more effectively. This is particularly important for improving the dynamic range and detail performance of audio signals, especially when processing music and sound effects containing rich transient information. By performing wavelet multi-resolution decomposition on the preliminarily adjusted window length and screening based on the principle of maximizing information entropy, the optimal window length can be selected from the multi-scale window length candidate set. This optimization selection mechanism makes the analysis window more suitable for the multi-scale characteristics of the audio signal, improving the flexibility and adaptability of audio analysis. By performing binary quantization on the optimized window length, the window length used in audio processing can be precisely controlled.

[0076] Preferably, step S25 includes the following steps:

[0077] Step S251: Calculate the spectral energy density of the initial spectrogram to obtain the spectral energy density distribution map, and perform spectral peak detection on the initial spectrogram to obtain the spectral peak feature set;

[0078] Specifically, the data of the initial spectrogram can be loaded, which contains the amplitude information of the audio signal at different frequencies and times. Square the amplitude of each frequency component at each time frame to obtain the energy value, and then sum the energy values of the same frequency component over all time frames to calculate the total energy of that frequency component. In this way, the spectral energy density distribution map is obtained, which shows the total energy distribution of the audio signal at different frequencies. Next, perform spectral peak detection on the spectral energy density distribution map. The purpose of spectral peak detection is to identify the significant peaks in the spectrum, which represent the most important frequency components in the audio signal. For this purpose, a threshold-based detection method is used, such as setting a global threshold or a local threshold, to distinguish noise from real spectral peaks. For example, set a threshold that is the average of all energy values in the spectral energy density distribution map plus a certain proportion of the standard deviation. All frequency components above this threshold will be considered spectral peaks and will be collected into the spectral peak feature set. Finally, output the spectral energy density distribution map and the spectral peak feature set.

[0079] Step S252: Determine the key frequency band energy threshold for the initial spectrogram based on the spectral energy density distribution map and the spectral peak feature set to obtain the key frequency band energy threshold;

[0080] Specifically, the spectral energy density distribution map and the spectral peak feature set can be loaded. Then, these data are analyzed to determine which frequency bands contain key energy information. This involves analyzing the spectral peak feature set to identify the most significant frequency range, or performing clustering analysis on the spectral energy density distribution map to identify the regions where energy is concentrated. Next, based on the analysis results, the energy threshold for the key frequency bands is determined. This threshold can be fixed or dynamic, depending on the characteristics of the audio signal and the required processing objectives. For example, a threshold is selected that is the median or average of the energy values within the key frequency bands, or a threshold that can distinguish between key and non-key frequency bands. Finally, the energy threshold for the key frequency bands is output.

[0081] Step S253: Mark and merge the high-energy frequency bands of the initial spectrogram according to the energy threshold of the key frequency bands to obtain preliminary key frequency band mapping data;

[0082] Specifically, the initial spectrogram and the energy threshold of the key frequency bands can be loaded. The energy value of each frequency component of the initial spectrogram is compared with the energy threshold of the key frequency bands. For the frequency bands with energy values higher than the threshold, they are marked as high-energy frequency bands. These high-energy frequency bands represent the most important parts of the audio signal, such as the melody line or rhythm strikes. Next, the marked high-energy frequency bands are merged. The purpose of the merge is to combine adjacent or overlapping high-energy frequency bands into a larger frequency band. The merge strategy is based on the frequency range, energy value, or time continuity of the frequency bands. For example, if two high-energy frequency bands are adjacent in frequency and their energy values are both significantly higher than the threshold, they are merged into a single preliminary key frequency band. Finally, the preliminary key frequency band mapping data are output, which include the information of all marked and merged high-energy frequency bands.

[0083] Step S254: Perform auditory weighting filtering on the initial spectrogram to obtain an auditory weighted spectrogram;

[0084] Specifically, the initial spectrogram can be loaded. An auditory weighting filter based on the human auditory system is used, which can simulate the sensitivity of the human ear to sounds of different frequencies. For example, the human ear is more sensitive to mid-frequency sounds than to low-frequency or high-frequency sounds, so the filter gives more weight to the mid-frequency components. The auditory weighting filter is applied to each frequency component of the initial spectrogram. The filtering process involves multiplying the energy value of each frequency component by the corresponding auditory weight, thereby adjusting the energy distribution of the spectrogram. In this way, some parts of the spectrogram are strengthened while other parts are weakened to better reflect the perceptual characteristics of human hearing. Finally, the auditory weighted spectrogram is output, which takes into account the non-linear characteristics of human hearing.

[0085] Step S255: Perform a perceptual importance scoring on the preliminary key frequency band mapping data based on the auditory weighted spectrogram to obtain a key frequency band importance matrix;

[0086] Specifically, the auditory weighted spectrogram and the preliminary key frequency band mapping data can be loaded. Analyze the auditory weighted spectrogram to determine the perceptual importance of each preliminary key frequency band. This involves considering the energy of the frequency band, the frequency position, and the relationship with other frequency bands. For example, give higher scores to those frequency bands with high energy and located within the frequency range sensitive to the human ear. Next, assign a perceptual importance score to each preliminary key frequency band, and this score reflects the contribution of the frequency band to the overall auditory perception. The score can be linear or non-linear, depending on the specific auditory model and algorithm. For example, use a psychoacoustic-based model, such as the Zwicker psychoacoustic model, to calculate the perceptual importance of each frequency band. Finally, output a key frequency band importance matrix that contains the perceptual importance scores of each preliminary key frequency band.

[0087] Step S256: Cluster the key frequency band importance matrix to obtain an optimized set of key frequency bands, and perform adaptive band-pass filtering on the optimized set of key frequency bands to obtain a key frequency band enhanced audio signal;

[0088] Specifically, the key frequency band importance matrix can be loaded. Use a clustering algorithm, such as K-means or hierarchical clustering, to group the key frequency bands. The purpose of clustering is to group the frequency bands with similar perceptual importance into one category. For example, group the frequency bands with scores higher than a certain threshold into the "high importance" category, and group the frequency bands with scores lower than the threshold into the "low importance" category. Next, perform adaptive band-pass filtering on each clustered optimized set of key frequency bands. The purpose of adaptive band-pass filtering is to enhance the signals of these frequency bands while reducing noise and other unimportant signals. Use an algorithm based on an adaptive filter, such as the LMS (Least Mean Square) algorithm, to dynamically adjust the parameters of the filter to adapt to the characteristics of the audio signal. Finally, output a key frequency band enhanced audio signal.

[0089] Step S257: Perform a time-frequency transform on the key frequency band enhanced audio signal to obtain key frequency band enhanced time-frequency spectrum data.

[0090] Specifically, the enhanced key frequency band audio signal can be loaded. Next, time-frequency transformation techniques, such as short-time Fourier transform (STFT), wavelet transform, or Wigner-Ville distribution, etc., are applied to convert the audio signal from the time domain to the time-frequency domain. This transformation can reveal the frequency components of the audio signal changing with time, providing richer time-frequency feature information. Taking STFT as an example, a suitable window function, such as the Hamming window, is selected, and the audio signal is windowed. Then, the windowed signal is segmented into short-time frames, and the fast Fourier transform (FFT) is performed on each short-time frame to calculate the spectral content of each frame. By changing the position of the window and repeating this process, a time-frequency spectrum can be constructed, which shows the amplitude and phase information of the frequency components changing with time. Finally, the time-frequency spectrum data of the enhanced key frequency band is output.

[0091] Through calculating the spectral energy density and detecting the spectral peaks of the initial spectrogram, the present invention can accurately identify the key frequency bands and significant features in the audio signal. By combining the spectral energy density distribution and the spectral peak feature set to determine the key frequency band energy threshold, the key frequency bands in the audio signal can be identified more precisely. This method based on the energy threshold improves the reliability of key frequency band identification. By marking and merging the high-energy frequency bands of the initial spectrogram, the high-energy parts in the audio signal can be identified and integrated. This helps to highlight the most important parts of the audio signal, making the audio processing more focused on the core content of the signal. Through auditory weighted filtering and perceptual importance scoring, the audio signal can be optimized according to the auditory characteristics of the human ear. This method makes the processing of the audio signal more in line with the auditory habits of humans, thus improving the auditory perception quality of the audio signal. Through clustering and adaptive band-pass filtering of the key frequency band importance matrix, the key frequency bands of the audio signal can be intelligently optimized. This not only improves the signal quality of the key frequency bands, but also makes the audio signal clearer and more natural. By performing time-frequency transformation on the key frequency band enhanced audio signal, the time-frequency features of the audio signal can be further enhanced. This helps to enhance the dynamic range and detail performance of the audio signal while maintaining the audio signal quality.

[0092] Preferably, step S27 includes the following steps:

[0093] Step S271: Perform logarithmic spectrum enhancement on the optimized spectrogram to obtain an enhanced logarithmic spectrogram;

[0094] Specifically, data of the optimized spectrogram can be loaded, which contains the amplitude information of the audio signal at different frequencies. Take the logarithm of the amplitude value of each frequency component to obtain a spectral representation that is closer to the perceived loudness of the human ear. The logarithmic transformation helps to compress the dynamic range of the audio signal, enhancing the low-energy frequency bands (which are usually less perceptible to the human ear) while avoiding distortion in the high-energy frequency bands. In specific operations, the natural logarithm or the logarithm with base 10 is used, and a certain offset processing is performed on the logarithmic spectrum to avoid the problem of negative infinity in logarithmic operations. For example, if the amplitude value range in the original spectrogram is from 0.01 to 1000, after taking the logarithm, these values will be converted to a more manageable range, such as from -4.6 to 3. Next, further enhancement processing is performed on the logarithmically transformed spectrum, including dynamic range compression or gain adjustment, to highlight the key features in the spectrum. Finally, an enhanced logarithmic spectrogram is output, which shows the spectrum of the audio signal after logarithmic transformation and enhancement processing.

[0095] Step S272: Perform cepstral fundamental frequency estimation on the enhanced logarithmic spectrogram to obtain an audio fundamental frequency estimation sequence, and calculate the theoretical harmonic frequencies of the optimized spectrogram based on the audio fundamental frequency estimation sequence to obtain a theoretical harmonic frequency map;

[0096] Specifically, the enhanced logarithmic spectrogram can be loaded. Perform cepstral analysis, which is a method of converting the logarithmic spectrum into a cepstrum. The cepstrum is a representation method that shows the periodic components of the spectrum and can reveal the fundamental frequency information in the audio signal. In specific operations, use the discrete cosine transform (DCT) to transform the logarithmic spectrum to obtain cepstral coefficients. The peak position of the cepstral coefficients corresponds to the fundamental frequency period of the audio signal. Identify these peaks and extract the corresponding fundamental frequency estimation sequence. Next, calculate the theoretical harmonic frequencies of the optimized spectrogram based on the audio fundamental frequency estimation sequence. This involves calculating the harmonic series for each fundamental frequency estimation value, that is, integer multiples of the fundamental frequency. Use a predefined harmonic range, such as up to the 20th harmonic, to calculate the harmonic frequencies of each fundamental frequency. Finally, output a theoretical harmonic frequency map, which shows the theoretical harmonic frequency distribution calculated based on the fundamental frequency estimation sequence.

[0097] Step S273: Detect harmonic peaks in the optimized spectrogram according to the theoretical harmonic frequency map to obtain an actual harmonic feature set;

[0098] Specifically, a theoretical harmonic frequency map and an optimized spectrogram can be loaded. The theoretical harmonic frequency map provides the expected harmonic frequency positions in the audio signal, while the optimized spectrogram contains the actual spectral data of the audio signal. Peak detection is performed at each theoretical harmonic frequency position in the optimized spectrogram. This involves searching for local maxima in the spectrogram near each theoretical harmonic frequency. To improve the accuracy of detection, a peak detection algorithm such as a threshold-based method or a local maximum-based method is used. For example, a threshold is set, and only when the amplitude value in the spectrum is higher than this threshold is a harmonic peak considered to exist at that position. Once the peaks are detected, the positions and amplitudes of these peaks are recorded and combined into an actual harmonic feature set.

[0099] Step S274: Calculate the harmonic energy ratio for the actual harmonic feature set to obtain a harmonic energy ratio matrix, and extract the time-varying harmonic structure from the actual harmonic feature set to obtain a time-varying harmonic structure sequence;

[0100] Specifically, the actual harmonic feature set can be loaded. Calculate the energy ratio for each detected harmonic peak. This involves dividing the amplitude value of each harmonic by the sum of the amplitudes of all harmonics to obtain the energy ratio of each harmonic. This ratio reflects the energy contribution of each harmonic in the overall audio signal. Next, a harmonic energy ratio matrix is constructed, which contains the energy ratio information of each harmonic. In addition, extract the time-varying harmonic structure from the actual harmonic feature set. This involves analyzing the variation of harmonic peaks over time to identify the time-varying harmonic structure in the audio signal. A time series analysis method such as a Hidden Markov Model (HMM) or a Recurrent Neural Network (RNN) is used to model and extract the time-varying harmonic structure. Finally, the harmonic energy ratio matrix and the time-varying harmonic structure sequence are output.

[0101] Step S275: Generate harmonic structure data from the harmonic energy ratio matrix and the time-varying harmonic structure sequence according to the enhanced logarithmic spectrogram to obtain fundamental frequency harmonic structure data; where the fundamental frequency harmonic structure data includes fundamental frequency trajectory data, harmonic frequency sequence, and harmonic intensity matrix;

[0102] Specifically, an enhanced logarithmic spectrogram, a harmonic energy ratio matrix, and a time-varying harmonic structure sequence can be loaded. Analyze these data to generate fundamental frequency harmonic structure data, including fundamental frequency trajectory data, harmonic frequency sequence, and harmonic intensity matrix. First, identify the fundamental frequency trajectory in the enhanced logarithmic spectrogram, that is, the variation of the most significant frequency component over time. Then, extract the harmonic frequency sequence, which is obtained based on the theoretical harmonic frequency diagram and the actual harmonic peak detection results. The harmonic intensity matrix is calculated according to the harmonic energy ratio matrix, representing the relative intensity of each harmonic. For example, assume there is an audio signal with a fundamental frequency of 100 Hz and 3 significant harmonics with frequencies of 200 Hz, 300 Hz, and 400 Hz respectively. Analyze the enhanced logarithmic spectrogram to determine the variation of these frequency components over time and record their amplitudes. In this way, the fundamental frequency trajectory data (e.g., the variation of the amplitude of 100 Hz over time), the harmonic frequency sequence ([200 Hz, 300 Hz, 400 Hz]), and the harmonic intensity matrix (the amplitude values of each harmonic) are obtained. Finally, the fundamental frequency harmonic structure data is output.

[0103] Step S276: Identify the phase relationship between the fundamental frequency trajectory data and the harmonic frequency sequence to obtain a harmonic phase relationship diagram, and perform harmonic spatial redistribution on the fundamental frequency harmonic structure data according to the harmonic phase relationship diagram and the harmonic intensity matrix to obtain redirected harmonic structure data;

[0104] Specifically, the fundamental frequency trajectory data and the harmonic frequency sequence can be loaded. Analyze these data to identify the phase relationship between the harmonics. This involves calculating the phase difference between each harmonic and its fundamental frequency and tracking the variation of these phase differences over time. For example, use the short-time Fourier transform (STFT) to calculate the phases of the fundamental frequency and each harmonic in each time frame, and then calculate the phase difference between them. By analyzing the variation of these phase differences over time, a harmonic phase relationship diagram is constructed, which shows the phase relationship between the fundamental frequency and each harmonic. Next, use the harmonic phase relationship diagram and the harmonic intensity matrix to perform harmonic spatial redistribution on the fundamental frequency harmonic structure data. This involves adjusting the phase and amplitude of the harmonics to improve the sound quality of the audio signal or to adapt to specific audio processing objectives. For example, increase the amplitude of a specific harmonic or change its phase to enhance a specific feature of the audio signal. Finally, the redirected harmonic structure data is output, which includes the adjusted fundamental frequency trajectory, harmonic frequency sequence, and harmonic intensity matrix.

[0105] Step S277: Perform harmonic synthesis on the redirected harmonic structure data to obtain a preliminary reconstructed audio signal, and perform envelope modulation on the preliminary reconstructed audio signal to obtain an envelope-modulated audio signal;

[0106] Specifically, redirected harmonic structure data can be loaded, which includes fundamental frequency trajectories, harmonic frequency sequences, and adjusted harmonic intensity matrices. These data are used to synthesize an audio signal. The synthesis process involves setting the frequency, amplitude, and phase information of each harmonic component sine wave according to the redirected harmonic structure data. In a specific operation, the following formula is applied to each harmonic component to generate a signal: where s(t) is the audio signal at time t, M is the total number of harmonics, A m is the amplitude of the m-th harmonic, y m is the frequency of the m-th harmonic, ψ m is the phase of the m-th harmonic, t is the time variable. After synthesizing the preliminary reconstructed audio signal, envelope modulation will be performed. Envelope modulation involves modulating the amplitude of each harmonic component according to an envelope function E(t), which can be expressed as:

[0107]

[0108] where A1 is the peak amplitude of the envelope; t is the time variable; e is the base of the natural logarithm; T attack is the attack time, that is, the time when the signal rises from zero; t decay is the decay time point, that is, the time from the start of the attack to the end of the decay; T decay is the decay time, that is, the time when the signal decays from the peak to the sustain level; t sustain is the sustain time point, that is, the time from the end of the decay to the start of the release; T release is the release time, that is, the time when the signal decays from the sustain level to zero. This envelope function can be the amplitude envelope of the audio signal, which describes the time-varying amplitude of the signal. For example, using an ADSR (Attack, Decay, Sustain, Release) envelope, which is a common envelope shape in music and audio processing, is used to simulate the natural start and end of a sound. Finally, the envelope-modulated audio signal is output, and this signal reflects the natural characteristics of the audio in the time domain, such as the dynamic changes in the start, sustain, and end phases of the sound.

[0109] Step S278: Perform phase correction on the envelope-modulated audio signal to obtain the reconstructed audio data.

[0110] Specifically, an envelope-modulated audio signal can be loaded. Analyze the phase content of the signal and compare it with an ideal phase reference to identify and correct any phase deviations or distortions. This involves using phase correction algorithms such as the Hilbert transform to estimate the instantaneous phase and amplitude of the signal. In specific operations, apply the Hilbert transform to the signal to obtain the analytic signal, and then calculate the instantaneous phase of the analytic signal. Next, compare this instantaneous phase with the theoretical phase (e.g., the phase of a sine wave) and calculate the difference between the two. Based on this difference, adjust the phase of the audio signal to align it with the theoretical phase. Finally, output the phase-corrected audio data, which reflects the ideal characteristics of the audio signal in both the frequency domain and the time domain.

[0111] The present invention can more deeply analyze the spectral characteristics of an audio signal by performing logarithmic spectral enhancement on an optimized spectrogram. This makes the details in the spectrogram more prominent and helps capture the subtle changes in the audio signal. Precise fundamental frequency and harmonic frequency information for the audio signal are provided through cepstral fundamental frequency estimation and theoretical harmonic frequency calculation. By performing harmonic peak detection on the optimized spectrogram and collecting actual harmonics, the harmonic components in the audio signal can be identified and enhanced. This helps improve the richness and expressiveness of the audio signal, making the reconstructed audio signal closer to the original sound quality. By calculating the harmonic energy ratio and extracting the time-varying harmonic structure, the energy distribution of the audio signal can be better understood and adjusted. This optimization helps improve the dynamic range and balance of the audio signal, making the audio output more natural and pleasant. By extracting and analyzing the sequence of time-varying harmonic structures, the structural characteristics of the audio signal changing over time can be captured. This time-varying analysis helps maintain the coherence and fluency of the audio signal and enhances the listening experience. By identifying the phase relationship and spatial reallocation of the fundamental frequency harmonic structure data, and subsequent harmonic synthesis and envelope modulation, the audio signal can be accurately reconstructed. This not only retains the key information of the audio signal but also improves the overall quality of the signal. The final phase correction step ensures the phase accuracy of the reconstructed audio signal, which is crucial for ensuring the clarity and stability of the sound quality. The phase-corrected audio signal can provide a more pure and accurate auditory experience.

[0112] Preferably, step S3 includes the following steps:

[0113] Step S31: Collect historical played tracks of the client to obtain a historical played track dataset, and identify the audio playback pattern of the user based on the historical played track dataset to obtain user audio playback feature data; wherein, the user audio playback feature data includes user style preference distribution data, user playback time pattern, and user music element preference data;

[0114] Specifically, a historical play track dataset can be collected from the music playback application on the client side. This involves accessing the database or reading the user's music playback log to obtain the list of music played by the user over a period of time. The detailed information of each song is recorded, such as the artist, album, number of plays, play duration, and play timestamp. Next, machine learning algorithms are used to analyze this data to identify the user's audio playback patterns. For example, clustering algorithms (such as K-means) are used to group the music played by the user to identify the music styles preferred by the user. In addition, the user's play time patterns are also analyzed, such as the time period of the day when the user usually listens to music and the duration of each listening session. Music Information Retrieval (MIR) techniques are used to analyze the characteristics of each song, such as rhythm, tonality, pitch, timbre, and intensity. A feature vector is constructed to represent these elements of each song, and these feature vectors are used to train a classifier, such as a Support Vector Machine (SVM), to identify the user's preferences for different music elements. Finally, the user audio playback feature data is output, including the user style preference distribution data, the user play time pattern, and the user music element preference data.

[0115] Step S32: Construct a user profile for the user based on the user style preference distribution data, the user play time pattern, and the user music element preference data to obtain a user audio profile, and use a preset preference recognition model to identify the user preference type of the user audio profile to obtain a user music preference type identifier; wherein, the user music preference type identifier is any one of a rhythm-oriented listener, a melody-oriented listener, and an atmosphere-oriented listener;

[0116] Specifically, the user style preference distribution data, the user play time pattern, and the user music element preference data can be loaded. Next, data mining techniques are used to construct a user audio profile. This involves using machine learning algorithms such as decision trees, random forests, or gradient boosting machines to analyze the user's data and find the patterns preferred by the user. For example, it is found that the user often listens to jazz at night and prefers music with a strong rhythm on certain days of the week. Then, a preset preference recognition model is used to classify the user audio profile. This model classifies users into different music preference types, such as rhythm-oriented listeners, melody-oriented listeners, and atmosphere-oriented listeners, based on previous user studies. A classifier, such as a Support Vector Machine (SVM) or a neural network, is trained to identify the user's music preference type based on the feature data of the user audio profile. For example, if the user's music element preference data shows that they often listen to music with complex melodies and the play time pattern shows that they usually listen to music when relaxing, then the model will classify this user as a melody-oriented listener. Finally, the user music preference type identifier is output.

[0117] Step S33: If the user's music preference type identifier indicates a rhythm-oriented listener, perform low-frequency phase consistency enhancement on the reconstructed audio data to obtain rhythm-enhanced audio data;

[0118] Specifically, perform low-frequency phase consistency enhancement on the reconstructed audio data of users identified as rhythm-oriented listeners to obtain rhythm-enhanced audio data. The reconstructed audio data can be loaded. Focus on the low-frequency part of the audio data, as these parts usually contain rhythm information. Use a band-pass filter to isolate the low-frequency band. For example, the cut-off frequency is set below 200 Hz. Next, perform phase consistency enhancement on these low-frequency components. This involves analyzing and adjusting the phase of adjacent samples to ensure the phase continuity and consistency of the rhythm components. One method is to use phase shaping techniques, such as the Hilbert transform, to generate an analytic signal and adjust the phase to enhance the rhythm perception. For example, if the phase of the rhythm components in the audio data is discontinuous in some parts, adjust the phase of these parts to smooth the transition, thereby enhancing the perception of rhythm. Finally, output the rhythm-enhanced audio data, which have enhanced phase consistency and amplitude in the low-frequency band.

[0119] Step S34: If the user's music preference type identifier indicates a melody-oriented listener, perform mid-high frequency phase relationship enhancement on the reconstructed audio data to obtain melody-prominent audio data;

[0120] Specifically, perform mid-high frequency phase relationship enhancement on the reconstructed audio data of users identified as melody-oriented listeners to obtain melody-prominent audio data. The reconstructed audio data can be loaded. Focus on the mid-high frequency part of the audio data, as these parts usually contain melody information. Use a band-pass filter to isolate the mid-high frequency band. For example, the cut-off frequency is set between 200 Hz and 4000 Hz. Next, perform phase relationship enhancement on these mid-high frequency components. This involves analyzing and adjusting the phase of adjacent frequency components to ensure the clarity and coherence of the melody line. One method is to use spectral editing techniques, such as phase-based spectral shaping, to enhance the melody components in a specific frequency range. For example, if the phase of the melody components in the audio data is inconsistent at some frequencies, adjust the phase of these frequencies to ensure the coherence of the melody line. Finally, output the melody-prominent audio data, which have enhanced phase relationship and amplitude in the mid-high frequency band.

[0121] Step S35: If the user's music preference type identifier indicates an atmosphere-oriented listener, perform full-frequency phase spatial sense enhancement on the reconstructed audio data to obtain spatially enhanced audio data;

[0122] Specifically, perform full-band phase spatial sense enhancement on the reconstructed audio data for users identified as ambience-oriented listeners to obtain spatially enhanced audio data. The reconstructed audio data can be loaded. Focus on the entire frequency band to enhance the spatial sense of the audio. Use stereo expansion techniques such as HRTF (Head-Related Transfer Function) or virtual speaker technology to simulate a wider sound field. Next, perform phase spatial sense enhancement on the audio data. This involves using cross-feedback delay networks to increase the perception of sound depth and width, or using environment-based filters to simulate the reflection and reverberation characteristics of different acoustic environments. For example, analyze each frequency component in the audio and adjust its phase and amplitude according to its frequency characteristics to simulate sound propagation in a specific environment. This includes adding early reflections and late reverberation to create a richer and more immersive listening experience. Finally, output the spatially enhanced audio data, which has enhanced phase spatial sense across the full frequency band.

[0123] Step S36: Perform phase modulation processing on the reconstructed audio data according to the rhythm-enhanced audio data or melody-prominent audio data or spatially enhanced audio data to obtain phase-optimized audio data.

[0124] Specifically, the reconstructed audio data and the corresponding enhanced audio data (rhythm-enhanced, melody-prominent, or spatially enhanced) can be loaded. Analyze this data and determine the phase modulation strategy to be applied. Next, perform phase modulation processing on the reconstructed audio data. This involves using phase modulation techniques such as phase shift or phase rotation to adjust the phase characteristics of the audio signal. For example, use a phase modulator that dynamically adjusts the phase of the original audio signal according to the characteristics of the enhanced audio data. If the enhanced audio data is rhythm-enhanced, focus on enhancing the phase consistency of the rhythm components. If the enhanced audio data is melody-prominent, enhance the phase clarity of the melody line. If the enhanced audio data is spatially enhanced, adjust the phase to expand the sound field. Finally, output the phase-optimized audio data.

[0125] By collecting the user's historical played tracks and identifying their audio playback patterns, the present invention can collect user audio playback feature data reflecting user preferences. This helps to more accurately understand the user's music listening habits and preferences, thereby providing audio content more in line with the user's taste. By constructing a user audio portrait using data such as user style preferences, playback time patterns, and music element preferences, the user classification becomes more detailed and accurate. According to the user's music preference type identifier (rhythm-oriented, melody-oriented, or atmosphere-oriented listener), targeted phase enhancement and modulation processing can be performed on the audio data. This makes the audio output better meet the auditory needs of different users. For rhythm-oriented listeners, enhancing the low-frequency phase consistency can make the rhythm more prominent; for melody-oriented listeners, enhancing the mid-high frequency phase relationship can make the melody clearer; for atmosphere-oriented listeners, enhancing the full-band phase spatial sense can provide a richer auditory background. By performing phase modulation processing on the reconstructed audio data, the perceived quality of the audio can be further improved. Whether it is the sense of rhythm, melody, or spatial sense, corresponding optimizations can be achieved, making the audio output more in line with the user's personal preferences.

[0126] Preferably, step S4 includes the following steps:

[0127] Step S41: Deploy a multi-point microphone array in the environment where the Bluetooth speaker is located, and perform spatial sound field sampling and positioning to obtain a spatial sampling point distribution map. Use the microphone array to perform broadband swept-frequency excitation on the environment where the Bluetooth speaker is located, and collect response signals to obtain multi-channel environmental response signals. Perform time-frequency identification on the multi-channel environmental response signals to obtain a time-varying spectrum atlas;

[0128] Specifically, a multi-point microphone array can be deployed inside the environment where the Bluetooth speaker is located. This array consists of four or more microphones, which are placed at different positions in the room to ensure comprehensive spatial sound field sampling. The microphones are connected to a central data acquisition system either wired or wirelessly. Next, the spatial sound field sampling and positioning process is initiated. This involves playing a known sound source, such as white noise or a maximum length sequence, and recording it through the microphone array. By analyzing these recordings, the positions of each microphone and the relative distances between them can be determined, thereby generating a spatial sampling point distribution map. Then, the microphone array is used to perform broadband swept-frequency excitation on the environment. This involves playing a series of pure tones at different frequencies, covering the audible range of the human ear (usually 20 Hz to 20 kHz). Each pure tone will generate reflections and reverberations in the environment, and the microphone array will capture these response signals. The acquired response signals are multi-channel, and each microphone channel will record a unique version of the signal. Time-frequency identification is performed on these multi-channel environmental response signals, usually using the short-time Fourier transform (STFT) or other time-frequency analysis methods, to obtain time-varying spectrogram atlases. These atlases reveal the response characteristics of the environment to sounds at different frequencies, including reflections, absorption, and reverberation, etc. Finally, the time-varying spectrogram atlases are output, which contain the response information of each sampling point in the environment to sounds at different frequencies.

[0129] Step S42: Extract spatial acoustic parameters from the time-varying spectrogram atlas for the environmental sound field to obtain a spatial acoustic feature matrix;

[0130] Specifically, digital signal processing techniques can be used to analyze each set of time-varying spectrograms (the time-varying spectrogram atlas contains multiple sets of time-varying spectrograms). One method is to estimate the reverberation time by measuring the attenuation slope of the signal at different time points. This can be achieved by performing a linear fit on the energy attenuation of each frequency component and then calculating the slope. The steeper the slope, the shorter the reverberation time; the flatter the slope, the longer the reverberation time. In addition, the resonance modes of the room can be determined through spectral envelope analysis. This involves analyzing the spectral energy of each time frame to identify regions where the energy is concentrated at specific frequencies, and these regions correspond to the resonance frequencies of the room. After extracting parameters such as the reverberation time and resonance modes, these parameters are organized into a spatial acoustic feature matrix. In this matrix, each row represents a sampling point, and each column represents a specific acoustic parameter, such as the reverberation time, sound pressure level distribution, and frequency response. In this way, a comprehensive dataset describing the environmental acoustic characteristics is obtained, which contains the key acoustic parameters at various positions in the environment.

[0131] Step S43: Identify the reverberation characteristics of the spatial acoustic feature matrix to obtain a reverberation feature vector;

[0132] Specifically, relevant acoustic parameters can be extracted from the spatial acoustic feature matrix, such as the reverberation time and frequency response at each sampling point. These parameters provide the basic data on the behavior of sound waves in the environment. Next, autocorrelation analysis is used to determine the time distribution of reverberation. The autocorrelation function can reveal the correlation of a signal with itself at different time delays, thereby identifying the duration of reverberation. By calculating the autocorrelation values at different time delays, the decay curve of reverberation can be estimated, and then the reverberation time can be obtained. In addition, to identify early reflections, the signal delays of different microphone channels can be compared. Early reflections refer to the sound waves that are reflected back after the sound waves first hit the room surface, and they have an important impact on the sense of sound localization and spatial perception. By measuring the time difference of the arrival of early reflections received by different microphones, the distances between the sound source and each reflecting surface can be estimated, and then the geometric structure of the room can be inferred. After the above analysis is completed, the obtained data is integrated into a reverberation feature vector. This vector contains key information describing the reverberation characteristics of the environment, such as the reverberation time, reflection intensity, and the performance of reverberation at different frequencies. The reverberation time reflects the duration of sound waves in the environment, the reflection intensity describes the intensity of early reflections, and the frequency content of reverberation reveals the distribution of reverberation at different frequencies.

[0133] Step S44: Conduct acoustic ray tracing simulation on the spatial sampling point distribution map to obtain a sound energy distribution model, and based on the sound energy distribution model, extract acoustic features and construct a database for the reverberation feature vector to obtain an environmental acoustic feature database;

[0134] Specifically, professional acoustic simulation software can be used for acoustic ray tracing simulation, such as Odeon or Raynoise. These software can simulate the propagation behavior of sound waves in three-dimensional space. First, import the spatial sampling point distribution map into the acoustic simulation software. This map contains the geometric structure of the room and the position information of the sampling points. Then, the software will simulate the propagation path of sound waves from the sound source to each sampling point according to the size and material properties of the room, including phenomena such as reflection, refraction, and diffraction of sound waves. During the simulation, the software will track each interaction of sound waves with the room interface and calculate the attenuation of sound waves on different paths. This involves considering the reflection coefficient of sound waves on different material surfaces and the energy loss during the propagation of sound waves. By this method, the distribution of sound energy at each point in the room can be predicted, and a detailed sound energy distribution model can be generated. Next, use the sound energy distribution model to extract the acoustic parameters in the reverberation feature vector. For example, the energy attenuation of sound waves at different frequencies in the model can be analyzed to estimate the reverberation time of the room. In addition, the diffusion degree of sound waves in the room can be evaluated to calculate parameters such as the clarity and intensity of the room. Finally, organize these extracted acoustic parameters into an environmental acoustic feature database. This database details the reverberation time, clarity, intensity, and other related acoustic characteristics of the room.

[0135] Step S45: Perform spectral compensation on the phase-optimized audio data based on the environmental acoustic feature database to obtain an environment-compensated spectrum;

[0136] Specifically, the environmental acoustic feature database can be accessed. This database contains detailed room acoustic parameters, such as reverberation time, absorption coefficient, diffusion characteristics, etc. Next, perform spectral analysis on the phase-optimized audio data, which is achieved through fast Fourier transform (FFT), converting the audio signal from the time domain to the frequency domain to obtain its spectral representation. The spectrum represents the energy distribution of the audio signal at different frequencies. According to the parameters in the environmental acoustic feature database, map how the acoustic characteristics of the room affect the audio signal. For example, a room with a longer reverberation time will cause the high-frequency components of the audio signal to attenuate, while materials with a higher absorption coefficient will reduce the energy at specific frequencies. Based on the above mapping, adjust the spectrum of the audio signal. If the reverberation time of the room is too long, enhance the high-frequency components of the audio signal to improve the clarity of speech. This can be achieved through a high-pass filter or gain boost. On the contrary, if the reverberation time is short, enhance the low-frequency components to increase the warmth of the sound, which can be achieved through a low-pass filter or gain boost. In this way, an environment-compensated spectrum can be generated, which takes into account the influence of the environment on the audio signal, thus providing a consistent auditory experience in different environments.

[0137] Step S46: Perform blind deconvolution on the environment compensation spectrum to obtain a reverberation-removed environment compensation spectrum, and perform spatial sound image reshaping on the reverberation-removed environment compensation spectrum to obtain an audio space optimized spectrum;

[0138] Specifically, the environment compensation spectrum can be analyzed in detail first to identify the characteristics of reverberation. This includes analyzing the delay response and attenuation characteristics in the spectrum, which are typical signs of reverberation. Use the Fourier transform to convert the spectrum into the time domain to identify the duration and intensity of reverberation. Based on the parameters in the environmental acoustic feature database, such as room dimensions, material properties, and known reverberation time, construct a mathematical model to describe the reverberation process. This model will predict the propagation and reflection behavior of sound waves in the room. Utilize the mathematical model of reverberation and apply a blind deconvolution algorithm to attempt to reverse the reverberation process. Blind deconvolution is a signal processing technique that attempts to recover the original signal from the observed signal even without knowing the specific details of the reverberation. Use algorithms such as Wiener filtering or the least squares method to estimate and eliminate the influence of reverberation, thereby recovering the original spectrum unaffected by reverberation. Next, perform spatial sound image reshaping on the reverberation-removed environment compensation spectrum. This involves using sound field analysis techniques, such as wave field synthesis or stereo processing techniques, to adjust the spatial characteristics of the audio signal. Analyze the geometric structure and material properties of the room, and then adjust the left and right channel balance, stereo width, and depth perception of the audio signal accordingly. This can be achieved by adjusting the delay, phase, and amplitude of the audio signal to simulate the effects of different sound source positions and environmental reflections. Finally, process the adjusted audio signal to output the audio space optimized spectrum.

[0139] Step S47: Perform an inverse phase correction transformation on the audio space optimized spectrum to obtain an audio time domain compensation signal, and perform loudness equalization on the audio time domain compensation signal to obtain the final optimized audio data.

[0140] Specifically, an audio spatial optimized spectrum can be loaded and an inverse transform, such as an inverse Fourier transform, can be applied to convert the frequency-domain signal back to the time-domain signal. During this process, special attention needs to be paid to correcting the phase of the signal to ensure that the converted time-domain signal has the correct timing and phase relationship. This involves analyzing the phase content of the spectrum and adjusting the phase as needed to compensate for any phase distortion introduced by previous processing steps. Next, loudness equalization processing is performed on the audio time-domain compensation signal. Loudness equalization techniques, such as dynamic range compression or multi-band compression, are used to adjust the loudness of the audio signal so that it can maintain a consistent perceived loudness at different volume levels. The loudness of the audio signal is analyzed and the amplitude of the signal is adjusted according to a loudness model, such as the Zwicker loudness model, to improve the audibility of the low-loudness part while preventing distortion in the high-loudness part. Finally, the final optimized audio data will be output, which has been spatially optimized and loudness equalized while maintaining the original audio characteristics.

[0141] Through the deployment of a multi-point microphone array and spatial sound field sampling and positioning, the present invention can accurately capture the acoustic characteristics of the environment and generate a detailed spatial sampling point distribution map. By performing time-frequency identification on the multi-channel environmental response signals, comprehensive parameters of the environmental sound field can be extracted to form a spatial acoustic feature matrix. This helps to deeply understand the impact of the environment on audio signals. By obtaining the reverberation eigenvector, the reverberation characteristics in the environment can be identified, which is crucial for adjusting audio signals to adapt to different environments. This identification helps to reduce the negative impact of reverberation on the sound quality. Based on the sound energy distribution model and the environmental acoustic feature database, spectral compensation can be performed on the audio data to adapt to different acoustic environments. This adaptability enables the audio output to maintain a high sound quality in different environments. Through blind deconvolution and spatial sound image reshaping, the influence of reverberation can be removed and the spatial performance of the audio can be optimized. This makes the positioning of the audio signal in space more accurate, enhancing the listener's immersion and sense of space. Through phase correction inverse transform and loudness equalization processing, it helps to improve the clarity and stability of the audio signal. After these processes, the audio signal can adapt to different volume levels while maintaining its original quality, providing a consistent auditory experience.

[0142] Preferably, the present invention also provides a system for processing Bluetooth speaker data, which is used to execute the method for processing Bluetooth speaker data as described above. The system for processing Bluetooth speaker data includes:

[0143] A data compression module, which is used to perform multi-modal data capture on the client to obtain original multi-modal client data; perform adaptive multi-resolution compression on the original multi-modal client data to obtain compressed multi-modal client data;

[0144] An audio reconstruction module, configured to transmit compressed multi-modal client data to a Bluetooth speaker through a client to obtain received compressed multi-modal data; decouple the received compressed multi-modal data to obtain separated audio data; perform dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data;

[0145] An audio phase optimization module, configured to determine the music preference type of a user to obtain a user music preference type identifier; perform a phase response on the reconstructed audio data according to the user music preference type identifier to obtain phase-optimized audio data;

[0146] An environmental acoustics compensation module, configured to construct an environmental acoustics feature database for the environment where the Bluetooth speaker is located to obtain an environmental acoustics feature database; perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustics feature database to obtain finally optimized audio data.

[0147] Preferably, a computer-readable storage medium stores a method for processing Bluetooth speaker data as described above that can be loaded and executed by a processor.

[0148] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be encompassed within the present invention.

[0149] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing Bluetooth speaker data, characterized in that, It includes the following steps: Step S1: Capture multi-modal data of the client to obtain original multi-modal client data; Perform adaptive multi-resolution compression on the original multi-modal client data to obtain compressed multi-modal client data; Step S2: Transmit the compressed multi-modal client data to the Bluetooth speaker through the client to obtain received compressed multi-modal data; perform multi-modal data decoupling on the received compressed multi-modal data to obtain separated audio data; perform dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data; Step S3: Determine the music preference type of the user to obtain a user music preference type identifier; Perform phase response on the reconstructed audio data according to the user music preference type identifier to obtain phase-optimized audio data; among them, step S3 includes the following steps: Step S31: Collect historical played tracks of the client to obtain a historical played track data set, and perform audio playback mode recognition on the user based on the historical played track data set to obtain user audio playback feature data; among them, the user audio playback feature data includes user style preference distribution data, user playback time pattern, and user music element preference data; Step S32: Construct a user audio portrait for the user according to the user style preference distribution data, user playback time pattern, and user music element preference data to obtain a user audio portrait, and use a preset preference recognition model to perform user preference type recognition on the user audio portrait to obtain a user music preference type identifier; among them, the user music preference type identifier is any one of a rhythm-oriented listener, a melody-oriented listener, and an atmosphere-oriented listener; Step S33: If the user music preference type identifier is a rhythm-oriented listener, perform low-frequency phase consistency enhancement on the reconstructed audio data to obtain rhythm-enhanced audio data; Step S34: If the user music preference type identifier is a melody-oriented listener, perform mid-high frequency phase relationship enhancement on the reconstructed audio data to obtain melody-prominent audio data; Step S35: If the user music preference type identifier is an atmosphere-oriented listener, perform full-band phase space sense enhancement on the reconstructed audio data to obtain space-sense enhanced audio data; Step S36: Perform phase modulation processing on the reconstructed audio data according to the rhythm-enhanced audio data or melody-prominent audio data or space-sense enhanced audio data to obtain phase-optimized audio data; Step S4: Construct an environmental acoustic feature database for the environment where the Bluetooth speaker is located to obtain an environmental acoustic feature database; perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustic feature database to obtain finally optimized audio data.

2. The method for processing Bluetooth speaker data according to claim 1, wherein, Step S1 includes the following steps: Step S11: Capture multi-modal data of the client to obtain original multi-modal client data; Step S12: Separate modal features of the original multi-modal client data to obtain client multi-modal feature data; among them, the client multi-modal feature data includes audio data, control instruction data, and device status data; Step S13: Perform run-length encoding and differential pulse coding on the control instruction data to obtain compressed control instruction data; Step S14: Perform incremental encoding based on change detection on the device status data to obtain compressed device status data; Step S15: Perform multi-scale signal decomposition on the audio data to obtain client multi-scale audio data, and perform wavelet transform on the client multi-scale audio data to obtain feature-enhanced audio data; Step S16: Perform enhanced Huffman encoding on the feature-enhanced audio data to obtain compressed audio data; Step S17: Perform embedded multi-modal fusion on the compressed control instruction data, compressed device status data, and compressed audio data to obtain compressed multi-modal client data.

3. The method for processing Bluetooth speaker data according to claim 1, wherein Step S2 includes the following steps: Step S21: Transmit the compressed multi-modal client data to the Bluetooth speaker through the client and receive the data to obtain received compressed multi-modal data; Step S22: Decouple the received compressed multi-modal data to obtain separated audio data, and extract time-frequency features from the separated audio data to obtain audio time-frequency feature data; Step S23: Adjust the preset initial analysis window length according to the audio time-frequency feature data to obtain an adjusted window length parameter; Step S24: Perform short-time Fourier transform on the separated audio data according to the adjusted window length parameter to obtain an initial spectrogram; Step S25: Identify the key frequency bands of the initial spectrogram to obtain audio key frequency band data, and perform high spectral density analysis on the audio key frequency bands in the separated audio data according to the audio key frequency band data to obtain key frequency band enhanced time-frequency spectrum data; Step S26: Update the initial spectrogram according to the key frequency band enhanced time-frequency spectrum data to obtain an optimized spectrogram; Step S27: Perform harmonic fundamental frequency analysis on the optimized spectrogram to obtain fundamental frequency harmonic structure data, and reconstruct the fundamental frequency harmonic structure data to obtain reconstructed audio data.

4. The method for processing Bluetooth speaker data according to claim 3, wherein Step S23 includes the following steps: Step S231: Identify the spectral centroid of the audio time-frequency feature data to obtain a spectral centroid distribution map, and perform time-varying spectral feature identification on the audio time-frequency feature data according to the spectral centroid distribution map to obtain an audio time-varying spectral feature matrix; Step S232: Identify the main frequency period of the audio time-frequency feature data according to the audio time-varying spectral feature matrix to obtain a main frequency period sequence; Step S233: Calculate the time-scale ratio of the window to the signal between the main frequency period sequence and the preset initial analysis window length to obtain a window-signal time-scale ratio coefficient; Step S234: Evaluate the transient capture efficiency of the preset initial analysis window length according to the window-signal time-scale ratio coefficient to obtain a transient capture efficiency index; Step S235: Determine the window length adjustment according to the transient capture efficiency index for the preset initial analysis window length to obtain a window length adjustment factor, and perform dynamic range compression on the preset initial analysis window length according to the window length adjustment factor to obtain a preliminary adjusted window length; Step S236: Perform binary quantization on the preliminary adjusted window length to obtain an adjusted window length parameter.

5. The method for processing Bluetooth speaker data according to claim 3, characterized in that Step S25 includes the following steps: Step S251: Calculate the spectral energy density of the initial spectrogram to obtain a spectral energy density distribution map, and perform spectral peak detection on the initial spectrogram to obtain a spectral peak feature set; Step S252: Determine the energy threshold of the key frequency band for the initial spectrogram based on the spectral energy density distribution map and the spectral peak feature set to obtain the key frequency band energy threshold; Step S253: Mark and merge the high-energy frequency bands of the initial spectrogram according to the key frequency band energy threshold to obtain preliminary key frequency band mapping data; Step S254: Perform auditory weighted filtering on the initial spectrogram to obtain an auditory weighted spectrogram; Step S255: Perform a perceptual importance score on the preliminary key frequency band mapping data based on the auditory weighted spectrogram to obtain a key frequency band importance matrix; Step S256: Cluster the key frequency band importance matrix to obtain an optimized key frequency band set, and perform adaptive band-pass filtering on the optimized key frequency band set to obtain a key frequency band enhanced audio signal; Step S257: Perform time-frequency transformation on the key frequency band enhanced audio signal to obtain key frequency band enhanced time-frequency spectrum data.

6. The method for processing Bluetooth speaker data according to claim 3, wherein, Step S27 includes the following steps: Step S271: Perform logarithmic spectrum enhancement on the optimized spectrogram to obtain an enhanced logarithmic spectrogram; Step S272: Perform cepstral fundamental frequency estimation on the enhanced logarithmic spectrogram to obtain an audio fundamental frequency estimation sequence, and calculate the theoretical harmonic frequency of the optimized spectrogram based on the audio fundamental frequency estimation sequence to obtain a theoretical harmonic frequency map; Step S273: Perform harmonic peak detection on the optimized spectrogram according to the theoretical harmonic frequency map to obtain an actual harmonic feature set; Step S274: Calculate the harmonic energy proportion of the actual harmonic feature set to obtain a harmonic energy proportion matrix, and extract the time-varying harmonic structure of the actual harmonic feature set to obtain a time-varying harmonic structure sequence; Step S275: Generate harmonic structure data for the harmonic energy proportion matrix and the time-varying harmonic structure sequence based on the enhanced logarithmic spectrogram to obtain fundamental frequency harmonic structure data; where the fundamental frequency harmonic structure data includes fundamental frequency trajectory data, harmonic frequency sequence, and harmonic intensity matrix; Step S276: Identify the phase relationship between the fundamental frequency trajectory data and the harmonic frequency sequence to obtain a harmonic phase relationship map, and perform harmonic spatial reassignment on the fundamental frequency harmonic structure data according to the harmonic phase relationship map and the harmonic intensity matrix to obtain redirected harmonic structure data; Step S277: Perform harmonic synthesis on the redirected harmonic structure data to obtain a preliminary reconstructed audio signal, and perform envelope modulation on the preliminary reconstructed audio signal to obtain an envelope modulated audio signal; Step S278: Perform phase correction on the envelope modulated audio signal to obtain reconstructed audio data.

7. The method for processing Bluetooth speaker data according to claim 1, wherein Step S4 includes the following steps: Step S41: Deploy a multi-point microphone array in the environment where the Bluetooth speaker is located, and perform spatial sound field sampling and positioning to obtain a spatial sampling point distribution map. Perform broadband swept-frequency excitation on the environment where the Bluetooth speaker is located through the microphone array, and collect response signals to obtain multi-channel environmental response signals. Perform time-frequency identification on the multi-channel environmental response signals to obtain a time-varying spectrogram atlas; Step S42: Extract spatial acoustic parameters from the ambient sound field based on the time-varying spectrum atlas to obtain a spatial acoustic feature matrix; Step S43: Identify the reverberation features of the spatial acoustic feature matrix to obtain a reverberation feature vector; Step S44: Perform acoustic ray tracing simulation on the spatial sampling point distribution map to obtain a sound energy distribution model, and extract acoustic features and construct a database for the reverberation feature vector based on the sound energy distribution model to obtain an environmental acoustic feature database; Step S45: Perform spectrum compensation on the phase-optimized audio data based on the environmental acoustic feature database to obtain an environmentally compensated spectrum; Step S46: Perform blind deconvolution on the environmentally compensated spectrum to obtain a dereverberated environmentally compensated spectrum, and perform spatial sound image reshaping on the dereverberated environmentally compensated spectrum to obtain an audio space-optimized spectrum; Step S47: Perform an inverse phase correction transformation on the audio space-optimized spectrum to obtain an audio time-domain compensated signal, and perform loudness equalization on the audio time-domain compensated signal to obtain the final optimized audio data.

8. A system for processing Bluetooth speaker data, characterized in that, A method for performing Bluetooth speaker data processing as described in claim 1, the Bluetooth speaker data processing system comprising: A data compression module, configured to capture multi-modal data from a client to obtain original multi-modal client data; perform adaptive multi-resolution compression on the original multi-modal client data to obtain compressed multi-modal client data; An audio reconstruction module, configured to transmit the compressed multi-modal client data to a Bluetooth speaker through the client to obtain received compressed multi-modal data; decouple the received compressed multi-modal data to obtain separated audio data; perform dynamic spectrum reconstruction optimization on the separated audio data to obtain reconstructed audio data; An audio phase optimization module, configured to determine the type of music preference of a user to obtain a user music preference type identifier; perform a phase response on the reconstructed audio data according to the user music preference type identifier to obtain phase-optimized audio data; An environmental acoustic compensation module, configured to construct an environmental acoustic feature database for the environment where the Bluetooth speaker is located to obtain an environmental acoustic feature database; perform acoustic environment compensation on the phase-optimized audio data according to the environmental acoustic feature database to obtain the final optimized audio data.

9. A computer-readable storage medium storing a method for Bluetooth speaker data processing as described in any one of claims 1 to 7 that can be loaded and executed by a processor.

Citation Information

Patent Citations

  • Data transmission method and device, storage medium and terminal device

    CN114006890A

  • Sound effect adjustment method and device of sound box, equipment and storage medium

    CN117130576A

  • Audio parameter adjustment method and device, equipment and storage medium

    CN118250603A