Audio signal data management method and system of audio chip
By implementing adaptive filtering, detail enhancement, dynamic spectrum feature mining and distributed storage on the audio chip, efficiency and quality problems in multi-channel audio signal processing and transmission are solved, and efficient and stable audio signal management and transmission are achieved.
Patent Information
- Application Number
- CN202510335099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to efficiently manage the data flow of audio signals in real-time processing of multi-channel audio signals, and there are limitations to optimize the transmission and storage of audio signals in different application scenarios.
By identifying the real-time multi-channel audio signals and network monitoring parameters of the audio chip, adaptive filtering and noise reduction and detail enhancement processing are adopted to perform band-by-band analysis and dynamic spectrum feature mining, dynamic compression coding and lossless compression, a distributed stream storage framework is built, and audio channel identification and phase alignment, deep audio semantic analysis and transmission load demand mining are carried out to make audio signal transmission decisions dynamically.
It realizes efficient processing and transmission of multi-channel audio signals, reduces network transmission delay, ensures high-quality output of audio signals, optimizes the utilization efficiency of system resources, and adapts to the noise suppression and signal enhancement requirements in complex environments.
Smart Images

Figure CN120089149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly relates to a method and system for managing audio signal data of an audio chip. Background Art
[0002] With the continuous development of technology and the rapid progress of artificial intelligence technology, audio signal processing technology has been widely applied in various application fields. Especially in scenarios such as smart home, speech recognition, intelligent customer service, voice communication, virtual reality, and smart speakers, the quality and processing efficiency of audio signals are crucial. As the core component of audio processing, the audio chip undertakes the key tasks of audio signal acquisition, processing, conversion, and output. With the continuous progress of audio technology, the functions of modern audio chips are becoming increasingly complex, involving multiple aspects such as multi-channel audio processing, noise reduction, echo cancellation, and voice enhancement. Especially in the real-time processing of multi-channel audio signals, how to efficiently manage the data stream of audio signals and how to optimize the transmission and storage of audio signals in different application scenarios have become technical problems that need to be solved urgently.
[0003] Traditional methods for managing audio signal data usually rely on hardware-based signal processing means, and different algorithms are used to process different types of audio signals. However, with the development of multi-channel audio signal processing and network transmission technology, traditional methods have certain limitations in terms of signal processing accuracy, real-time performance, and utilization efficiency of system resources. For example, in the management of multi-channel audio signals, how to effectively reduce network transmission delay, ensure high-quality output of audio signals, and how to deal with noise suppression and signal enhancement in complex environments have become urgent problems to be solved. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a method and system for managing audio signal data of an audio chip to solve at least one of the above technical problems.
[0005] To achieve the above object, the present invention provides a method for managing audio signal data of an audio chip, including the following steps: Step S1: Identify the real-time multi-channel audio signals and real-time network monitoring parameters of the audio chip; perform adaptive filtering noise reduction and detail enhancement processing on the real-time multi-channel audio signals to obtain detail-enhanced real-time audio signals; Step S2: Define a per-band analysis window for the detail-enhanced real-time audio signals, and perform dynamic spectrum feature mining to generate dynamic audio signal features for each band; Step S3: Perform dynamic compression encoding and lossless compression audio encoding on the detail-enhanced real-time audio signals based on the dynamic audio signal features of each band, and construct a distributed stream storage framework; Step S4: Identify the audio channels of the real-time multi-channel audio signal and perform phase alignment processing to obtain a channel phase-aligned audio signal; Step S5: Perform in-depth audio semantic analysis on the channel phase-aligned audio signal and mine the transmission load requirements to generate application scenario transmission load requirement characteristics; Step S6: Make a dynamic audio signal transmission decision based on the application scenario transmission load requirement characteristics, and perform collaborative management optimization based on the distributed stream storage framework to build an intelligent audio collaborative management engine.
[0006] The present invention uses adaptive filtering and noise reduction processing to systematically remove background noise and unnecessary interference, thereby improving the clarity and quality of audio signals. This is particularly important for high-precision audio processing, especially in a multi-channel audio signal environment. Reducing noise and enhancing details can ensure the recognizability of the audio. Through detail enhancement, the tiny details in the audio can be retained and strengthened, thereby improving the perceived quality of the audio. Through frequency-by-frequency analysis, the system can deeply understand the characteristics of audio signals in different frequency bands, especially in complex audio signals, where different frequency bands contain different useful information. The extraction of dynamic spectrum features helps the system identify frequency band features and perform accurate audio signal processing. Dynamic spectrum feature mining reveals changes in frequency response and The frequency distribution of audio signals can be used to optimize the dynamic processing strategy of audio signals. Different noise reduction or gain strategies are adopted for different frequency bands. Through dynamic compression and lossless compression encoding, the system can save audio data with minimal storage requirements while avoiding loss of audio quality (lossless compression). Dynamic compression can adjust the compression strategy according to the real-time characteristics of the audio signal, reduce redundant data, and further improve transmission efficiency. Building a distributed stream storage framework can achieve efficient storage and distributed access of audio data, optimize the access speed and reliability of audio data, and ensure that multi-channel audio signals can be transmitted quickly and efficiently between different locations or devices. In multi-channel audio signals, the phases of each channel are different, resulting in audio signals. The incoordination of audio channels is eliminated through audio channel identification and phase alignment, ensuring that the signals of each channel are aligned on the time axis, thereby reducing audio distortion or mutual interference caused by phase differences. After phase alignment, the audio signal will be reconstructed more accurately, thereby reducing reverberation and phase distortion between different channels, ensuring that the audio restoration is more realistic. Through deep audio semantic analysis, the system not only analyzes the physical characteristics of the audio, but also understands its potential semantic content. In this way, when transmitting and processing audio, it makes intelligent decisions based on its content, optimizes the transmission path and processing strategy, and through load demand mining, it can optimize the transmission plan according to the audio characteristics of different application scenarios (such as audio bandwidth requirements, real-time requirements, etc.). This means that the system can optimize the transmission plan according to the audio characteristics of different application scenarios (such as audio bandwidth requirements, real-time requirements, etc.). The system dynamically adjusts the transmission bandwidth and delay according to the real-time demand of the frequency, thereby achieving more efficient utilization of network resources. Based on the dynamic characteristics of the audio signal and the needs of the application scenarios, the system can make intelligent audio transmission decisions, which not only improves the transmission efficiency of the audio signal, but also reduces the waste of network resources. Especially when the network bandwidth is limited, the audio transmission strategy is dynamically adjusted according to actual needs. Through the distributed stream storage framework and collaborative management engine, the processing, storage and transmission of audio data can collaborate more efficiently. Collaborative management optimization can ensure that the various modules within the system (such as audio acquisition, processing, storage and transmission) can interact in real time, automatically adjust and optimize the workflow, and improve the response speed and stability of the entire system.
[0007] Preferably, step S1 comprises the following steps: Step S11: Identify the real-time multi-channel audio signal and real-time network monitoring parameters of the audio chip; Step S12: Perform in-depth audio feature analysis on the real-time multi-channel audio signal to extract real-time audio depth features; Step S13: Based on the real-time audio depth features, perform refined signal classification to obtain audio background noise data; Step S14: Perform adaptive filtering and noise reduction on the audio background noise data to generate a noise-reduced and optimized audio signal; Step S15: Perform detail enhancement processing on the noise-reduced and optimized audio signal to obtain a detail-enhanced real-time audio signal.
[0008] In the present invention, by identifying the real-time multi-channel audio signal of the audio chip, the system can obtain the original data of each audio channel, which is the basis for subsequent signal processing, ensuring the comprehensiveness and accuracy of audio information. In-depth audio feature analysis extracts more detailed features (such as frequency, time domain, voice features, emotional color, etc.) from the audio signal, providing detailed information for subsequent signal processing. Through in-depth analysis of multi-channel signals, the system can identify different patterns in the audio (such as speech, music, background noise, etc.). This in-depth analysis effectively separates useful information and noise, providing key data for subsequent noise reduction, signal enhancement and other processing steps. Precise feature extraction helps to optimize the subsequent audio processing effect. Through refined classification of audio depth features, the system can distinguish and extract the components of background noise. Background noise is usually an interference source in the audio signal. Precise identification of noise components is crucial for subsequent noise reduction processing. The noise data obtained through classification can provide information for targeted processing of the adaptive filtering algorithm, making the noise reduction process more efficient and accurate. The adaptive filter can automatically adjust the filtering parameters according to the noise characteristics to remove or suppress unnecessary background noise. This process can significantly improve the clarity of the audio signal, especially for audio containing complex noise sources (such as traffic noise, wind noise, etc.). The advantage of adaptive filtering is that it can be dynamically adjusted according to the noise characteristics, so that while removing noise, it maximally retains useful audio information and avoids losing audio details due to overprocessing. By performing detail enhancement processing on the noise-reduced audio signal, the system further enhances the subtle differences in the audio signal, such as the spatial sense of sound, the richness of sound quality, etc., which significantly improves the auditory experience, especially for application scenarios such as speech recognition, music playback, etc.
[0009] Preferably, the specific steps of step S14 are as follows: Perform short time series segment division on the audio background noise data to obtain multiple segments of short time series background noise data; Perform power spectral density calculation on each segment of the short time series background noise data to generate the power spectral density of each segment; Perform amplitude envelope analysis on multi-segment short-time series background noise data, and extract the amplitude envelope features of each segment; Calculate the adaptive filtering parameters based on the amplitude envelope features of each segment and the power spectral density of each segment, so as to obtain the adaptive filtering parameters of each segment of noise; Perform real-time filtering and suppression processing on the audio background noise data according to the adaptive filtering parameters of each segment of noise, so as to generate a noise-reduced and optimized audio signal.
[0010] In the present invention, by dividing the audio signal into multiple short-time series segments, it helps to analyze and process the time-varying characteristics of noise more precisely. The characteristics of noise are different in different time periods. Therefore, dividing the audio background noise into short-time series segments helps to capture and process these time-varying characteristics, improving the pertinence and effect of noise reduction processing. The power spectral density is an important index describing the energy distribution of a signal in the frequency domain. Calculating the power spectral density segment by segment provides frequency domain characteristics for each time period, helping the system to understand the intensity of noise in different frequency ranges, so as to perform noise suppression more accurately. Calculating the power spectral density of each segment of the signal can reveal the characteristics of noise in different frequency bands. Through the frequency domain characteristics, it is possible to design filters better, especially to suppress the specific frequency components of noise and avoid unnecessary signal distortion. Amplitude envelope analysis reveals the change trend of the signal in time. Especially in the background noise signal, the amplitude envelope features can describe the dynamic changes of the noise. By extracting the amplitude envelope features of each segment, the system can identify the intensity change pattern of the noise. The adaptive filter dynamically adjusts its filtering parameters according to the characteristics of each segment of noise (such as amplitude envelope and power spectral density), so as to ensure that the filter processes different noise components more precisely. The frequency and amplitude of the noise change in different time periods. Therefore, adaptive parameter adjustment can respond to these changes in real time, ensuring that the noise reduction effect is always the best. Real-time filtering and suppression can dynamically remove noise in the audio signal, ensuring the clarity of the audio signal. In a changing noise environment, real-time processing can ensure that the noise reduction effect remains effective. Especially in real-time voice, radio and other application scenarios, it can maintain the stability of the audio quality. By performing filtering processing according to the adaptive filtering parameters of each segment of noise, the system accurately removes the noise without affecting the quality of the original audio signal. This ensures that the details and authenticity of the audio signal are retained, thus providing a higher quality audio experience.
[0011] Preferably, the specific steps of step S2 are as follows: Step S21: Perform multi-time-frequency decomposition on the detail-enhanced real-time audio signal to generate audio signals in multiple frequency bands; Step S22: Identify the transient frequency changes of the audio signals in multiple frequency bands, and identify the transient frequency change characteristics of the audio signals; Step S23: Perform signal stationarity analysis based on the transient frequency change characteristics of the audio signal, so as to obtain the stationarity characteristics of the audio signal; Step S24: Define analysis windows for each frequency band based on the stationarity characteristics of the audio signal, and generate the analysis window length for each frequency band; Step S25: Mine the dynamic spectrum characteristics of the audio signals in multiple frequency bands based on the analysis window length of each frequency band, and generate the dynamic audio signal characteristics for each frequency band.
[0012] The present invention performs multi-time-frequency decomposition on the real-time audio signal with enhanced details, and the system can decompose the signal into components of different frequency bands and further analyze the characteristics of each frequency band. After the audio signal is divided into multiple frequency bands, each frequency band can be processed in a targeted manner. Audio signals in different frequency bands have different characteristics, and segmented processing can improve the control and optimization effect of audio details, especially in tasks such as noise reduction and signal enhancement. Transient frequency change recognition can capture sudden changes or rapid changes in audio signals, such as syllable changes in speech or strong and weak fluctuations in music. The changing characteristics of transient frequency are crucial to the time domain changes and emotional expression of audio signals, and can improve the precision of audio processing. For dynamically changing audio signals (such as speech, music, etc.), transient frequency change recognition can provide real-time feedback on the changing trend of the signal, helping the system to better understand the structure and behavior of the audio signal. The recognition of transient features provides an important basis for subsequent processing steps (such as signal stability analysis and dynamic spectrum feature mining). The stability of an audio signal refers to whether the signal remains constant over a certain period of time, and a stable signal usually has relatively fixed statistical characteristics. Through signal stationarity analysis, we can understand the time-varying nature of the audio signal and identify which parts belong to the stationary segment and which parts are transient or changing segments. The stationarity characteristics of the signal guide the analysis window settings in the subsequent steps. When the signal is stationary, a longer analysis window is used to capture more frequency domain information, while a shorter window is used for rapid response in the transient part. This helps to achieve refined signal analysis and improve the accuracy of overall audio processing. According to the stationarity characteristics of each frequency band, the system defines a suitable analysis window length for each frequency band. For stationary signals, a longer analysis window is selected to capture more information; for transient parts or signals with drastic frequency changes, a shorter window is selected for rapid response. This method of dynamically adjusting the window length can accurately process audio signals with different characteristics and improve processing efficiency. Dynamic windows are applied in each frequency band for spectrum analysis to accurately mine the dynamic spectrum characteristics of the audio signal. These dynamic features can reflect the changing trend of the signal and the time-varying characteristics of the frequency components, which are of great significance for the understanding and optimization of audio signals. Dynamic spectrum feature mining enables the system to adjust the processing strategy in real time according to the actual needs of each frequency band. In this way, the fast-changing parts of the audio (such as transients and burst sounds) can be processed more efficiently, while the stable parts can be carefully analyzed and optimized. This is important for the efficient management and real-time optimization of audio signals.
[0013] Preferably, step S3 specifically comprises the following steps: Step S31: performing audio bandwidth classification on the detail-enhanced real-time audio signal based on the dynamic audio signal characteristics of each frequency band, and generating a plurality of audio spectrum sub-bands, wherein the plurality of audio spectrum sub-bands include high frequency band sub-bands and low frequency band sub-bands; Step S32: Calculate the frequency amplitude change of the high-frequency sub-band to generate the amplitude change feature of the high-frequency sub-band; Step S33: Dynamically adjust the coding rate and compression ratio of the amplitude change feature of the high-frequency sub-band to generate dynamic compression ratio parameters; Step S34: Perform dynamic compression coding on the high-frequency sub-band based on the dynamic compression ratio parameters to obtain dynamic compressed audio coding; Step S35: Perform lossless compression coding on the low-frequency sub-band to generate lossless compressed audio coding; Step S36: Store the lossless compressed audio coding and the dynamic compressed audio coding in a distributed coding stream to construct a distributed stream storage framework.
[0014] In the present invention, by classifying the bandwidth of the audio signal, the signal is divided into multiple spectral sub-bands, and more detailed processing is performed according to the characteristics of different frequency bands. The high-frequency band and the low-frequency band usually have different dynamic characteristics and processing requirements. Separating the processing helps to improve the overall signal optimization effect. The audio signal in the high-frequency band (such as the clarity in speech or the brightness in music) usually has a relatively fast amplitude change. Calculating the amplitude change of the high-frequency band accurately captures these dynamic changes and identifies the detailed changes in the signal. According to the amplitude change characteristics of the high-frequency band, the coding rate and compression ratio are dynamically adjusted to make the compression process more intelligent. For the part with a large amplitude change (i.e., the important audio feature), a lower compression ratio is used to ensure that the audio quality is not affected; for the part with a small change, a higher compression ratio is used to save storage space and bandwidth. Through the dynamic compression ratio parameters based on the amplitude change characteristics, the system can perform more accurate compression coding on the high-frequency band audio signal. This process ensures that while compressing the audio, the impact on the audio quality is minimized, especially in the retention of high-frequency details. The low-frequency band audio (such as the bass part, low-frequency noise, etc.) has a greater impact on the audio quality. Therefore, using lossless compression coding can completely retain the quality of the low-frequency signal. This is because the low-frequency band is usually crucial for the auditory experience, and any distortion affects the naturalness and depth of the audio. Lossless compression can ensure that no audio information is lost during the compression process, especially in the low-frequency band, effectively preventing the degradation of audio quality and maintaining the integrity of the audio. By storing the audio signals of lossless compression coding and dynamic compression coding in a distributed manner, the system can better utilize network and storage resources. Distributed storage can provide better redundancy backup to ensure data security. At the same time, through the mechanism of distributed stream storage, the efficiency of data access and processing is improved. The distributed stream storage framework can distribute and access audio data between different storage nodes, which not only improves the utilization rate of storage capacity but also provides a more efficient solution for the processing, management, and transmission of large-scale audio data. The system can dynamically adjust the storage location according to needs to optimize the storage and transmission efficiency.
[0015] Preferably, the specific steps of step S4 are as follows: Step S41: Identify the audio channels of the real-time multi-channel audio signal and extract each audio channel; Step S42: Calculate the phase difference for each audio channel to generate an inter-channel phase difference parameter; Step S43: Perform unified phase compensation calculation based on the inter-channel phase difference parameter to obtain the phase compensation value for each channel; Step S44: Perform phase alignment processing on the real-time multi-channel audio signal based on the phase compensation value of each channel to obtain a channel-phase-aligned audio signal.
[0016] In the present invention, by identifying and extracting each audio channel, subsequent processing (such as phase alignment, noise reduction, enhancement, etc.) is performed on independent channels, which is more flexible and precise. For complex audio signals (such as stereo or multi-channel audio), this channel-level processing method can improve the efficiency and quality of overall audio management. In a multi-channel audio system, the audio signals of different channels may have phase mismatches due to different hardware, transmission delays, or environmental factors. By calculating the phase difference of each audio channel, the phase mismatches between channels can be accurately identified to ensure signal synchronization. The phase difference will affect the overall quality of the audio signal, especially the spatial positioning and clarity of the audio. By identifying the phase difference between channels, it can provide a basis for subsequent phase compensation and alignment, which helps to restore the correctness and consistency of the audio signal. By calculating and applying phase compensation, the phase offsets between channels due to hardware differences, delays, or transmission are eliminated, so that the audio signals of each channel are resynchronized. Phase compensation can ensure that each channel of the audio signal is at the same phase reference point, thus avoiding sound quality distortion caused by phase differences. By performing phase alignment processing on each audio channel, the audio signals of all channels will achieve phase consistency, which not only ensures the coordination of the audio signal but also prevents sound blurring or interference caused by inconsistent phases of different channels.
[0017] Preferably, the specific steps of step S5 are as follows: Step S51: Perform in-depth audio semantic analysis on the channel-phase-aligned audio signal to generate audio signal semantic features; Step S52: Identify the application scenarios of multiple audio segments based on the audio signal semantic features to generate the application scenarios of multiple audio segments; Step S53: Analyze the audio scene transmission requirements of the application scenarios of multiple audio segments to obtain the audio transmission requirements of each application scenario; Step S54: Mine the transmission load requirements of the audio transmission requirements of each application scenario to generate application scenario transmission load requirement features. Through in-depth audio semantic analysis, the system of the present invention can identify different components in the audio signal, such as speech, music, environmental noise, etc., and extract the semantic features of these components, so as to deeply understand the actual content of the audio signal, thereby providing a basis for subsequent processing. The semantic features of the audio signal provide a more detailed basis for audio classification for the system, enabling different processing based on the content type, such as setting different encoding, compression, or transmission strategies for speech signals and music signals respectively. Through the semantic features of the audio signal, the system automatically identifies the scenarios of different audio segments in actual applications, such as voice calls, music playback, environmental noise detection, etc. This automated scenario recognition can significantly improve the adaptability and processing efficiency of the system in different application scenarios. According to different application scenarios, the system formulates personalized processing strategies for each audio segment. Speech signals require higher clarity and low-latency processing, while music signals focus more on sound quality fidelity and dynamic range. This provides a guarantee for the efficient management of audio signals in different scenarios. Through the analysis of the application scenarios of each audio segment, the system evaluates the transmission requirements of the audio signal in different scenarios. Voice communication has lower requirements for bandwidth but higher requirements for latency and clarity, while music streaming requires higher bandwidth to ensure sound quality. This step helps to determine the specific transmission requirements of each application scenario. By accurately analyzing the transmission requirements of each application scenario, the system reasonably allocates bandwidth in different application scenarios, avoiding resource waste. In the case of limited bandwidth, the system preferentially guarantees high-priority applications (such as real-time voice communication) while appropriately reducing the transmission quality or bandwidth occupancy of other applications. By mining the characteristics of transmission load requirements, the system predicts the load conditions of network and hardware resources according to the audio transmission requirements of different application scenarios. This helps to understand in advance the load generated by each scenario during transmission, avoiding overload or resource waste. Through the mining of load requirements, the system can dynamically schedule network and storage resources based on the actual situation. During peak network hours, the system automatically adjusts the transmission quality of low-priority scenarios to ensure the real-time performance and reliability of high-priority applications.
[0018] Preferably, the specific steps of step S6 are as follows: Step S61: Calculate the bandwidth utilization rate of the real-time network monitoring parameters to obtain the current network bandwidth utilization rate; Step S62: Analyze the transmission delay of the real-time network monitoring parameters to generate real-time network transmission delay parameters; Step S63: Evaluate the real-time network status of the current network bandwidth utilization rate and the real-time network transmission delay parameters to obtain the real-time network status characteristics of the audio chip; Step S64: Make a dynamic audio signal transmission decision on the application scenario transmission load requirement characteristics based on the real-time network status characteristics of the audio chip to generate a dynamic transmission strategy; Step S65: Perform collaborative management optimization according to the dynamic transmission strategy and the distributed stream storage framework, and construct an intelligent audio collaborative management engine.
[0019] Through calculating the bandwidth utilization rate, the system can understand the current network bandwidth usage in real time, which helps predict the possibility of network congestion and avoid audio signal transmission under saturated bandwidth, thus avoiding the decline of audio quality. The calculation of bandwidth utilization rate provides a basis for network resource allocation. The system can reasonably schedule resources according to the actual bandwidth utilization situation to ensure that audio signals can be transmitted when the network conditions are optimal, and avoid packet loss or delay of audio signals caused by excessive bandwidth usage. Real-time transmission delay analysis helps identify delay problems existing in network transmission, such as delay fluctuations caused by factors such as network congestion, unstable routing, and hardware failures. By generating delay parameters, the system can accurately reflect the real-time status of network transmission. For audio transmission applications that require low latency (such as voice calls, online meetings, etc.), delay analysis can timely detect network bottlenecks and provide real-time feedback to ensure that audio signals reach the destination within the required time window, thus guaranteeing communication quality. By combining the bandwidth utilization rate and delay parameters, the system can comprehensively evaluate the overall state of the current network. This comprehensive evaluation helps identify potential transmission problems and provides valuable network condition information for subsequent decision-making. Real-time network state evaluation provides accurate network quality data, enabling the audio chip to dynamically adjust transmission parameters according to the current network conditions. This can effectively cope with network fluctuations and ensure the reliability and stability of audio signal transmission. According to the real-time network state characteristics of the audio chip, the system can make dynamic decisions on the transmission requirements of different application scenarios. If the network condition is poor, the system automatically reduces the audio quality or adjusts the compression algorithm to reduce network load and avoid packet loss or delay; if the network bandwidth is sufficient, it gives priority to ensuring audio quality. Through real-time monitoring and analysis of the network state, the system can intelligently adjust the audio transmission strategy, such as dynamically selecting the encoding format, adjusting the bit rate, or optimizing the packet size according to bandwidth and delay, so that the audio signal can still maintain the best transmission effect under changing network conditions. Based on the dynamic transmission strategy and the distributed stream storage framework, the system can optimize the storage and transmission management of audio data. Through collaborative management, the load of different storage nodes is dynamically adjusted according to the real-time network conditions to avoid performance degradation or data loss caused by node overload. The distributed storage framework can improve data redundancy and availability. When the network load is high or a certain node fails, the system quickly switches to the standby storage node to ensure the transmission stability and high availability of audio data.
[0020] In this specification, an audio signal data management system for an audio chip is provided, which is used to execute the audio signal data management method of the audio chip as described above, including: An audio enhancement module, which is used to identify the real-time multi-channel audio signals and real-time network monitoring parameters of an audio chip; perform adaptive filtering and noise reduction and detail enhancement processing on the real-time multi-channel audio signals to obtain detail-enhanced real-time audio signals; A dynamic audio feature module, which is used to define analysis windows for each frequency band of the detail-enhanced real-time audio signals, and perform dynamic spectrum feature mining to generate dynamic audio signal features for each frequency band; A compression coding module, which is used to perform dynamic compression coding and lossless compression audio coding on the detail-enhanced real-time audio signals based on the dynamic audio signal features of each frequency band, and construct a distributed stream storage framework; A phase alignment module, which is used to identify audio channels of the real-time multi-channel audio signals and perform phase alignment processing to obtain channel-phase-aligned audio signals; A load demand module, which is used to perform in-depth audio semantic analysis on the channel-phase-aligned audio signals and perform transmission load demand mining to generate application scenario transmission load demand features; A collaborative management module, which is used to make dynamic audio signal transmission decisions on the application scenario transmission load demand features according to the application scenario transmission load demand features, and perform collaborative management optimization based on the distributed stream storage framework to construct an intelligent audio collaborative management engine.
[0021] The present invention removes interfering noises in real time through adaptive filtering to ensure that the audio signal remains clear in complex environments. Detail enhancement can improve the dynamic range and spatial sense of the audio, enhancing the user's auditory experience. By combining noise reduction and detail enhancement, the impact of environmental noise on the audio signal is reduced, thereby improving the accuracy of speech recognition systems or voice communication systems. Whether in a noisy environment or in an environment with high-quality audio requirements, the audio enhancement module can effectively optimize the signal quality. Through per-band analysis, the system extracts fine-grained features in the audio signal, enhancing the understanding and processing ability of the audio signal. This helps to process different types of audio signals (such as speech, music, environmental sounds) more precisely. The dynamic spectrum feature mining provides rich feature data support for subsequent processing such as compression coding, noise reduction, and enhancement, and can accurately adapt to different audio requirements. The combination of dynamic compression coding and lossless compression enables efficient transmission of audio data in different frequency bands under limited network bandwidth, reducing bandwidth pressure. For application scenarios with high requirements for sound quality (such as music or high-fidelity audio), lossless compression ensures that the sound quality will not degrade due to compression, meeting the needs of high-quality audio transmission. Phase alignment processing ensures that there are no phase differences and interferences in the multi-channel audio signal during audio output, making the audio output clearer and more stereo, suitable for use in stereo or surround sound systems. Performing phase alignment processing on multi-channel signals helps to reduce phase distortion and echo phenomena in the audio, improving the auditory experience. In terms of the spatial positioning and directivity of the audio signal, phase alignment helps to improve the positioning accuracy of the audio, enabling users to obtain a better sense of immersion in applications such as virtual reality and audio positioning. By deeply analyzing the semantic features of the audio signal, the system can identify different requirements for audio transmission in different scenarios. Voice communication has high requirements for latency, while music streaming is more concerned about the fidelity of sound quality. By mining the load requirements of the application scenario, the system can accurately predict the required resources, avoiding bandwidth overload or resource waste, thereby improving the transmission efficiency and quality. The dynamic transmission decision automatically adjusts the transmission strategy of the audio signal according to real-time network conditions, application scenarios, and system load, reducing human intervention and improving the processing efficiency. Through collaborative management and optimization, the system can dynamically balance the resource requirements of different application scenarios, avoid system overload, and at the same time ensure the timely and stable transmission of the audio signal. Through distributed storage and collaborative management optimization, the system better supports the transmission and management of audio signals across devices and network environments, improving the integration ability of multi-terminal systems. Brief Description of the Drawings
[0022] Figure 1 It is a schematic flow chart of the steps of a method for managing audio signal data of an audio chip according to the present invention; Figure 2 It is a schematic detailed implementation step flow chart of step S1; Figure 3It is a schematic diagram of the detailed implementation steps of step S2; Figure 4 It is a schematic diagram of the detailed implementation steps of step S3. Specific implementation manner
[0023] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] The embodiments of the present application provide a method and system for managing audio signal data of an audio chip. The execution subjects of the method and system for managing audio signal data of the audio chip include, but are not limited to, the following general computing nodes that carry the system: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. The data processing platform includes, but is not limited to, at least one of an audio image management system, an information management system, and a cloud data management system.
[0025] Please refer to Figures 1 to 4 , the present invention provides a method for managing audio signal data of an audio chip. The method for managing audio signal data of the audio chip includes the following steps: Step S1: Identify the real-time multi-channel audio signals and real-time network monitoring parameters of the audio chip; perform adaptive filtering and noise reduction and detail enhancement processing on the real-time multi-channel audio signals to obtain detail-enhanced real-time audio signals; Step S2: Define a per-band analysis window for the detail-enhanced real-time audio signals, and perform dynamic spectrum feature mining to generate dynamic audio signal features for each frequency band; Step S3: Perform dynamic compression coding and lossless compression audio coding on the detail-enhanced real-time audio signals based on the dynamic audio signal features of each frequency band, and construct a distributed stream storage framework; Step S4: Identify the audio channels of the real-time multi-channel audio signals and perform phase alignment processing to obtain channel-phase-aligned audio signals; Step S5: Perform in-depth audio semantic analysis on the channel-phase-aligned audio signals and perform transmission load demand mining to generate application scenario transmission load demand features; Step S6: Make a dynamic audio signal transmission decision based on the application scenario transmission load demand features, and perform collaborative management optimization based on the distributed stream storage framework to construct an intelligent audio collaborative management engine.
[0026] The present invention uses adaptive filtering and noise reduction processing to systematically remove background noise and unnecessary interference, thereby improving the clarity and quality of audio signals. This is particularly important for high-precision audio processing, especially in a multi-channel audio signal environment. Reducing noise and enhancing details can ensure the recognizability of the audio. Through detail enhancement, the tiny details in the audio can be retained and strengthened, thereby improving the perceived quality of the audio. Through frequency-by-frequency analysis, the system can deeply understand the characteristics of audio signals in different frequency bands, especially in complex audio signals, where different frequency bands contain different useful information. The extraction of dynamic spectrum features helps the system identify frequency band features and perform accurate audio signal processing. Dynamic spectrum feature mining reveals changes in frequency response and The frequency distribution of audio signals can be used to optimize the dynamic processing strategy of audio signals. Different noise reduction or gain strategies are adopted for different frequency bands. Through dynamic compression and lossless compression encoding, the system can save audio data with minimal storage requirements while avoiding loss of audio quality (lossless compression). Dynamic compression can adjust the compression strategy according to the real-time characteristics of the audio signal, reduce redundant data, and further improve transmission efficiency. Building a distributed stream storage framework can achieve efficient storage and distributed access of audio data, optimize the access speed and reliability of audio data, and ensure that multi-channel audio signals can be transmitted quickly and efficiently between different locations or devices. In multi-channel audio signals, the phases of each channel are different, resulting in audio signals. The incoordination of audio channels is eliminated through audio channel identification and phase alignment, ensuring that the signals of each channel are aligned on the time axis, thereby reducing audio distortion or mutual interference caused by phase differences. After phase alignment, the audio signal will be reconstructed more accurately, thereby reducing reverberation and phase distortion between different channels, ensuring that the audio restoration is more realistic. Through deep audio semantic analysis, the system not only analyzes the physical characteristics of the audio, but also understands its potential semantic content. In this way, when transmitting and processing audio, it makes intelligent decisions based on its content, optimizes the transmission path and processing strategy, and through load demand mining, it can optimize the transmission plan according to the audio characteristics of different application scenarios (such as audio bandwidth requirements, real-time requirements, etc.). This means that the system can optimize the transmission plan according to the audio characteristics of different application scenarios (such as audio bandwidth requirements, real-time requirements, etc.). The system dynamically adjusts the transmission bandwidth and delay according to the real-time demand of the frequency, thereby achieving more efficient utilization of network resources. Based on the dynamic characteristics of the audio signal and the needs of the application scenarios, the system can make intelligent audio transmission decisions, which not only improves the transmission efficiency of the audio signal, but also reduces the waste of network resources. Especially when the network bandwidth is limited, the audio transmission strategy is dynamically adjusted according to actual needs. Through the distributed stream storage framework and collaborative management engine, the processing, storage and transmission of audio data can collaborate more efficiently. Collaborative management optimization can ensure that the various modules within the system (such as audio acquisition, processing, storage and transmission) can interact in real time, automatically adjust and optimize the workflow, and improve the response speed and stability of the entire system.
[0027] In the embodiment of the present invention, refer to Figure 1, which is a schematic diagram of the step flow of a method of the present invention. In this example, the steps of the method include: Step S1: Identify the real-time multi-channel audio signal and real-time network monitoring parameters of the audio chip; perform adaptive filtering and noise reduction and detail enhancement processing on the real-time multi-channel audio signal to obtain a detail-enhanced real-time audio signal; In this embodiment, multi-channel audio signals are collected in real time to ensure that the sampling rate (e.g., 48 kHz) and bit depth (e.g., 24 bits) of the audio signals meet the standards. Configure audio acquisition software (such as LabVIEW or MATLAB) to monitor multiple channels to ensure that the data of each channel can be transmitted and processed in real time. Implement real-time network monitoring, use network analysis tools (such as Wireshark or Nagios) to obtain network parameters, including bandwidth utilization, latency, packet loss rate, etc. Set the monitoring period (e.g., collect once per second) to obtain the real-time network status, and synchronize the network monitoring parameters with the audio signal acquisition for subsequent analysis to ensure that the network condition is grasped while the audio is being processed. Store the collected audio signals and network monitoring parameters in the memory buffer to ensure that the data can be processed in a timely manner. Use data structures (such as arrays or data frames) to store the audio signals of each channel and the corresponding network parameters for subsequent processing. Select a suitable adaptive filtering algorithm, such as the least mean square error (LMS) or recursive least squares (RLS). These algorithms can dynamically adjust the parameters of the filter according to the characteristics of the input signal to achieve efficient noise suppression. Determine the order and learning rate of the filter. These parameters will affect the convergence speed and stability of the filter. In the experiment, the filter order is set to 32 and the learning rate is set to 0.01. First, estimate the background noise by analyzing the silent segments of the audio signal to obtain the noise characteristics. Calculate the power spectral density for each audio channel to determine the spectral characteristics of the background noise. Apply the adaptive filter to the real-time audio signal for noise reduction. The input of the filter is the current audio signal and the estimated noise signal, and the output is the filtered audio signal. Monitor the quality of the audio signal after noise reduction, use the signal-to-noise ratio (SNR) as the evaluation index to ensure that the noise reduction effect is significant. Visualize by plotting waveform and spectrum diagrams to intuitively show the effects before and after noise reduction. Adjust the filter parameters according to the real-time feedback to ensure that the system can adaptively respond to environmental changes. Select a suitable detail enhancement technique, such as high-pass filtering or dynamic range compression. High-pass filtering can enhance the high-frequency components to make the audio signal clearer, while dynamic range compression balances the volume and enhances the sense of detail. Determine the cut-off frequency of the high-pass filter (e.g., 300 Hz) to avoid over-enhancing the low-frequency noise. Apply high-pass filtering to the audio signal after adaptive filtering and noise reduction to extract the high-frequency components. Use digital signal processing tools (such as SciPy) to implement the design and application of the high-pass filter. Further apply dynamic range compression, set the compression ratio (e.g., 4:1) and threshold to enhance the dynamic performance of the signal and ensure that the bass part is still clearly audible. Store the real-time audio signal after detail enhancement in the memory and prepare to output it to the audio playback device or subsequent processing module. Select to save the processed audio signal in the WAV format for subsequent analysis and use. Monitor the quality of the processed audio signal to ensure that it meets the application requirements, such as clarity, loudness, and sound quality.
[0028] Step S2: Define a per-band analysis window for the detail-enhanced real-time audio signal, and perform dynamic spectrum feature mining to generate the dynamic audio signal features for each frequency band; In this embodiment, the size and overlap degree of the analysis window are determined. Usually, the selected window size is 1024 samples, and the overlap rate is set to 50% (i.e., 512 samples). Such a setting ensures the resolution of the spectrum analysis and can effectively capture the dynamic changes of the signal. The selection of the window function is also important. Commonly used window functions include the Hamming window or the Hanning window. These window functions can reduce the spectrum leakage phenomenon and improve the accuracy of the frequency-domain analysis. Apply the defined window function to the detail-enhanced real-time audio signal. Specifically, when implementing, use the signal.hamming() or signal.hanning() function in the NumPy or SciPy library to generate the window, and apply the window to each segment of the audio signal by multiplying them segment by segment. Slice the enhanced real-time audio signal according to the set window size. During the slicing process, it is necessary to ensure that there is an overlap between each slice so that the subsequent spectrum analysis can have a smooth transition. Record the start time and end time of each slice so that the subsequent feature extraction can correspond to the original signal. Apply the fast Fourier transform (FFT) to each slice to convert the time-domain signal into a frequency-domain signal. Use the np.fft.fft() function in NumPy to perform the FFT calculation, ensuring that the length of the FFT is consistent with the window size to ensure the accuracy of the spectrum. Calculate the power spectral density (PSD) of each frequency band for subsequent feature extraction. Use the np.abs() function to calculate the magnitude of the complex spectrum and square it to obtain the power spectrum. Extract dynamic features from the spectrum, such as peak frequency, spectral centroid, and spectral flatness. Store the dynamic spectrum features of each slice in structured data, such as a dictionary or a data frame (using the Pandas library). Record the feature values of each frequency band and associate them with timestamps for subsequent analysis and application. Ensure that the feature data can reflect the dynamic changes of the audio signal, which is convenient for subsequent audio analysis and processing. Use visualization libraries such as Matplotlib to plot the dynamic spectrum features into charts for intuitive analysis. Plot spectrograms, time-frequency diagrams, etc. to help understand the frequency-domain characteristics of the audio signal. Analyze the feature changes in different frequency bands to identify which frequency bands change significantly during a specific time period for more in-depth audio feature analysis. Provide feedback for subsequent audio processing according to the analysis results of the spectrum features, and adjust the audio signal processing strategy to improve the quality. Record the abnormal situations that occur during the analysis process and make corresponding adjustments to ensure the accuracy and reliability of the feature extraction.
[0029] Step S3: Perform dynamic compression encoding and lossless compression audio encoding on the detail-enhanced real-time audio signal based on the dynamic audio signal features of each frequency band, and construct a distributed stream storage framework;In this embodiment, a suitable dynamic compression coding algorithm is selected, such as Adaptive Bitrate Coding (ABR) or Dynamic Range Compression (DRC) based on spectral characteristics. In this step, we adopt the DRC method to dynamically adjust the coding parameters according to the real-time audio characteristics, set the compression ratio and threshold parameters. The compression ratio is set to 4:1 and the threshold is -10 dB to ensure that the data volume can be effectively reduced while maintaining the audio quality. Analyze the dynamic audio signal characteristics of each frequency band, identify key features (such as peaks, spectral centroids, etc.), and adjust the compression strategy according to these features. Use audio processing libraries such as librosa for dynamic range compression. During the processing, first detect the instantaneous amplitude of the signal, dynamically adjust the gain according to the set threshold, and encode the audio signal after dynamic compression into a format suitable for transmission and storage (such as AAC or Opus). Select appropriate coding parameters, such as bitrate (e.g., 128 kbps) and sampling rate (e.g., 48 kHz), to ensure the quality of the compressed audio signal. Store the encoded audio data in memory or write it to a file for subsequent lossless compression processing. Select a suitable lossless compression coding format, such as FLAC or ALAC, which can effectively reduce the data volume without loss of audio quality. Determine the compression parameters, such as setting the compression level of FLAC (0 - 8), and minimize the encoding and decoding time delay while ensuring the compression efficiency. Select the compression level as 5 to achieve a better balance between compression effect and efficiency. Use pyFLAC or other lossless audio coding libraries to encode the audio signal with enhanced details. Specifically, during implementation, first read the audio data after dynamic compression, and then apply the lossless coding algorithm for compression. During the encoding process, record parameters such as the sampling rate and number of channels of the original audio to ensure that the original quality can be restored during decoding. Save the audio data after lossless compression to a distributed storage system, such as using Amazon S3, Google Cloud Storage, or a self-built HDFS. Through object storage, more efficient data access and management can be achieved. Set the data storage strategy, such as storing by timestamp or frequency band classification, for subsequent retrieval and analysis. Design a distributed stream storage framework to manage the audio data after dynamic compression and lossless compression. Select a suitable stream processing technology, such as Apache Kafka or Apache Pulsar, to support the real-time stream processing and management of audio data. Determine the data stream architecture, including data producers (audio signal sources), data processing nodes (audio encoding and storage), and data consumers (audio playback and analysis modules). Integrate the dynamic compression and lossless compression audio coding modules with the distributed storage framework to ensure that after the audio data is encoded, it can be automatically pushed to the distributed storage system. Configure the parameters of data transmission, such as batch size, latency time, etc., to optimize the performance of the data stream. Set the batch size to 1MB and the transmission latency to 100 ms.To ensure smooth audio transmission, a real-time monitoring system is implemented to track the transmission and storage status of audio data, ensure data integrity and availability, use tools such as Prometheus and Grafana for monitoring, and optimize storage policies and encoding parameters according to real-time monitoring results to ensure that the system can still maintain stable performance under high load.
[0030] Step S4: Identify the audio channels of the real-time multi-channel audio signal and perform phase alignment processing to obtain a channel-phase-aligned audio signal; In this embodiment, a multi-channel audio signal is collected in real time using an audio interface or audio processing software (such as PyAudio or Librosa). The sampling rate is set to 48 kHz to ensure that the audio signal of each channel can be accurately captured. The audio signals of each channel are stored as independent arrays or data frames for subsequent processing. Feature extraction is performed on each audio channel for channel identification. The extracted features include time-domain features of the audio (such as mean and variance) and frequency-domain features (such as MFCC and spectral centroid). The MFCC features are extracted using Librosa. Parameters such as the window length are set to 25 ms and the overlap rate is 50%. The MFCC features of each channel are extracted through the librosa.feature.mfcc() function and stored as a feature matrix. By comparing the extracted features, each channel is matched and identified. The dynamic time warping (DTW) algorithm is used to evaluate the similarity between different channels. The DTW algorithm can effectively process time series data and identify similar audio patterns. According to the recognition result, the identifier of each channel is determined for subsequent phase alignment processing. The phase difference between each channel is calculated. First, a short-time Fourier transform (STFT) is performed on each channel to obtain the frequency-domain representation. The STFT calculation is performed using the librosa.stft() function. The window length is set to 1024 samples and the overlap rate is 50%. The phase information of each frequency point is calculated, and the phase difference between channels is calculated through the phase difference formula: phase difference = arg(X1(f)) - arg(X2(f)), where X1 and X2 are the frequency-domain representations of different channels. According to the calculated phase difference, phase correction is performed on one of the channels using the method of phase delay compensation. The specific method is to apply the corresponding phase delay to the channel to be corrected, y(t) = x(t - τ), where τ is the calculated delay time. This operation is implemented using the lfilter function in the scipy.signal library. All the audio channels that have undergone phase alignment processing are combined into a stereo signal. The audio signals of each channel are combined into a multi-channel audio signal (such as stereo or surround sound) using NumPy. The final channel phase-aligned audio signal is output to an audio file and saved in a standard format (such as WAV or FLAC). The scipy.io.wavfile.write() function is used to implement the writing of the audio file. When saving, appropriate sampling rates (such as 48 kHz) and bit depths (such as 24 bits) are set to ensure the audio quality.
[0031] Step S5: Perform in-depth audio semantic parsing on the channel phase-aligned audio signal and conduct transmission load requirement mining to generate application scenario transmission load requirement features; In this embodiment, first, feature extraction is performed on the audio signals with aligned channel phases for subsequent semantic parsing. The extracted features include MFCC (Mel Frequency Cepstral Coefficients), pitch, loudness, etc. The librosa library is used to extract MFCC, with the window length set to 25 ms and the overlap rate set to 50%, and the feature matrix of each audio signal is extracted. To increase the depth of semantic parsing, other features are combined, such as timbre features and rhythm features. These features help to understand the content such as speech, music, or environmental sounds in the audio signal. A suitable deep learning model is selected for audio semantic parsing, and a convolutional neural network (CNN) or a recurrent neural network (RNN) is used to process the audio signals. When designing the network structure, the audio features are used as the input. For example, CNN is used for feature extraction, and then the extracted features are passed to the RNN to capture the temporal information. The model is trained using a labeled audio dataset (such as UrbanSound or LibriSpeech), and the training parameters are set, including the batch size (such as 32), the learning rate (such as 0.001) and the number of training rounds (such as 50 rounds) to ensure that the model can effectively learn the semantic information in the audio signal. Use the trained model to infer the audio signal with channel phase alignment, extract the semantic information in the audio, identify whether the audio signal is speech, music or environmental noise, and extract relevant context information. Record the parsing results in structured data, including information such as audio type, sentiment analysis, keywords, etc. for subsequent analysis. Based on the results of deep audio semantic parsing, establish a transmission load demand model. By analyzing different types of audio content and their corresponding network requirements, identify the sensitivity of the audio signal to bandwidth, latency, and packet loss rate. Set model parameters, such as bandwidth requirements (e.g., 128 kbps for speech, 320 kbps for high-quality music), latency requirements (speech requires less than 150 ms, music can accept up to 300 ms), etc. Collect the transmission performance of the audio signal in different network environments, record the latency, packet loss rate, and bandwidth utilization during audio transmission. This can be done by conducting experiments under different network conditions to collect the transmission data of the audio signal, analyze the performance of different audio types under specific network conditions, and conduct a comprehensive evaluation in combination with the results of semantic parsing. Use statistical analysis methods (such as regression analysis) to reveal the relationship between audio content and network requirements. Store the extracted application scenario transmission load demand characteristics in the database for subsequent optimization and decision-making. Record the bandwidth, latency, and packet loss rate requirements corresponding to each audio type to form a complete feature set. Combine real-time network monitoring data to create a dynamic load demand model to support audio transmission decisions under different network conditions. Verify the accuracy of the established transmission load demand model through experiments. Conduct multiple audio transmission experiments to measure whether the actual network performance is consistent with the requirements predicted by the model. Adjust the model parameters according to the verification results to improve the prediction accuracy. Use cross-validation techniques to ensure the robustness and accuracy of the model. According to the transmission load demand characteristics, optimize the transmission strategy of the audio signal. In case of insufficient bandwidth, select to reduce the audio quality or adjust the encoding parameters to ensure the smooth transmission of the audio signal. Implement real-time monitoring and dynamically adjust the transmission strategy to adapt to changes in network conditions.
[0032] Step S6: Make dynamic audio signal transmission decisions based on the application scenario transmission load demand characteristics, and perform collaborative management optimization based on the distributed stream storage framework to build an intelligent audio collaborative management engine.
[0033] In this embodiment, a dynamic audio signal transmission decision model is designed. Based on the transmission load demand characteristics of the application scenarios extracted previously, the model should consider the audio type, network conditions (bandwidth, latency, packet loss rate), and user requirements (such as audio quality requirements), select an appropriate decision algorithm, such as Support Vector Machine (SVM) or Decision Tree, set model parameters, such as feature selection (bandwidth, latency, packet loss rate, etc.) and the scale of the training set, to ensure the accuracy of the model. During the decision-making process, real-time network status data, including the current bandwidth, latency, and packet loss rate, is obtained, and network monitoring tools (such as Prometheus and Grafana) are used to monitor these parameters and transmit the data to the decision model. The feature data of the audio signal (such as audio type and quality requirements) is combined with the real-time network status and input into the decision model for dynamic decision-making. According to the output of the decision model, the transmission strategy of the audio signal is determined, including selecting an appropriate coding format (such as AAC or Opus), adjusting the bit rate (such as 128 kbps or 320 kbps), and selecting an appropriate transmission protocol (such as UDP or TCP). After implementing the decision, record the decision-making process and results for subsequent analysis and optimization. In a distributed stream storage system (such as Apache Kafka or Google Cloud Storage), implement a collaborative management strategy. According to the transmission decision and load demand characteristics of the audio signal, dynamically adjust the storage strategy, such as selecting an appropriate number of replicas and sharding strategy, determine the load balancing strategy of the storage nodes, ensure that under high load conditions, the audio signal can be transmitted and accessed quickly, use the consistent hashing algorithm to optimize data distribution, monitor the status of the audio data stream, ensure the consistency and availability of data in the storage system, use a stream processing framework (such as Apache Flink or Apache Spark) to process and analyze the audio data stream, calculate the performance metrics of storage and transmission in real-time. For network conditions with high latency or packet loss, automatically adjust the storage strategy, such as selecting a more efficient compression algorithm or adjusting the data storage location to reduce access latency. Design the architecture of the intelligent audio collaborative management engine, including a data input module (receiving audio signals and network status data), a decision module (executing transmission decisions), a storage management module (optimizing data storage and stream management), and a feedback module (collecting and analyzing transmission performance), determine the communication mechanism between each module, and use a message queue (such as RabbitMQ or Kafka) to achieve asynchronous communication between modules to ensure the efficiency and stability of the system. Integrate each module into a complete system and conduct system testing. By simulating different network conditions and audio signal types, verify the performance and decision accuracy of the collaborative management engine, record the system operation data, analyze the efficiency and quality of audio signal transmission, optimize system parameters, and continuously optimize the intelligent audio collaborative management engine according to the test results and real-time data feedback, and adjust the parameters and algorithms of the decision model.To improve the accuracy of transmission decisions and the overall performance of the system, regular system evaluations and updates are implemented to ensure that the engine can still operate effectively under changing network environments and user requirements.
[0034] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Step S11: Identify the real-time multi-channel audio signal and real-time network monitoring parameters of the audio chip; Step S12: Perform in-depth audio feature analysis on the real-time multi-channel audio signal to extract real-time audio depth features; Step S13: Based on the real-time audio depth features, perform refined signal classification to obtain audio background noise data; Step S14: Perform adaptive filtering and noise reduction on the audio background noise data to generate a noise-reduced and optimized audio signal; Step S15: Perform detail enhancement processing on the noise-reduced and optimized audio signal to obtain a detail-enhanced real-time audio signal.
[0035] In this embodiment, ensure that the audio chip is correctly connected to a multi-channel audio input device (such as a microphone array). Configure the audio chip, set appropriate sampling rates (such as 44.1 kHz or 48 kHz) and the number of channels (such as stereo or four-channel), start the audio acquisition program, and use an audio programming interface (such as PortAudio or JACK) to capture audio signals in real time. Ensure that the audio signals of each channel can be synchronously acquired. At the same time, monitor network performance parameters such as latency, bandwidth, and packet loss rate, which is achieved through network monitoring tools (such as Wireshark or SNMP). Ensure the integrity of network data. Integrate the real-time audio signals with the network monitoring parameters to form a multi-dimensional data set. Each data point should contain a timestamp, audio signal, and network parameters for subsequent analysis. Monitor the data acquisition process in real time, record any abnormal situations (such as signal distortion, increased network latency, etc.), ensure the stability of data acquisition, and perform troubleshooting if necessary. Select a suitable audio feature extraction method, such as Mel Frequency Cepstral Coefficients (MFCC), spectrogram features, or pitch analysis, etc. Select the most appropriate features according to the application scenario (such as speech recognition or music classification). Determine the parameter settings for feature extraction, such as the window function type (such as Hamming window or Hanning window), window length (such as 256 ms or 512 ms), and overlap rate (such as 50%). Use an audio processing library (such as Librosa) to process the real-time audio signals, implement the Short-Time Fourier Transform (STFT), and convert the time-domain signal into a frequency-domain signal. During the feature extraction process, extract the MFCC, spectral features, and time-domain features of each audio frame and store them as a feature matrix. Organize the extracted audio features into structured data for subsequent signal classification and analysis. Record the timestamp of each feature and the corresponding audio channel information. Select a suitable classification algorithm according to the nature of the feature data, such as Support Vector Machine (SVM), Random Forest, or Convolutional Neural Network (CNN). The selected method should be adaptable to multi-channel data and be able to handle high-dimensional feature inputs. Prepare the hyperparameter settings for the classification model, such as learning rate, regularization parameter, and number of training epochs. Use the labeled training data (including background noise and other audio signals) to train the selected classification model, ensuring the use of sufficient sample data to improve the generalization ability of the model. Evaluate the performance of the model by analyzing the classification results through cross-validation and confusion matrix to ensure that the model can accurately distinguish background noise from other audio signals. Apply the trained model to the real-time audio deep feature data, classify each audio frame, record the classification results, and identify the data belonging to background noise. Conduct a statistical analysis of the classification results to evaluate the frequency and duration of background noise for subsequent processing. Select a suitable adaptive filtering algorithm according to the background noise characteristics, such as the Least Mean Square (LMS) algorithm or the Recursive Least Squares (RLS) algorithm. The selected algorithm should be able to dynamically adapt to signal changes. Determine parameters such as the order and step size of the filter to optimize the noise reduction effect. Design an adaptive filter.And it is applied to real-time audio signals. Using the identified background noise as a reference signal, filtering is performed, and the convergence speed and stability during the filtering process are monitored to ensure that the filter can quickly adapt to changes in the input signal. The filtered signal is compared with the original signal, the noise reduction effect is recorded, the signal-to-noise ratio (SNR) and total harmonic distortion (THD) of the noise reduction result are evaluated, a noise-reduced and optimized audio signal is generated and saved as an audio file for subsequent playback and analysis. A suitable detail enhancement algorithm is selected, such as dynamic range compression, equalizer adjustment, or frequency-domain enhancement technology. The enhancement target (such as enhancing high frequencies or low frequencies) is determined according to the processing requirements, and the enhancement parameter settings, such as gain, frequency band range, and threshold, are determined. The selected detail enhancement algorithm is applied to the noise-reduced and optimized audio signal, and the enhancement parameters are adjusted in real time to ensure that the quality of the output signal meets the expectations. During the processing, the spectral changes of the audio signal are monitored to ensure the naturalness and clarity of the enhanced signal. Subjective and objective evaluations are performed on the enhanced audio signal, and the signal quality is analyzed using listening tests and signal processing tools to evaluate the effect of detail enhancement. Parameter adjustments are made according to the evaluation results to optimize the detail enhancement effect to ensure that the final output real-time audio signal reaches the best quality.,
[0036] In this embodiment, the specific steps of step S14 are as follows: The audio background noise data is divided into short time series segments to obtain multiple segments of short time series background noise data; The power spectral density of each segment is calculated for multiple segments of short time series background noise data to generate the power spectral density of each segment; The amplitude envelope analysis is performed on multiple segments of short time series background noise data to extract the amplitude envelope characteristics of each segment; Based on the amplitude envelope characteristics of each segment and the power spectral density of each segment, the adaptive filtering parameters are calculated to obtain the adaptive filtering parameters of each segment of noise; According to the adaptive filtering parameters of each segment of noise, the audio background noise data is subjected to real-time filtering and suppression processing to generate a noise-reduced and optimized audio signal.
[0037] In this embodiment, load the audio background noise data, use an audio processing library (such as Librosa or PySoundFile) to read the data, ensure that the data is stored in an appropriate format (such as WAV or FLAC), determine the sampling rate of the audio signal (such as 44.1 kHz) for reasonable time division in subsequent processing, set the length and overlap rate of short time segments. Usually, the length of short time segments is set to 100 ms to 500 ms, and the overlap rate is set to 50% or 75% to ensure sufficient overlap between each segment and capture the continuity of the signal. If the segment length is set to 200 ms and the overlap rate is 50%, the interval between each segment is 100 ms. Write code to implement the division of the audio signal into short time segments. By looping through the audio data and using the set segment length and interval, the audio signal is sliced into multiple short time segments, and record the start and end times of each short time segment to ensure accurate association of the data for each segment during subsequent analysis. Store the multiple divided short time segments in a list or array for subsequent processing. Each segment should retain its timestamp and original data for subsequent power spectral density calculation and amplitude envelope analysis. Select a suitable method for calculating the power spectral density, such as the Welch method. This method calculates the power spectral density by segmenting the signal, applying a window function, and performing a Fourier transform, which can effectively reduce spectral leakage. Set the window function type (such as Hamming window or Hanning window) and window length to determine the accuracy and stability of the calculation. Apply the power spectral density calculation to each stored short time segment. During the calculation process, segment using the set window function and overlap rate, record the power spectral density value of each segment, and store it in an array for subsequent analysis. Check the calculation results of the power spectral density for each segment to ensure that the spectral range and amplitude of each segment meet the expectations. Visualize by plotting the power spectral density graph to quickly identify anomalies. Select a suitable method for extracting the amplitude envelope, such as the Hilbert transform. This method can effectively extract the envelope characteristics of the signal by calculating the instantaneous amplitude of the signal. Determine the parameter settings for the analysis, such as the sampling frequency and the window length for envelope calculation, to ensure the accuracy of the extraction results. Apply the Hilbert transform to each short time segment to calculate its envelope signal. By performing a complex Fourier transform on the short time segment, obtain the instantaneous amplitude, and record the envelope characteristic data of each segment, including the maximum value, minimum value, and root mean square (RMS) value, etc. Store the extracted amplitude envelope characteristics in a structured data format for convenient subsequent use. Ensure that each segment's characteristics match its corresponding short time segment. According to the amplitude envelope characteristics and power spectral density, select a suitable adaptive filtering model. Commonly used models include the least mean square error (LMS) and the recursive least squares (RLS). Determine the hyperparameters such as the order of the filter and the learning rate to optimize the filtering effect. For each short time segment, calculate the adaptive filtering parameters by combining its amplitude envelope characteristics and power spectral density. Use the maximum value of the envelope and the mean value of the power spectrum to set the gain of the filter.Record the adaptive filtering parameters for each segment, including the step size and filter coefficients, etc., for subsequent real-time filtering applications. Verify the effectiveness of the calculated adaptive filtering parameters to ensure that they can adapt to different noise characteristics. Test through simulated data and evaluate their performance during the noise reduction process. Design a filter based on the adaptive filtering parameters and select a suitable implementation method, such as digital filter design tools (e.g., SciPy or MATLAB) to implement the adaptive filter. Determine the type of filter (e.g., FIR or IIR) and set the corresponding implementation parameters. Apply the calculated adaptive filter to each short time series segment to suppress background noise in real-time. Implement the filtering process by calling a real-time audio processing library (e.g., PyAudio). Monitor the quality of the filtered signal and record the signal-to-noise ratio (SNR) and other quality metrics of the signal. Save the filtered audio signal as a file (e.g., WAV format) or play it in real-time for subsequent analysis and use. Compare the audio signals before and after noise reduction processing to ensure that the noise reduction effect is significant and can meet the application requirements.
[0038] In this embodiment, refer to Figure 3 , which is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Step S21: Perform multi-time-frequency decomposition on the detail-enhanced real-time audio signal to generate audio signals in multiple frequency bands; Step S22: Identify the transient frequency changes of the audio signals in multiple frequency bands and identify the transient frequency change characteristics of the audio signals; Step S23: Perform signal stationarity analysis based on the transient frequency change characteristics of the audio signals to obtain the stationarity characteristics of the audio signals; Step S24: Define the analysis window for each frequency band based on the stationarity characteristics of the audio signals to generate the analysis window length for each frequency band; Step S25: Mine the dynamic spectral characteristics of the audio signals in multiple frequency bands based on the analysis window length of each frequency band to generate the dynamic audio signal characteristics of each frequency band.
[0039] In this embodiment, a real-time audio signal with enhanced loading details is used. An audio processing library (such as Librosa or PySoundFile) is used to read the audio data, ensuring that the data format is compatible (such as WAV or FLAC). The sampling rate of the audio signal (such as 44.1 kHz) is determined for reasonable time-frequency analysis in subsequent processing. A suitable time-frequency decomposition method is selected, such as the short-time Fourier transform (STFT), wavelet transform, or Hilbert-Huang transform (HHT). These methods can effectively analyze the frequency-domain characteristics of non-stationary signals. The parameter settings for the decomposition are determined, such as the window function type, window length, and overlap rate. A Hamming window is used, with the window length set to 512 points and the overlap rate set to 50%. The selected time-frequency decomposition method is applied to process the audio signal. During this process, the audio signal is decomposed into multiple frequency bands, and the frequency range and time information of each frequency band are recorded.
[0040] import numpy as np import librosa y, sr = librosa.load('audio_file.wav', sr=None) D = librosa.stft(y, n_fft=512, hop_length=256, window='hann') According to the decomposition results, the required frequency bands are extracted. The frequency bands are divided into low frequency (0 - 200 Hz), middle frequency (200 - 2000 Hz), and high frequency (2000 - 20000 Hz). The audio signals of each frequency band are recorded for subsequent analysis. A suitable transient frequency identification method is selected, such as transient frequency estimation or frequency modulation analysis. It is implemented using the Hilbert transform combined with frequency extraction technology. The analysis parameters are determined, such as the time window and frequency range for transient definition. The transient frequency change identification algorithm is applied to each extracted frequency band. The instantaneous phase of the signal is calculated using the Hilbert transform, and then the transient frequency is calculated through differentiation.
[0041] from scipy.signal import hilbert instantaneous_phase = np.angle(hilbert(y_segment)) instantaneous_frequency = np.diff(instantaneous_phase) * (sr / (2.0 *np.pi)) Record the transient frequency change characteristics of each frequency band, including the maximum value, minimum value, and change rate of the transient frequency, etc. Ensure that the characteristics of each frequency band are sorted out for subsequent analysis. Define the stationarity characteristics of the signal, such as mean, variance, frequency change rate, etc. Select a suitable stationarity analysis method, such as autocorrelation analysis or unit root test. Determine the length of the analysis window to effectively evaluate the stationarity in each frequency band. Conduct a stationarity analysis on the transient frequency change characteristics of each frequency band. Calculate the mean, variance, and autocorrelation coefficient of each frequency band to evaluate the stationarity of the signal. Record the analysis results, including the numerical values of the stationarity characteristics and the corresponding frequency band information. Organize the results of the stationarity analysis into a table or data frame for visualization and subsequent analysis. Ensure that the stationarity characteristics of each frequency band correspond to its transient frequency change characteristics. According to the analysis results of the stationarity characteristics, set the length of the analysis window for each frequency band. Generally, for frequency bands with better stationarity, a shorter window length is used, while for non-stationary frequency bands, a longer window length is required. Use empirical rules or set the window length according to the standard deviation of the stationarity characteristics to ensure that the window length can effectively capture the changes in the signal. For each frequency band, calculate the window length according to its stationarity characteristics. Set the window length to 100 ms, 200 ms or longer, depending on the characteristics of the signal. Record the window length of each frequency band and ensure that the appropriate window is used in subsequent analysis. Store the analysis window length of each frequency band in a structured data format for convenient use in subsequent dynamic spectrum feature mining. Select a suitable dynamic spectrum feature mining method, such as instantaneous energy, spectral centroid, spectral width, etc. Select the most suitable features according to the application scenario. Determine the parameter settings for mining, such as window length, overlap rate, etc., to ensure the effectiveness of feature extraction. For the audio signal of each frequency band, apply the set analysis window for dynamic spectrum feature mining. Use the short-time Fourier transform to calculate the spectrum of each window and extract the required dynamic features. Record the dynamic features of each frequency band, including instantaneous energy, spectral centroid, and spectral width, etc., to ensure that the data is complete and structured. Organize the mined dynamic audio signal features into a table or data frame for convenient subsequent analysis and visualization. Ensure that each feature matches the corresponding frequency band information.
[0042] In this embodiment, refer to Figure 4 , which is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Step S31: Classify the detail-enhanced real-time audio signal based on the dynamic audio signal features of each frequency band to generate a plurality of audio spectrum sub-bands, and the plurality of audio spectrum sub-bands include high-frequency sub-bands and low-frequency sub-bands; Step S32: Calculate the frequency amplitude change of the high-frequency sub-band to generate the amplitude change characteristics of the high-frequency sub-band; Step S33: Dynamically adjust the coding rate and compression ratio according to the amplitude change characteristics of the high-frequency sub-band, so as to generate dynamic compression ratio parameters; Step S34: Perform dynamic compression coding on the high-frequency sub-band based on the dynamic compression ratio parameters to obtain dynamically compressed audio coding; Step S35: Perform lossless compression coding on the low-frequency sub-band to generate lossless compressed audio coding; Step S36: Perform distributed coding stream storage on the lossless compressed audio coding and the dynamically compressed audio coding to construct a distributed stream storage framework.
[0043] In this embodiment, the frequency band division standard is set according to the audio characteristics. Generally, the frequency band is divided into low frequency (0 - 300 Hz), medium frequency (300 - 2000 Hz), and high frequency (above 2000 Hz). Dynamic features, such as spectral centroid, are used to judge the main frequency distribution of the audio signal, so as to more accurately determine the frequency band. Write code to implement the frequency band classification of the audio signal. According to the set frequency band range, the signal is assigned to the corresponding high-frequency sub-band and low-frequency sub-band, and the audio signal data of each frequency band is recorded to ensure the accuracy of the classification result.
[0044] low_band = signal[(frequency>= 0)&(frequency<300)] high_band = signal[(frequency>= 2000)&(frequency<20000)] Define the calculation method of amplitude change, such as the difference between instantaneous amplitude and average amplitude. Use the short-time Fourier transform (STFT) to calculate the amplitude change of the high-frequency sub-band. Determine the calculation parameters, such as window function type, window length, and overlap rate, to ensure the accuracy of the results. Apply STFT to the high-frequency sub-band and calculate the amplitude value of each time frame. Record the instantaneous amplitude and average amplitude of each time frame for subsequent amplitude change analysis.
[0045] D_high = librosa.stft(high_band, n_fft=512, hop_length=256) amplitude = np.abs(D_high) Generate the amplitude change characteristics of the high-frequency sub-band according to the calculation results, including the maximum amplitude, minimum amplitude, average amplitude, and amplitude change rate, etc. Record these characteristics as structured data for subsequent use. Define the calculation method of the compression ratio, usually set according to the amplitude and frequency characteristics of the amplitude change characteristics. Set a higher compression ratio for the audio segment with larger amplitude changes, and use a lower compression ratio for the segment with smaller amplitude changes. Determine the reference parameters, such as the initial compression ratio (e.g., 4:1) and the maximum compression ratio (e.g., 8:1). Calculate the dynamic compression ratio according to the amplitude characteristics of the high-frequency band. Design a simple rule that if the amplitude change rate exceeds a certain threshold, increase the compression ratio; if it is insufficient, decrease the compression ratio. Select a suitable audio encoding tool, such as AAC, MP3, or Opus. Select a suitable encoding format according to the dynamic compression ratio to ensure the required compression effect while retaining the sound quality. Determine the encoding parameter settings, such as the bit rate and sampling rate, to adapt to the needs of dynamic compression. Perform dynamic compression encoding on the high-frequency sub-band using the selected encoding tool. Adjust the encoding settings according to the calculated dynamic compression ratio parameters. Record the output information during the encoding process, including the encoded file format, size, and compression ratio, etc. Save the dynamically compressed audio signal as a file (such as MP3 or AAC format) and play it for verification to ensure that the sound quality meets the expectations. Select a suitable lossless compression encoding tool, such as FLAC, ALAC, or WAVPACK. These formats can retain all the information of the audio signal and are suitable for processing the low-frequency band. Determine the encoding parameter settings, such as the sampling rate and number of channels, to ensure the effectiveness of lossless compression. Encode the low-frequency sub-band using the lossless compression tool to ensure that the sound quality is not lost. Record the file information output during the encoding process, including the encoded file format and size. After the encoding is completed, perform a playback verification to ensure that the quality of the audio signal meets the expectations. Select a suitable distributed storage framework, such as Apache Kafka, Amazon S3, or Google Cloud Storage. These frameworks can support large-scale data and provide high availability. Determine the storage structure, including data sharding, redundant backup, and access policies, to improve the storage efficiency. Upload the lossless compressed audio encoding and the dynamic compressed audio encoding to the selected distributed storage framework. Ensure that the metadata of the file, including the file name, size, and encoding method, is recorded during the upload process. Upload the data stream to the storage service by writing an upload script and using the API interface. Verify the integrity of the uploaded file by comparing the original file and the uploaded file to ensure that the data is not damaged. Implement a monitoring mechanism to regularly check the health status and storage capacity of the storage framework.
[0046] In this embodiment, step S4 includes the following steps: Step S41: Identify the audio channels of the real-time multi-channel audio signal and extract each audio channel; Step S42: Calculate the phase difference for each audio channel to generate the inter-channel phase difference parameter; Step S43: Perform unified phase compensation calculation based on the inter-channel phase difference parameter to obtain the phase compensation value for each channel; Step S44: Perform phase alignment processing on the real-time multi-channel audio signal based on the phase compensation value of each channel, so as to obtain the channel phase-aligned audio signal.
[0047] In this embodiment, an audio processing library (such as Librosa or PyAudio) is used to load real-time multi-channel audio signals, ensuring that the sampling rate (such as 44.1 kHz or 48 kHz) and bit depth (such as 16 bits or 24 bits) of the audio signals meet the requirements. For multi-channel signals, ensure that the signal format is stereo (such as two channels) or a higher number of channels (such as four channels or eight channels). Traverse each channel in the audio signal. Usually, each channel of the audio signal is stored in the form of an array or matrix in the data structure, and the index of the channel corresponds to different channels of the multi-channel signal. Extract each audio channel as an independent signal, and record the timestamp and corresponding original data of each channel for subsequent phase difference calculation and processing. Store each extracted channel signal in a list or dictionary for subsequent processing and management, ensuring that the index of the channel and the corresponding audio data can be in one-to-one correspondence. Use the NumPy array format to store the data of each channel for convenient subsequent mathematical operations and processing. Perform frequency-domain analysis on the audio signal of each channel using the fast Fourier transform (FFT). The FFT can convert the time-domain signal into a frequency-domain signal, facilitating the calculation of its phase information. Use an appropriate FFT implementation library (such as NumPy or SciPy), and set an appropriate FFT length (such as 1024 or 2048) to ensure the accuracy of the frequency-domain analysis. Perform the FFT on each audio channel, extract its complex frequency-domain representation, and calculate the phase angle of the complex number (usually using np.Using the angle() function, obtain the frequency-domain phase information of each channel, calculate the phase difference between channels, and by subtracting the phase information of the reference channel (such as the first channel), obtain the phase difference of each channel relative to the reference channel. Store the calculated phase differences in an array to ensure that the phase difference parameters of each channel can correspond one-to-one with their respective channels for subsequent compensation calculations. Select a suitable phase compensation algorithm. Usually, the compensation value is directly taken as the negative phase difference value to facilitate adjusting the phase to be consistent. Determine the frequency-domain range of the compensation calculation (such as from 0 Hz to half of the sampling rate) to ensure that all frequencies can be compensated accordingly. For the phase difference parameters of each channel, calculate their corresponding compensation values, and store the compensation values of each channel in a structured array or list for subsequent phase alignment processing. Check the results of the compensation calculation to ensure that the compensation values are appropriate and do not exceed the valid range, and verify the effectiveness of the calculation by comparing the phase consistency before and after compensation. Select a suitable phase alignment method. Usually, use the inverse Fourier transform (IFFT) of the phase compensation and the complex multiplication of the phase adjustment to ensure that when performing phase alignment, complex signals can be processed and the amplitude information can be retained. For the audio signal of each channel, apply the calculated phase compensation value. By multiplying the frequency-domain signal of each channel by the compensation value, the effect of phase alignment is achieved. Perform the inverse Fourier transform (IFFT) to convert the frequency-domain signal back to the time domain to obtain the phase-aligned audio signal. Store the phase-aligned audio signals of all channels in a new data structure for subsequent processing and use. Select to save it as a multi-channel audio file (such as WAV format) for further analysis and processing.
[0048] In this embodiment, the specific steps of step S5 are as follows: Step S51: Perform in-depth audio semantic parsing on the channel phase-aligned audio signal to generate audio signal semantic features; Step S52: Based on the audio signal semantic features, perform multi-audio segment application scenario recognition to generate the application scenarios of multiple audio segments; Step S53: Analyze the audio scene transmission requirements for the application scenarios of multiple audio segments to obtain the audio transmission requirements for each application scenario; Step S54: Mine the transmission load requirements for the audio transmission requirements of each application scenario to generate the application scenario transmission load requirement features. In this embodiment, an audio processing library (such as Librosa or PyAudio) is used to load the audio signal with aligned channel phases into memory, ensuring that the format and sampling rate of the audio signal (such as 44.1 kHz) meet the standards. The audio signal is normalized to eliminate amplitude differences caused by different recording devices. A suitable audio feature extraction method is selected, such as Mel Frequency Cepstral Coefficients (MFCC), pitch features, time-domain features, and spectral features. An appropriate feature set is selected according to specific application requirements, and extraction parameters are determined, such as window length (e.g., 25 ms), overlap rate (e.g., 50%), etc., to improve the accuracy of feature extraction. The audio processing library (such as Librosa) is used to perform a Short-Time Fourier Transform (STFT) on the audio signal, and features such as MFCC, spectral centroid, and spectral roll-off are extracted. The extracted features are organized into feature vectors for subsequent semantic parsing. The feature vectors should contain descriptive information for each feature to ensure the traceability of subsequent analysis. A deep learning model (such as a Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN)) is combined to train the features for semantic parsing of the audio signal. A labeled audio dataset (such as Freesound or Google AudioSet) is used for model training. After training, the model is used to infer the extracted features to generate semantic features of the audio signal, and specific sounds (such as human voices, music, environmental noise, etc.) contained in the audio signal are identified. According to application requirements, multiple audio scenarios are defined, such as "music playback", "video conferencing", "environmental sound effects", etc. Corresponding audio samples are prepared for each scenario for model training. A scene label dataset is created to associate the semantic features of the audio signal with their corresponding scenarios. A suitable algorithm (such as a Support Vector Machine (SVM) or a deep learning model) is selected for application scenario recognition. The semantic features of the audio signal are used as input, and the scene labels are used as output for model training. The cross-validation method is used to evaluate the accuracy of the model to ensure the robustness of recognition. The ratio of the training set to the test set is set (such as 80% for training and 20% for testing). The model is used to infer the semantic features of each audio segment to identify its corresponding application scenario. The recognition results of each audio segment are recorded and associated with the original data of the audio segment to generate application scenario information for multiple audio segments, facilitating subsequent analysis of transmission requirements and mining of load requirements. The key parameters affecting audio transmission requirements are determined, such as bandwidth, latency, packet loss rate, and packet size. According to the characteristics of different scenarios, thresholds for each parameter are set. For the "video conferencing" scenario, the latency requirement is usually less than 150 ms, while the "music playback" scenario can tolerate higher latency. For each identified application scenario, its audio characteristics (such as sampling rate, number of channels, bit rate, etc.) are analyzed, and the required bandwidth is calculated. For a 16-bit stereo (2-channel) audio signal, bandwidth = sampling rate × bit depth × number of channels / 8. The audio transmission requirements for each application scenario are recorded.Include metrics such as the required bandwidth and latency to determine the key parameters affecting the transmission load, such as the number of concurrent users, the number of audio streams, and the data packet transmission frequency. Set corresponding thresholds according to the actual application situation. Through statistical analysis of historical data, determine the load patterns of each application scenario. For each application scenario, calculate the total transmission load by combining the number of concurrent users and the number of audio streams. For the number of simultaneously played audio streams and the bandwidth required for each stream, the total load = bandwidth × number of concurrent users. Record the transmission load demand characteristics of each application scenario, including metrics such as the maximum load and average load, to ensure that the actual requirements of the scenario can be reflected. Organize the mined transmission load demand characteristics into structured data for subsequent optimization and decision support, and ensure that the load demand of each scenario can correspond to its audio transmission demand.
[0049] In this embodiment, the specific steps of step S6 are as follows: Step S61: Calculate the bandwidth utilization rate of the real-time network monitoring parameters to obtain the current network bandwidth utilization rate; Step S62: Analyze the transmission latency of the real-time network monitoring parameters to generate real-time network transmission latency parameters; Step S63: Evaluate the real-time network status based on the current network bandwidth utilization rate and the real-time network transmission latency parameters to obtain the real-time network status characteristics of the audio chip; Step S64: Make a dynamic audio signal transmission decision on the transmission load demand characteristics of the application scenario based on the real-time network status characteristics of the audio chip to generate a dynamic transmission strategy; Step S65: Perform collaborative management optimization according to the dynamic transmission strategy and the distributed stream storage framework to build an intelligent audio collaborative management engine.
[0050] In this embodiment, a network monitoring tool (such as Wireshark or Nagios) is used to collect network data in real time, including total bandwidth, used bandwidth, and idle bandwidth, ensuring that the data obtained during data collection can reflect the network status in real time. Determine the monitoring period (such as once per second) for subsequent real-time calculation and analysis. The bandwidth utilization rate = (used bandwidth / total bandwidth) × 100%. By reading the network parameters in real time, substitute the used bandwidth and total bandwidth into the formula to calculate the current bandwidth utilization rate. Store the calculated bandwidth utilization rate result in the database and record the timestamp for subsequent trend analysis. Monitor whether the bandwidth utilization rate exceeds the set threshold (such as 80%) to detect network congestion problems in a timely manner. Use the Ping command or more advanced network performance monitoring tools (such as iperf or MTR) to measure network latency. These tools can send data packets to the target address and record the round-trip time (RTT). Determine the monitoring frequency (such as once every 5 seconds) to obtain real-time latency data. For each Ping test, record the round-trip time and calculate the average latency. The average latency = number of tests / total latency. Through multiple tests, ensure that the obtained latency data is representative. Store the real-time network transmission latency parameters in the database and record the timestamp for subsequent analysis. Monitor whether the latency exceeds the set threshold (such as 100 ms) to detect network performance problems in a timely manner. Determine the evaluation metrics, including bandwidth utilization rate and transmission latency, and set reasonable thresholds for status evaluation. A bandwidth utilization rate exceeding 80% is regarded as a warning state, and a latency exceeding 100 ms is regarded as a bad state. Based on the real-time monitoring parameters, comprehensively analyze the bandwidth utilization rate and transmission latency, and use the weighted method or logistic regression model to build a comprehensive evaluation model to obtain the network status score. Store the evaluation result in the database and record the timestamp for subsequent analysis of the historical records of the network status. Judge whether the network status is good according to the score and generate corresponding status characteristics (such as "good", "warning", "bad"). Conduct correlation analysis between the transmission load demand characteristics of the application scenario and the real-time network status characteristics to identify network requirements in different scenarios. The video conferencing scenario is more sensitive to latency, while the music playback scenario has a higher demand for bandwidth. Design a decision model and formulate a dynamic transmission strategy based on the real-time network status and transmission demand characteristics. Use decision trees or reinforcement learning models to automate the decision-making process. Set decision rules that when the bandwidth utilization rate is higher than 80% and the latency is lower than 100 ms, give priority to transmitting high-priority audio streams. Record the generated dynamic transmission strategy in the database and implement corresponding strategy adjustments to ensure the efficient transmission of audio signals. Monitor the network performance after the implementation of the strategy for dynamic adjustment and optimization. Select a suitable distributed stream storage framework (such as Apache Kafka or RabbitMQ) to support the efficient transmission and storage of audio data. These frameworks can handle high-concurrent audio streams and ensure the reliability and real-time nature of the data.Integrate the dynamic transmission strategy with the distributed stream storage framework to ensure that audio signals can be transmitted and stored through the optimal path. Set the load balancing strategy to ensure that in high-load situations, each audio stream can evenly allocate storage resources. Develop an intelligent audio collaborative management engine that continuously optimizes the transmission strategy and storage management using real-time data analysis and machine learning techniques. The engine should be able to analyze the network status and application requirements in real time, automatically adjust the transmission strategy, monitor the performance of the engine, and ensure that it can adaptively manage in different scenarios to improve the overall audio transmission efficiency.
[0051] In this embodiment, an audio signal data management system for an audio chip is provided, which is used to execute the audio signal data management method of the audio chip as described above, and includes: An audio enhancement module, which is used to identify the real-time multi-channel audio signals and real-time network monitoring parameters of the audio chip; perform adaptive filtering and noise reduction and detail enhancement processing on the real-time multi-channel audio signals to obtain detail-enhanced real-time audio signals; A dynamic audio feature module, which is used to define analysis windows for each frequency band of the detail-enhanced real-time audio signals, and perform dynamic spectrum feature mining to generate dynamic audio signal features for each frequency band; A compression coding module, which is used to perform dynamic compression coding and lossless compression audio coding on the detail-enhanced real-time audio signals based on the dynamic audio signal features of each frequency band, and construct a distributed stream storage framework; A phase alignment module, which is used to identify audio channels of the real-time multi-channel audio signals and perform phase alignment processing to obtain channel-phase-aligned audio signals; A load demand module, which is used to perform in-depth audio semantic parsing on the channel-phase-aligned audio signals and perform transmission load demand mining to generate application scenario transmission load demand features; A collaborative management module, which is used to make dynamic audio signal transmission decisions on the application scenario transmission load demand features according to the application scenario transmission load demand features, and perform collaborative management optimization based on the distributed stream storage framework to construct an intelligent audio collaborative management engine.
[0052] The present invention removes interfering noises in real time through adaptive filtering to ensure that the audio signal remains clear in complex environments. Detail enhancement can improve the dynamic range and spatial sense of the audio, enhancing the user's auditory experience. By combining noise reduction and detail enhancement, the impact of environmental noise on the audio signal is reduced, thereby improving the accuracy of speech recognition systems or voice communication systems. Whether in a noisy environment or an environment with high-quality audio requirements, the audio enhancement module can effectively optimize the signal quality. Through per-band analysis, the system extracts fine-grained features in the audio signal, enhancing the understanding and processing ability of the audio signal. This helps to more precisely process different types of audio signals (such as speech, music, environmental sounds). The dynamic spectrum feature mining provides rich feature data support for subsequent processing such as compression coding, noise reduction, and enhancement, and can accurately adapt to different audio requirements. The combination of dynamic compression coding and lossless compression enables efficient transmission of audio data in different frequency bands under limited network bandwidth, reducing bandwidth pressure. For application scenarios with high requirements for sound quality (such as music or high-fidelity audio), lossless compression ensures that the sound quality will not degrade due to compression, meeting the requirements of high-quality audio transmission. Phase alignment processing ensures that there are no phase differences and interferences in the multi-channel audio signal during audio output, making the audio output clearer and more stereoscopic, suitable for stereo or surround sound systems. Performing phase alignment processing on multi-channel signals helps to reduce phase distortion and echo phenomena in the audio, improving the auditory experience. In terms of the spatial positioning and directionality of audio signals, phase alignment helps to improve the positioning accuracy of audio, enabling users to obtain a better sense of immersion in applications such as virtual reality and audio positioning. By deeply analyzing the semantic features of audio signals, the system can identify different requirements for audio transmission in different scenarios. Voice communication has high requirements for latency, while music streaming is more concerned with the fidelity of sound quality. By mining the load requirements of application scenarios, the system can accurately predict the required resources, avoiding bandwidth overload or resource waste, thereby improving the transmission efficiency and quality. The dynamic transmission decision automatically adjusts the transmission strategy of the audio signal according to real-time network conditions, application scenarios, and system load, reducing manual intervention and improving processing efficiency. Through collaborative management and optimization, the system can dynamically balance the resource requirements of different application scenarios, avoid system overload, and at the same time ensure the timely and stable transmission of audio signals. Through distributed storage and collaborative management optimization, the system better supports the transmission and management of audio signals across devices and network environments, improving the integration ability of multi-terminal systems.
[0053] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to encompass all changes that fall within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0054] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A method for managing audio signal data of an audio chip, characterized in that: The following steps are involved: Step S1: Identify the real-time multi-channel audio signal and real-time network monitoring parameters of the audio chip; Performing adaptive filtering, noise reduction and detail enhancement processing on the real-time multi-channel audio signal to obtain a detail-enhanced real-time audio signal; Step S2: defining a frequency band-by-frequency band analysis window for the detail-enhanced real-time audio signal, and mining dynamic spectrum features to generate dynamic audio signal features for each frequency band; Step S3: Based on the dynamic audio signal characteristics of each frequency band, the detail-enhanced real-time audio signal is dynamically compressed and encoded and lossless compressed audio is encoded to construct a distributed stream storage framework; Step S4: performing audio channel identification on the real-time multi-channel audio signal and performing phase alignment processing to obtain a channel phase-aligned audio signal; Step S5: performing deep audio semantic analysis on the channel phase aligned audio signal and mining the transmission load demand, thereby generating transmission load demand characteristics of the application scenario; Step S6: Make dynamic audio signal transmission decisions based on the transmission load demand characteristics of the application scenario, and perform collaborative management optimization based on the distributed stream storage framework to build an intelligent audio collaborative management engine.
2. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S1 are: Step S11: Identify the real-time multi-channel audio signal and real-time network monitoring parameters of the audio chip; Step S12: performing deep audio feature analysis on the real-time multi-channel audio signal to extract real-time audio deep features; Step S13: performing refined signal classification based on real-time audio depth features to obtain audio background noise data; Step S14: Adaptively filter and reduce noise on the audio background noise data, thereby generating a noise reduction optimized audio signal; Step S15: performing detail enhancement processing on the noise reduction optimized audio signal to obtain a detail enhanced real-time audio signal.
3. The audio signal data management method of the audio chip according to claim 2, characterized in that: The specific steps of step S14 are: Divide the audio background noise data into short time series segments to obtain multiple segments of short time series background noise data; Calculate the power spectrum density of multiple short-time series background noise data segment by segment to generate the power spectrum density of each segment; Perform amplitude envelope analysis on multiple short-time series background noise data to extract the amplitude envelope characteristics of each segment; The adaptive filtering parameters are calculated based on the amplitude envelope characteristics of each section and the power spectrum density of each section, so as to obtain the adaptive filtering parameters of each section of noise; The audio background noise data is filtered and suppressed in real time according to the adaptive filtering parameters of each noise segment, thereby generating a noise reduction optimized audio signal.
4. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S2 are: Step S21: performing multi-time-frequency decomposition on the detail-enhanced real-time audio signal, thereby generating audio signals of multiple frequency bands; Step S22: performing transient frequency change recognition on audio signals of multiple frequency bands to identify transient frequency change characteristics of the audio signals; Step S23: performing signal stability analysis according to the transient frequency change characteristics of the audio signal, thereby obtaining the stability characteristics of the audio signal; Step S24: defining the analysis window for each frequency band based on the stationary characteristics of the audio signal, and generating the analysis window length for each frequency band; Step S25: mining dynamic spectrum features of audio signals of multiple frequency bands based on the analysis window length of each frequency band, and generating dynamic audio signal features of each frequency band.
5. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S3 are: Step S31: performing audio bandwidth classification on the detail-enhanced real-time audio signal based on the dynamic audio signal characteristics of each frequency band, and generating a plurality of audio spectrum sub-bands, wherein the plurality of audio spectrum sub-bands include high frequency band sub-bands and low frequency band sub-bands; Step S32: Calculate the frequency amplitude change of the high frequency sub-band to generate the amplitude change characteristics of the high frequency sub-band; Step S33: dynamically adjusting the encoding bit rate and compression ratio based on the amplitude variation characteristics of the high frequency sub-band, thereby generating a dynamic compression ratio parameter; Step S34: dynamically compressing and encoding the high frequency sub-band based on the dynamic compression ratio parameter to obtain a dynamically compressed audio code; Step S35: performing lossless compression encoding on the low frequency sub-band, thereby generating lossless compression audio encoding; Step S36: Perform distributed coding stream storage on the lossless compression audio coding and the dynamic compression audio coding, and build a distributed stream storage framework.
6. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S4 are: Step S41: performing audio channel identification on the real-time multi-channel audio signal to extract each audio channel; Step S42: Calculate the phase difference of each audio channel to generate a phase difference parameter between channels; Step S43: performing unified phase compensation calculation according to the phase difference parameter between channels to obtain a phase compensation value for each channel; Step S44: performing phase alignment processing on the real-time multi-channel audio signal based on the phase compensation value of each channel, so as to obtain a channel phase-aligned audio signal.
7. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S5 are: Step S51: performing deep audio semantic analysis on the channel phase aligned audio signal to generate audio signal semantic features; Step S52: performing application scenario recognition of multiple audio segments based on the semantic features of the audio signal, thereby generating application scenarios of multiple audio segments; Step S53: performing audio scene transmission requirement analysis on application scenarios of multiple audio segments to obtain audio transmission requirements of each application scenario; Step S54: mining the transmission load requirements of the audio transmission requirements of each application scenario, thereby generating application scenario transmission load requirement characteristics.
8. The audio signal data management method of the audio chip according to claim 1, characterized in that: The specific steps of step S6 are: Step S61: Calculating the bandwidth utilization of the real-time network monitoring parameters to obtain the current network bandwidth utilization; Step S62: performing transmission delay analysis on the real-time network monitoring parameters to generate real-time network transmission delay parameters; Step S63: performing real-time network status evaluation on the current network bandwidth utilization and real-time network transmission delay parameters to obtain real-time network status characteristics of the audio chip; Step S64: making a dynamic audio signal transmission decision based on the real-time network status characteristics of the audio chip and the transmission load demand characteristics of the application scenario to generate a dynamic transmission strategy; Step S65: Optimize collaborative management based on dynamic transmission strategies and distributed stream storage framework, and build an intelligent audio collaborative management engine.
9. An audio signal data management system for an audio chip, characterized in that: The method for managing audio signal data of an audio chip according to claim 1 comprises: The audio enhancement module is used to identify the real-time multi-channel audio signal of the audio chip and the real-time network monitoring parameters; and to perform adaptive filtering, noise reduction and detail enhancement processing on the real-time multi-channel audio signal to obtain a detail-enhanced real-time audio signal; Dynamic audio feature module, which is used to define the frequency band analysis window for detail-enhanced real-time audio signals, and to mine dynamic spectrum features to generate dynamic audio signal features for each frequency band; The compression coding module is used to perform dynamic compression coding and lossless compression audio coding on detail-enhanced real-time audio signals based on the dynamic audio signal characteristics of each frequency band, and to build a distributed stream storage framework; A phase alignment module, used to perform audio channel identification on the real-time multi-channel audio signal and perform phase alignment processing to obtain a channel phase-aligned audio signal; The load demand module is used to perform deep audio semantic analysis on the channel phase-aligned audio signal and to mine the transmission load demand, thereby generating the transmission load demand characteristics of the application scenario; The collaborative management module is used to make dynamic audio signal transmission decisions based on the transmission load demand characteristics of the application scenario, and to perform collaborative management optimization based on the distributed stream storage framework to build an intelligent audio collaborative management engine.
Citation Information
Cited By
Satellite communication voice adaptive adjustment method based on environmental perception
CN121664270A