Digital sound data acquisition method and system
By acquiring and fusing acoustic data from multiple auxiliary microphones, environmental reflections and external noise are identified and separated, solving the sound calibration problem of digital audio systems in complex public places and improving the immersive listening experience and adaptability.
Patent Information
- Application Number
- CN202511458154.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-12-30
AI Technical Summary
Existing digital audio systems cannot accurately separate environmental reflections and external noise in complex public spaces, leading to sound calibration errors and reducing the immersive listening experience and adaptability.
By acquiring acoustic change information, location information, and timestamps from multiple auxiliary microphones, the acoustic profile is fused and updated to identify abnormal acoustic components, obtain high-precision acquired sound data, and use the predicted sound data and acquired sound data to identify reflection components and external environmental noise information.
It achieves accurate description and dynamic updates of real soundscapes, enhancing the immersive listening experience and adaptability of digital audio systems in complex public spaces.
Smart Images

Figure CN121240004A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data acquisition technology, specifically to a digital audio data acquisition method and system. Background Technology
[0002] Modern digital sound systems are installed in public spaces such as high-end retail spaces or interactive museum exhibition halls to provide an immersive auditory experience. The operating environment of public spaces is highly variable; factors such as merchandise displays, temporary structures, customer density, exhibit changes, and visitor numbers frequently alter the acoustic characteristics of the space. Traditional static calibration methods are inadequate to adapt to these continuous changes and cannot be repeatedly used for calibrating against interfering test tones in public environments.
[0003] By deploying a network of auxiliary microphones within public spaces, the data collected by the auxiliary microphones while the digital sound system continuously plays background audio constitutes a complex, mixed sound field. This field includes both the sound output after the system interacts with its environment and actual external environmental sounds, such as conversations, footsteps, exhibit sounds, and crowd noise. However, currently, the sound output after environmental interactions is easily misinterpreted as environmental changes, making it difficult to separate environmental reflections and external noise. This leads to incorrect sound adjustments and further reduces the immersive experience. Furthermore, different brands and types of auxiliary microphones exhibit inconsistencies in frequency response curves, sensitivity levels, signal-to-noise ratios, and dynamic ranges. Integrating the raw data from multiple auxiliary microphones introduces significant errors, resulting in inaccurate descriptions of the real soundscape. It becomes difficult to accurately separate environmental reflections and external noise from the mixed sound field, hindering precise sound calibration of the digital sound system and reducing its immersive listening experience and adaptability in complex public spaces. Summary of the Invention
[0004] The purpose of this invention is to provide a digital audio data acquisition method and system to solve the problem that existing technologies are unable to accurately separate environmental reflections and external noise from a mixed sound field, and thus cannot accurately calibrate the sound of digital audio.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a digital audio data acquisition method, comprising the following steps: Acquire acoustic change information, position information, and timestamps of multiple auxiliary microphones during digital audio playback; All acoustic change information, location information, and timestamps are fused together to obtain fused acoustic information; The preset acoustic profile is updated using fused acoustic information to obtain an updated acoustic profile; After identifying the abnormal acoustic component information in the acoustic update profile, the collected sound data is obtained by the auxiliary microphone corresponding to the abnormal acoustic component information. Based on the collected sound data and the updated acoustic profile, the predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information is determined. Based on predicted sound data and collected sound data, the collected results are used to identify information on the reflection components of audio played by digital speakers and information on external environmental noise.
[0006] Furthermore, the step of acquiring acoustic change information from multiple auxiliary microphones during digital audio playback includes: Acquire the initial sound data collected by each auxiliary microphone over a period of time and the raw audio information of the digital audio playback during digital audio playback; The initial sound data is processed to obtain frequency distribution information and root mean square energy information; Based on frequency distribution information, root mean square energy information, original audio information, and preset audio rules, acoustic change information containing acoustic event summary information or state characteristic information is determined.
[0007] Furthermore, after identifying anomalous acoustic component information in the acoustic update profile, the step of obtaining the acquired sound data from the auxiliary microphone corresponding to the anomalous acoustic component information includes: After identifying the abnormal acoustic component information in the acoustic update profile, a data acquisition command with detection mode parameters is sent to the auxiliary microphone corresponding to the abnormal acoustic component information; Based on the data acquisition command, acquire the sound data, which includes high-resolution data information and high-precision timestamp information, collected by the auxiliary microphone corresponding to the abnormal acoustic component information.
[0008] Furthermore, the step of determining the predicted sound data collected by the auxiliary microphone corresponding to the anomalous acoustic component information based on the collected sound data and the acoustic updated profile includes: The fine-tuned response data acquired by the auxiliary microphone in response to the attached detection mode parameters is used as the acquired sound data; Based on the collected sound data and the updated acoustic profile, the predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information is determined.
[0009] Furthermore, the step of identifying the acquisition results, including the reflection component information of the audio played by the digital speaker and the external environmental noise information, based on the predicted sound data and the acquired sound data includes: After determining the residual signal between the predicted sound data and the acquired sound data, it is confirmed that the residual signal has a quasi-periodic component information with a time synchronization causal relationship when fine-tuning with the original audio information; After removing the quasi-periodic component information from the residual signal, time-frequency analysis is performed on the removed residual signal to confirm the energy distribution of the removed residual signal within multiple preset time windows and frequency sub-bands. Confirm the cross-correlation coefficient between the energy distribution of the stripped residual signal and the energy distribution of the original audio information; The time delay component whose cross-correlation coefficient exceeds the first preset threshold is taken as the collection result of the reflection component information; The spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the stripped residual signal corresponding to cross-correlation coefficients below the second preset threshold are matched with preset environmental noise feature templates to identify the acquisition results of external environmental noise.
[0010] Furthermore, in the step of using the time delay component with a cross-correlation coefficient exceeding a first preset threshold as the acquisition result of the reflection component information, the step of constructing the first preset threshold includes: Based on the acoustic update profile, confirm the reverberation time, spatial geometry changes, or reflective surface material information recorded in the environmental acoustic characteristics description; The preset base threshold is adjusted by using the recorded reverberation time, spatial geometric changes, or reflective surface material information to obtain the first preset threshold.
[0011] Furthermore, the acquisition result of using time delay components with cross-correlation coefficients exceeding a first preset threshold as reflection component information includes: Spatial consistency verification is performed on the cross-correlation coefficients to obtain coefficient verification information; The cross-correlation coefficient after verification is greater than the preset verification threshold; The time delay components whose cross-correlation coefficients exceed the first preset threshold after testing are taken as the results of the collection of reflection component information.
[0012] Furthermore, the step of performing spatial consistency verification on the cross-correlation coefficients to obtain coefficient verification information includes: The cross-correlation coefficients were subjected to data quality assessment, and the quality assessment results were obtained. When the quality assessment results are abnormal, the cross-correlation coefficient is corrected to obtain the corrected cross-correlation coefficient. Spatial consistency checks are performed on the corrected cross-correlation coefficients to obtain coefficient verification information.
[0013] Furthermore, the steps for identifying the acquisition results of external environmental noise by matching the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the stripped residual signal corresponding to cross-correlation coefficients below a second preset threshold with preset environmental noise feature templates include: The residual signals after stripping, corresponding to those with cross-correlation coefficients lower than the second preset threshold, are spatially weighted and averaged to obtain the weighted average residual signals. Extract the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope rate of change from the weighted average residual signal; The spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate in the weighted average residual signal are matched with a preset environmental noise feature template to identify the acquisition results of external environmental noise.
[0014] The present invention also provides a digital audio data acquisition system, the system comprising: The first acquisition module is used to acquire acoustic change information of multiple auxiliary microphones, position information of multiple auxiliary microphones and timestamps during digital audio playback; The fusion module is used to fuse all acoustic change information, location information, and timestamps to obtain fused acoustic information. The update module is used to update the preset acoustic profile using fused acoustic information to obtain an updated acoustic profile. The second acquisition module is used to identify abnormal acoustic component information in the acoustic update profile and then acquire the acquired sound data obtained by the auxiliary microphone corresponding to the abnormal acoustic component information. The determination module is used to determine the predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information based on the collected sound data and the acoustic updated profile. The identification and acquisition module is used to identify the acquisition results, including the reflection component information of the audio played by the digital speaker and the external environmental noise information, based on the predicted sound data and the acquired sound data.
[0015] Compared with the prior art, the digital audio data acquisition method and system of the present invention have the following advantages: This invention acquires acoustic change information, location information, and timestamps from multiple auxiliary microphones during digital audio playback, and performs fusion processing to update a preset acoustic profile, enabling dynamic perception and description of environmental acoustic characteristics. When abnormal acoustic components are identified, high-precision sound data can be acquired specifically, and based on the acquired sound data and the updated acoustic profile, the corresponding sound data can be predicted. By comparing the predicted sound data and the acquired sound data, the reflection components and external environmental noise information of the audio played by the digital audio system can be accurately identified. Through non-invasive and continuous data acquisition and refined analysis, the challenges posed by differences in auxiliary microphone hardware characteristics and uncertainties in data transmission are overcome, thereby achieving accurate description and dynamic updating of the real sound scene. This facilitates subsequent precise sound calibration of the digital audio system, significantly improving the immersive listening experience and adaptability of the digital audio system in complex public places. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the specific embodiments will be briefly described below. In all the drawings, the elements or parts are not necessarily drawn to scale.
[0017] Figure 1 This is a flowchart of a digital audio data acquisition method according to the present invention.
[0018] Figure 2 This is a structural block diagram of a digital audio data acquisition system according to the present invention.
[0019] In the diagram: 210, First Acquisition Module; 220, Fusion Module; 230, Update Module; 240, Second Acquisition Module; 250, Determination Module; 260, Identification and Acquisition Module.
[0020] The implementation and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] The following drawings disclose several embodiments of the present invention. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential. Furthermore, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.
[0022] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0023] Furthermore, in this invention, the use of terms such as "first" and "second" is for descriptive purposes only and does not specifically refer to any order or sequence, nor is it intended to limit the invention. They are merely used to distinguish components or operations described using the same technical terms, and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but only if they are feasible for those skilled in the art. If a combination of technical solutions is contradictory or impossible to implement, such a combination should be considered nonexistent and not within the scope of protection claimed by this invention.
[0024] Existing digital audio systems typically employ static calibration methods when calibrating acoustic environments in public spaces. This approach struggles to adapt to changes in acoustic characteristics caused by factors such as merchandise displays, temporary structures, and customer density. Furthermore, different auxiliary microphones are affected by uncertainties in hardware characteristics, placement, and data transmission, making it impossible to accurately separate environmental reflections and external noise from the mixed sound field. If these issues are not addressed, the system may misinterpret its own output sound as environmental changes, leading to incorrect acoustic environment adjustments and further diminishing the user's immersive experience.
[0025] To further understand the content, features, and effects of this invention, the following embodiments are provided, and detailed descriptions are given below in conjunction with the accompanying drawings: Please see Figure 1 This invention provides a digital audio data acquisition method, comprising the following steps: S100: Acquire acoustic change information, position information, and timestamps of multiple auxiliary microphones during digital audio playback. Here, auxiliary microphones refer to microphones used to collect sound data in the digital audio system environment, which may have different hardware characteristics and placement locations. Acoustic change information refers to parameters reflecting changes in environmental acoustic characteristics, such as frequency distribution and root-mean-square energy. Specifically, the steps of acquiring acoustic change information, position information, and timestamps of multiple auxiliary microphones during digital audio playback can be implemented in various ways. For example, each auxiliary microphone can periodically collect sound data from its surroundings and upload it to the central processing unit. Simultaneously, the installation location information of the auxiliary microphones can be pre-stored in the system, with a corresponding timestamp attached each time data is uploaded. Alternatively, the auxiliary microphones can be configured to only begin collecting sound data and recording their position and precise timestamp when a digital audio playback event is detected.
[0026] S200. All acoustic change information, location information, and timestamps are fused to obtain fused acoustic information. Fusion acoustic information refers to the information obtained by comprehensively processing acoustic change information, location information, and timestamps from multiple auxiliary microphones, used to more comprehensively describe the environmental acoustic state. Specifically, in the step of fusing all acoustic change information, location information, and timestamps to obtain fused acoustic information, various data fusion techniques can be used, such as weighted averaging, Kalman filtering, or deep learning. For example, acoustic change information from different auxiliary microphones can be weighted and averaged, where the weights can be allocated according to the microphone's signal-to-noise ratio, distance from the digital audio system, or historical data quality. Simultaneously, location information and timestamps can be used to construct a spatiotemporal model of acoustic events to ensure data synchronization and spatial consistency.
[0027] S300. Update the preset acoustic profile using fused acoustic information to obtain an updated acoustic profile. The preset acoustic profile describes the initial or historical acoustic characteristics of the target acoustic environment, including information such as reverberation time, spatial geometry, and reflective surface materials. Specifically, in the step of updating the preset acoustic profile using fused acoustic information, adaptive filtering, machine learning models, or rule-based inference systems can be employed. For example, the fused acoustic information can be input into a pre-trained neural network model, which can adjust parameters in the preset acoustic profile, such as reverberation time, absorption coefficient, or scattering coefficient, based on new acoustic data. Alternatively, the acoustic profile can be gradually corrected using preset update rules based on the differences between the fused acoustic information and the preset acoustic profile.
[0028] S400: After identifying abnormal acoustic component information in the acoustic update profile, acquire the acquired sound data obtained by the auxiliary microphone corresponding to the abnormal acoustic component information. Abnormal acoustic component information refers to the portion of the acoustic update profile that significantly differs from the expected or normal acoustic characteristics, indicating new events or changes in the environment. Specifically, in the step of acquiring the acquired sound data obtained by the auxiliary microphone corresponding to the abnormal acoustic component information after identifying abnormal acoustic component information in the acoustic update profile, threshold comparison, pattern recognition, or anomaly detection algorithms can be used. For example, various parameters in the acoustic update profile can be continuously monitored, and when a parameter (such as local reverberation time or energy in a specific frequency range) exceeds a preset normal fluctuation range, it is identified as abnormal acoustic component information. Once an anomaly is identified, a command is sent to the auxiliary microphone closest to or most relevant to the anomaly area, requesting it to perform more refined sound acquisition to obtain high-resolution acquired sound data.
[0029] S500. Based on the acquired sound data and the updated acoustic profile, determine the predicted sound data collected by the auxiliary microphone corresponding to the anomalous acoustic component information. Specifically, in the step of determining the predicted sound data collected by the auxiliary microphone corresponding to the anomalous acoustic component information based on the acquired sound data and the updated acoustic profile, acoustic modeling, inverse filtering, or wave field synthesis techniques can be employed. For example, the environmental geometry and material properties described in the updated acoustic profile can be used, combined with the original audio information from the digital audio system, to predict the sound data that should be received at the location of the auxiliary microphone corresponding to the anomalous acoustic component information through an acoustic simulation model. The predicted sound data will serve as a benchmark for subsequent comparison with the actual acquired sound data.
[0030] S600: Based on the predicted sound data and the acquired sound data, identify the acquired results including the reflection component information of the digital audio playback and the external environmental noise information. The reflection component information refers to the sound component of the digital audio playback that reaches the microphone after reflection in the environment. The external environmental noise information refers to all other sounds besides the digital audio playback and its reflection components, such as human voices and footsteps. Specifically, in the step of identifying the acquired results including the reflection component information of the digital audio playback and the external environmental noise information based on the predicted sound data and the acquired sound data, signal differential analysis, correlation analysis, or blind source separation techniques can be used. For example, by subtracting the predicted sound data from the acquired sound data, a residual signal can be obtained, which mainly contains the reflection component information and the external environmental noise information. Subsequently, the residual signal can be further analyzed, for example, by using time delay estimation and spectral analysis to distinguish the reflection component information and the external environmental noise information.
[0031] In this embodiment, by fusing acoustic change information, location information, and timestamps from multiple auxiliary microphones, an acoustic profile is constructed and continuously updated, thereby achieving real-time perception and dynamic adaptation to environmental acoustic characteristics. Compared with existing technologies, this application can non-invasively identify abnormal acoustic components in the acoustic profile and acquire high-resolution sound data in a targeted manner. Simultaneously, by determining predicted sound data based on the acquired sound data and the acoustic profile, and comparing it with the actual acquired sound data, this application can accurately identify the reflection components and external environmental noise information of the audio played by the digital speaker. This overcomes the challenges posed by differences in auxiliary microphone hardware characteristics and uncertainties in data transmission, thus achieving an accurate description and dynamic update of the real soundscape. This facilitates subsequent precise sound calibration of the digital speaker, significantly improving the immersive listening experience and adaptability of the digital speaker system in complex public spaces.
[0032] In some embodiments described above in this application, the present invention also provides a step for acquiring acoustic change information of multiple auxiliary microphones during digital audio playback, comprising: This process acquires initial sound data collected by each auxiliary microphone over a period of time during digital audio playback, as well as the raw audio information played by the digital audio system. Initial sound data refers to the raw acoustic signals directly collected by the auxiliary microphones during digital audio playback, encompassing all sound components in the environment. Raw audio information refers to the audio content currently being played by the digital audio system, serving as a baseline for subsequent acoustic analysis.
[0033] The initial sound data is processed to obtain frequency distribution information and root-mean-square (RMS) energy information. Specifically, various signal processing techniques, such as Fourier transform or wavelet analysis, can be used to process the initial sound data to extract its frequency distribution information. Simultaneously, the RMS energy information, reflecting the loudness or energy level of the sound, can be obtained by calculating the short-time energy or RMS value.
[0034] Based on frequency distribution information, root mean square energy information, raw audio information, and preset audio rules, acoustic change information, including acoustic event summary information or state feature information, is determined. The preset audio rules can be understood as a set of predefined logic or models used to analyze and judge acoustic events or states. For example, they may include energy thresholds within a specific frequency range, matching templates for specific sound patterns, or classifiers trained based on machine learning. These rules guide how to identify meaningful acoustic changes from the processed audio data. Furthermore, the acoustic change information can include acoustic event summary information, such as identifying human voices, percussion sounds, and environmental noise types, or state feature information, such as the reverberation level of the environment and the stability of background noise, thus enabling a concise and effective description of the acoustic dynamics in a digital audio playback environment.
[0035] In this embodiment, initial sound data acquired by an auxiliary microphone and raw audio information played from digital speakers provide the foundational data for subsequent acoustic analysis. Subsequently, by extracting the frequency distribution and root-mean-square energy from the initial sound data, key spectral and energy features can be separated from the original complex acoustic signal. These features, combined with the raw audio information and preset audio rules, can be used to specifically identify and quantify acoustic events or state changes in the environment, thereby generating acoustic change information with higher precision and richer detail. Through step-by-step processing and multi-information fusion, a comprehensive and accurate perception of changes in the acoustic environment is ensured.
[0036] In some embodiments described above in this application, the present invention also provides a step of obtaining the acquired sound data obtained by using an auxiliary microphone corresponding to the abnormal acoustic component information after identifying abnormal acoustic component information in the acoustic update profile, including: After identifying anomalous acoustic components in the updated acoustic profile, a data acquisition command with detection mode parameters is sent to the auxiliary microphone corresponding to the anomalous acoustic components. Specifically, the detection mode parameters can be understood as configuration information used to guide the auxiliary microphone to perform data acquisition in a specific mode. For example, these parameters may include, but are not limited to, sampling rate, bit depth, gain settings, specific frequency response curves, acquisition duration, or triggering conditions. This ensures that the auxiliary microphone can acquire data in an optimized manner for the detected anomalous acoustic components, thereby capturing richer and more accurate acoustic details.
[0037] Based on data acquisition commands, the system acquires sound data, including high-resolution data and high-precision timestamp information, from auxiliary microphones corresponding to abnormal acoustic components. High-resolution data refers to sound data with a sampling rate and bit depth exceeding conventional acquisition standards, capable of reflecting sound wave amplitude and frequency variations more precisely. For example, it can be acquired using 24-bit or 32-bit depth and a sampling rate of 96kHz or 192kHz. High-precision timestamp information refers to time stamps with extremely high time accuracy, recorded synchronously with the acquired sound data. Its accuracy can reach microseconds or even nanoseconds, ensuring time synchronization between different auxiliary microphones and between the microphone and the original audio played back by the digital audio system. This facilitates subsequent time delay analysis and spatial localization of acoustic events.
[0038] Specifically, during digital audio playback, the acoustic profile update identifies an anomalous acoustic component in a certain area with an extremely short duration and high frequency. This could be a faint reflection or a transient high-frequency noise. To accurately analyze this anomalous component, a data acquisition command is sent to the auxiliary microphone corresponding to that area. This command may include parameters such as increasing the sampling rate to 192kHz, setting the bit depth to 24 bits, enabling a gain mode for the high-frequency band, and requiring the acquired data to include a microsecond-level timestamp. Based on this command, the auxiliary microphone will acquire sound data in a high-specification mode, obtaining sound data containing rich high-frequency details and precise time synchronization information. For example, in this way, the precise start time, duration, and spectral characteristics of the high-frequency anomalous acoustic component can be clearly captured. This allows for accurate time-delay comparison with the original audio from the digital audio system in subsequent analysis to identify whether it is a reflection from a specific surface, or to match it with a preset noise template to determine whether it is a specific type of external environmental noise, such as high-frequency electronic interference. Targeted, high-precision data acquisition greatly improves the ability to identify complex acoustic events.
[0039] In this embodiment, after identifying abnormal acoustic components, instead of directly performing conventional data acquisition, a data acquisition command with detection mode parameters is first sent to the corresponding auxiliary microphone. This allows the auxiliary microphone to adjust its acquisition behavior based on these parameters, such as increasing the sampling rate, increasing the bit depth, or focusing on a specific frequency band. This targeted, high-specification data acquisition ensures that the acquired sound data contains high-resolution data and high-precision timestamp information. The high-resolution data captures subtle features of abnormal acoustic components, such as weak reflections or noise at specific frequencies, while the high-precision timestamp information ensures accurate alignment of these data in the time dimension, thus providing a solid foundation for subsequent accurate calculation of sound wave propagation delay, identification of reflection paths, and differentiation of transient noise.
[0040] In some embodiments described above in this application, the present invention also provides the step of determining the predicted sound data collected by the auxiliary microphone corresponding to the anomalous acoustic component information based on the collected sound data and the acoustic updated profile, including: The fine-tuned response data acquired by the auxiliary microphone in response to the accompanying detection mode parameters is used as the acquired sound data. Specifically, the fine-tuned response data refers to the sound data acquired by the auxiliary microphone under the guidance of specific detection mode parameters, which has been finely adjusted and optimized. The detection mode parameters are a set of preset configurations or instructions used to guide the auxiliary microphone to adopt specific operating modes when acquiring sound data, such as adjusting the sampling rate, gain, filtering characteristics, beamforming direction, or dynamic range, to more effectively capture information about anomalous acoustic components. By introducing detection mode parameters, the acquisition behavior of the auxiliary microphone can be made more targeted, thereby obtaining higher-quality sound data that is more correlated with anomalous acoustic components.
[0041] Based on the collected sound data and the updated acoustic profile, the predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information is determined.
[0042] In this embodiment, before determining the predicted sound data, the fine-tuned response data acquired by the auxiliary microphone corresponding to the attached detection mode parameters is used as the acquired sound data. Due to the introduction of the detection mode parameters, the auxiliary microphone can adjust its acquisition strategy based on the characteristics of the anomalous acoustic components identified in the acoustic update profile. For example, high-resolution acquisition can be performed on specific frequency ranges or time windows of the anomalous acoustic components, or beamforming technology can be used to focus on the direction of the anomalous sound source to enhance the target signal and suppress background noise. This targeted data acquisition method allows the obtained fine-tuned response data to more accurately and precisely reflect the true characteristics of the anomalous acoustic components, while effectively suppressing irrelevant noise and interference, thus providing high-quality input for subsequent determination of the predicted sound data. This significantly improves the accuracy and robustness of determining the predicted sound data based on the acquired sound data and the acoustic update profile. Compared to directly using unoptimized acquired sound data, it can more accurately predict the sound data corresponding to anomalous acoustic components, thus laying a solid foundation for subsequent identification of the reflection components and external environmental noise information of digital audio playback, effectively improving the overall performance and recognition accuracy of the entire digital audio data acquisition method.
[0043] In some embodiments described above in this application, the present invention also provides the step of identifying the acquisition results, including reflection component information of digital audio playback and external environmental noise information, based on predicted sound data and acquired sound data, comprising: After determining the residual signal between the predicted sound data and the acquired sound data, it is confirmed that the residual signal has a quasi-periodic component information with a time-synchronous causal relationship with the original audio information during fine-tuning. The residual signal can be understood as the difference between the actually acquired sound data and the sound data predicted based on the acoustic profile update, including reflection components of the audio played by the digital speaker, external environmental noise information, and other acoustic events not covered by the prediction model. Quasi-periodic component information refers to a signal that has a certain time delay from the original audio information, but whose waveform or spectral characteristics are highly similar to the original audio information. It is typically the signal of the audio played by the digital speaker reaching the auxiliary microphone after reflection in the environment. Through fine-grained time synchronization and correlation analysis of the residual signal, reflection components can be effectively identified.
[0044] After removing the quasi-periodic component information from the residual signal, time-frequency analysis is performed on the removed residual signal to confirm its energy distribution within multiple preset time windows and frequency sub-bands. Specifically, the purpose of removing the quasi-periodic component is to separate the reflected signal from the residual signal, so that the remaining signal mainly contains external environmental noise. Time-frequency analysis, such as using methods like Short-Time Fourier Transform (STFT), can reveal the energy distribution characteristics of the signal at different times and frequencies, facilitating subsequent identification of different types of noise. The division of preset time windows and frequency sub-bands allows for refined analysis of the signal, capturing the specific time-frequency characteristics that different noise sources may possess.
[0045] The cross-correlation coefficient between the energy distribution of the stripped residual signal and the energy distribution of the original audio information is determined. This cross-correlation coefficient is used to further evaluate whether any components related to the original audio information still exist in the stripped residual signal. Although preliminary quasi-periodic component stripping has been performed, some complex or weakly correlated reflection components may still exist. By using the cross-correlation coefficient, the correlation can be quantified, thereby more accurately distinguishing reflection components from external environmental noise.
[0046] The time-delay components with cross-correlation coefficients exceeding a first preset threshold are used as the acquisition results of reflection component information. The first preset threshold is a pre-defined discrimination criterion used to define the degree of cross-correlation that indicates the signal is a reflection of audio played from a digital audio system. When the cross-correlation coefficient between the stripped residual signal and the original audio information is higher than this threshold, it indicates that the signal is likely still a delayed version of the original audio, i.e., a reflection component.
[0047] The spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the stripped residual signal corresponding to cross-correlation coefficients below a second preset threshold are matched with a preset environmental noise feature template to identify the acquired external environmental noise. The second preset threshold is used to distinguish signals that are almost unrelated to the original audio information; these signals are more likely to be external environmental noise. For signals with low correlation, by extracting their acoustic features such as spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate, and matching them with the preset environmental noise feature template, different types of external environmental noise, such as fan noise, human voices, and traffic noise, can be effectively identified.
[0048] Specifically, a digital audio system is deployed in a large conference room with a long reverberation time. Multiple auxiliary microphones are deployed within the large conference room to collect sound data. First, based on the content played by the digital audio system and the room's acoustic model, the sound data that each auxiliary microphone should receive is predicted. Simultaneously, the auxiliary microphones actually collect sound data. The predicted sound data is compared with the collected sound data to obtain a residual signal. Next, this residual signal is analyzed to identify quasi-periodic components that have a time delay but similar waveform to the original music signal. These components are identified as echoes or reflections within the room. For example, these reflections can be identified by calculating the cross-correlation function between the residual signal and the original music signal and finding time points with significant peaks in the delay. After removing these reflection components, time-frequency analysis is performed on the remaining residual signal to confirm its cross-correlation coefficient with the original music signal. If the cross-correlation coefficient exceeds a first preset threshold, it is further confirmed as a reflection component. For signals with cross-correlation coefficients below a second preset threshold, features such as spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate are extracted and matched with pre-trained environmental noise templates (such as air conditioner noise, fan noise, and external traffic noise). This allows for accurate identification of digital speaker reflections and external environmental noise within the room, providing a precise basis for subsequent acoustic optimization.
[0049] In this embodiment, by refining the residual signal between the predicted sound data and the acquired sound data, quasi-periodic components with a time-synchronous causal relationship with the original audio information are identified and removed, thereby effectively separating the reflection components of the digital audio playback. Subsequently, time-frequency analysis and cross-correlation coefficient calculation are performed on the removed residual signal to further confirm and extract the reflection components. For the remaining signals with extremely low correlation to the original audio information, their unique acoustic features are extracted and matched with a preset environmental noise feature template to accurately identify external environmental noise. Through a step-by-step and refined processing flow, the problem of traditional methods being unable to distinguish between reflections and noise in complex acoustic environments can be effectively solved.
[0050] In some embodiments described above in this application, the present invention further provides a step in which the time delay component with a cross-correlation coefficient exceeding a first preset threshold is used as the acquisition result of the reflection component information. The step of constructing the first preset threshold includes: Based on the updated acoustic profile, the reverberation time, spatial geometry changes, and reflective surface material information recorded in the environmental acoustic characteristic description are confirmed. The environmental acoustic characteristic description, a part of the updated acoustic profile, details the acoustic characteristics of the current environment, such as reverberation time, spatial geometry changes, and reflective surface material information. Reverberation time refers to the time required for sound to decay to a specific level in space after it ceases to emit sound; it is an important indicator for measuring the acoustic characteristics of a space. Spatial geometry changes refer to changes in the spatial layout of the environment, such as room size, shape, and furniture placement, which affect the propagation and reflection paths of sound. Reflective surface material information refers to the material properties of various surfaces in the environment, such as sound-absorbing and reflective materials, which determine the degree to which sound is absorbed or reflected when it encounters these surfaces.
[0051] The first preset threshold is obtained by adjusting the preset base threshold using recorded reverberation time, spatial geometric changes, or reflective surface material information. The preset base threshold is an initial, universal cross-correlation coefficient threshold used to preliminarily determine the presence of reflective components. The first preset threshold is a dynamic threshold, adjusted based on the environmental acoustic characteristics, to better suit the current environment.
[0052] Specifically, a digital sound system is deployed in a large conference room with a long reverberation time. During playback, acoustic information collected by multiple auxiliary microphones is fused and used to update the acoustic profile. The environmental acoustic characteristics recorded in this updated profile indicate a long reverberation time and numerous hard reflective surfaces in the conference room. Based on these environmental characteristics, a preset base threshold (e.g., 0.7) is adjusted. Specifically, due to the long reverberation time, reflected signals may attenuate more, resulting in a relatively lower cross-correlation coefficient, thus lowering the base threshold to 0.65. Conversely, if the digital sound system is deployed in a small, well-absorbed recording studio, the updated acoustic profile indicates a short reverberation time and highly absorbent reflective surfaces. The reflected signals may be very weak, resulting in an even lower cross-correlation coefficient, potentially leading to a further reduction in the base threshold to 0.6 or even lower to ensure the capture of weak reflected components. In this way, the first preset threshold can adaptively match different acoustic environments, thereby optimizing the identification of reflected components.
[0053] In this embodiment, the reverberation time, spatial geometry changes, and reflective surface material information of the current environment can be obtained by using the environmental acoustic characteristic description contained in the acoustic update profile. This information directly reflects the propagation and reflection characteristics of sound waves in the current environment. For example, in a space with a long reverberation time, the cross-correlation coefficient of the reflection component may be relatively low, and the threshold needs to be appropriately lowered to avoid missed detections; while in a space with strong sound absorption of the reflective surface material, the reflection component may be weak, and corresponding threshold adjustments are also required. By using this real-time environmental acoustic characteristic information to adjust the preset base threshold, a first preset threshold that highly matches the current environment can be obtained, achieving dynamic adjustment of the first preset threshold. Therefore, when identifying the cross-correlation coefficient, it is possible to more accurately determine which time delay components are true reflection components, thereby effectively distinguishing the reflection component information of digital audio playback from background noise.
[0054] In some embodiments described above in this application, the present invention also provides the following acquisition results for using time delay components with cross-correlation coefficients exceeding a first preset threshold as reflection component information: Spatial consistency verification is performed on the cross-correlation coefficients to obtain coefficient verification information. Specifically, the cross-correlation coefficient refers to the degree of cross-correlation between the energy distribution of the stripped residual signal and the energy distribution of the original audio information. Spatial consistency verification can be understood as the process of cross-verifying and comparing the cross-correlation coefficients collected by different auxiliary microphones in the spatial dimension. Its purpose is to ensure that the identified reflection component information is spatially reasonable and consistent, avoiding misjudgments due to local anomalies or noise interference from a single auxiliary microphone. The coefficient verification information is the result of spatial consistency verification and is used to quantify the reliability or consistency of the cross-correlation coefficients in space.
[0055] The cross-correlation coefficients after verification are those that have passed the spatial consistency check and are confirmed to be greater than the preset verification threshold. The preset verification threshold is a predetermined value used to determine whether the coefficient verification information meets the spatial consistency requirements. When the coefficient verification information is greater than this threshold, it indicates that the cross-correlation coefficients have high spatial consistency and can therefore be considered reliable. The verified cross-correlation coefficients refer to the cross-correlation coefficients that have passed the spatial consistency check and are confirmed to meet the check conditions. These coefficients are considered to more accurately reflect the presence of the reflection component.
[0056] The time delay components whose cross-correlation coefficients exceed the first preset threshold after testing are taken as the results of the collection of reflection component information.
[0057] Specifically, during digital audio playback, multiple auxiliary microphones are deployed in different locations. After initially confirming the cross-correlation coefficients between the signals collected by each auxiliary microphone and the original audio information, it is found that the cross-correlation coefficients of certain time delay components exceed a first preset threshold. At this point, these time delay components are not immediately identified as reflection components. Instead, spatial consistency checks are further performed on these cross-correlation coefficients. For example, the similarity of cross-correlation coefficients between adjacent auxiliary microphones can be analyzed, or a spatial model can be constructed to predict the distribution of cross-correlation coefficients at different locations and compared with actual measurements. If a high cross-correlation coefficient reported by an auxiliary microphone is spatially inconsistent with the measurement results of its surrounding auxiliary microphones, or deviates significantly from the expected spatial distribution model, the coefficient verification information corresponding to the high cross-correlation coefficient will be lower than the preset verification threshold, and thus the time delay component will be excluded. Conversely, if multiple auxiliary microphones consistently report high cross-correlation coefficients spatially, and the coefficient verification information is greater than the preset verification threshold, then the time delay components corresponding to these verified cross-correlation coefficients will be ultimately identified as reflection components. For example, when a reflected wave is reflected from a wall, it is usually captured by multiple auxiliary microphones located on the reflection path with similar time delay and intensity. By checking the spatial consistency, this spatial correlation can be confirmed, thereby accurately identifying the reflected component.
[0058] In this embodiment, spatial consistency verification is introduced to address the potential for misjudgment when relying solely on a single cross-correlation coefficient. Specifically, when the reflection component information of audio played by a digital audio system truly exists, it typically exhibits certain propagation patterns and correlations in space. That is, the reflection signals collected by different auxiliary microphones at different locations, after appropriate time delay compensation, should have high spatial consistency in their cross-correlation coefficients. By performing spatial consistency verification on these cross-correlation coefficients, false high cross-correlation coefficients caused by factors such as local noise, sporadic interference, or sensor anomalies can be effectively filtered out, thereby ensuring that the identified reflection component information is genuine and spatially reliable. Therefore, only those time delay components that not only have cross-correlation coefficients exceeding a first preset threshold but are also spatially verified as consistent by multiple auxiliary microphones will be ultimately confirmed as reflection component information, significantly improving the accuracy and robustness of the identification.
[0059] In some embodiments described above in this application, the present invention also provides the step of performing spatial consistency verification on the cross-correlation coefficients to obtain coefficient verification information, which includes: Data quality assessment is performed on cross-correlation coefficients to obtain the assessment results. Specifically, data quality assessment refers to checking the reliability or validity of the calculated cross-correlation coefficients. For example, the amplitude, stability, continuity, or consistency with other relevant parameters of the cross-correlation coefficients can be evaluated. The purpose is to identify cross-correlation coefficients that may be unreliable due to measurement errors, environmental interference, or calculation anomalies.
[0060] When anomalies are found in the quality assessment results, the cross-correlation coefficients are corrected to obtain corrected cross-correlation coefficients. The quality assessment result refers to the output of the data quality assessment process, used to indicate the quality status of the cross-correlation coefficients. For example, the quality assessment result can be a Boolean value (normal / abnormal), a score, or a confidence index. Anomalies in the quality assessment results mean that the cross-correlation coefficients may be inaccurate or unreliable. Anomalies may include cross-correlation coefficients with excessively low amplitude (potentially indicating a weak signal or lack of correlation), excessively high amplitude (potentially indicating saturation or errors), drastic fluctuations, or significant inconsistencies with cross-correlation coefficients from other microphone arrays.
[0061] Spatial consistency verification is performed on the corrected cross-correlation coefficients to obtain coefficient verification information. Specifically, correcting the cross-correlation coefficients to obtain corrected cross-correlation coefficients refers to taking appropriate measures to improve their quality when anomalies are detected. For example, correction methods may include, but are not limited to, interpolating outliers (e.g., using cross-correlation coefficients from adjacent time points or adjacent microphones), filtering (e.g., low-pass filtering to smooth fluctuations), recalculating (if possible), or correcting based on statistical models. The aim is to eliminate or mitigate the impact of abnormal cross-correlation coefficients on subsequent spatial consistency verification, thereby improving data reliability. Therefore, corrected cross-correlation coefficients refer to cross-correlation coefficients whose reliability and accuracy have been improved after data quality assessment and correction processing, providing more reliable input for subsequent spatial consistency verification.
[0062] In this embodiment, by introducing a data quality assessment and correction step before spatial consistency verification of cross-correlation coefficients, the problem of poor data quality caused by various interferences is effectively solved. Specifically, data quality assessment can proactively identify cross-correlation coefficients that may be unreliable or biased, avoiding the direct use of low-quality data for verification. When anomalies are identified, the cross-correlation coefficients are corrected, for example, by using interpolation, filtering, or recalculation, to effectively eliminate or mitigate the impact of these anomalies on subsequent processing, thereby ensuring that the cross-correlation coefficients input to the spatial consistency verification stage are optimized and more representative data. Because the cross-correlation coefficients are preprocessed and optimized, the subsequent spatial consistency verification can be based on more reliable data, thereby significantly improving the accuracy and robustness of reflection component information identification. In other words, this application can effectively improve the accuracy and reliability of reflection component information identification in digital audio data acquisition. By conducting data quality assessment and correction on the cross-correlation coefficients before spatial consistency verification, misjudgments or omissions caused by quality problems of the original cross-correlation coefficients can be avoided, ensuring the validity of the verification results. This not only improves the system's ability to identify reflective components in complex acoustic environments, but also provides more accurate basic data for subsequent applications such as acoustic scene modeling, sound source localization, or echo cancellation, thereby significantly enhancing the practicality and performance of the entire digital audio data acquisition method.
[0063] In some embodiments described above in this application, the present invention also provides a step of matching the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the stripped residual signal corresponding to a cross-correlation coefficient lower than a second preset threshold with a preset environmental noise feature template to identify the acquisition results of external environmental noise, including: The residual signals after stripping, corresponding to cross-correlation coefficients below a second preset threshold, are spatially weighted and averaged to obtain a weighted average residual signal. The stripped residual signal refers to the signal components remaining after removing reflection components from the digital audio playback. Spatial weighted averaging involves combining the stripped residual signals from multiple auxiliary microphones and assigning different weights to each microphone based on factors such as its spatial location, distance from the source of anomalous acoustic components, or the acoustic characteristics of its local environment. For example, auxiliary microphones closer to the source of anomalous acoustic components or with higher signal quality can be given greater weight to ensure their data dominates the averaging process. Through this weighted averaging process, a weighted average residual signal that better represents the overall environmental noise characteristics can be obtained.
[0064] This method extracts the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope rate of change from the weighted average residual signal. The spectral centroid describes the center frequency of the signal spectrum, reflecting the brightness or sharpness of the sound. Spectral flatness measures the flatness of the spectrum, distinguishing between noise and tonal sounds. The zero-crossing rate indicates the number of times the signal crosses the zero axis per unit time, often used to differentiate speech from noise. The energy envelope rate of change describes how quickly the signal energy changes over time, reflecting the dynamic characteristics of the sound. This method comprehensively characterizes the acoustic properties of environmental noise.
[0065] The spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the weighted average residual signal are matched with a preset environmental noise feature template to identify the acquired external environmental noise. The preset environmental noise feature template is a pre-established database containing acoustic features of various typical environmental noises (e.g., fan noise, air conditioner noise, and background human voices). The matching process can be achieved by calculating the distance (e.g., Euclidean distance) or similarity (e.g., cosine similarity) between feature vectors. When the matching degree reaches a preset threshold, the acquired external environmental noise can be identified.
[0066] Specifically, multiple auxiliary microphones are deployed in the digital audio playback environment. Upon identifying anomalous acoustic components in the acoustic profile, corresponding sound data is acquired, and residual signals are determined and stripped. For the stripped residual signals with cross-correlation coefficients below a second preset threshold, a spatially weighted average is applied. Specifically, different weights are assigned based on the location of each auxiliary microphone, its distance from the digital audio system, and the acoustic characteristics of its local environment (e.g., local reverberation or occlusion information obtained from the acoustic profile). For example, microphones closer to anomalous acoustic components or in more stable acoustic environments can be assigned higher weights. This weighted average yields a weighted average residual signal that better represents the overall environmental noise characteristics. Subsequently, acoustic features such as spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate are extracted from this weighted average residual signal. These extracted features are then matched against a preset environmental noise feature template, for example, by calculating the Euclidean distance or cosine similarity between feature vectors to determine the degree of matching. When the matching degree reaches the preset threshold, the collection results of external environmental noise can be identified, thereby achieving accurate identification and separation of environmental noise.
[0067] In this embodiment, by performing spatial weighted averaging on the stripped residual signals corresponding to cross-correlation coefficients below a second preset threshold, acoustic information from different auxiliary microphones is effectively integrated. Weighted averaging effectively suppresses potential local noise interference or anomalous data from individual microphones, highlighting environmental noise characteristics prevalent across multiple microphones, thereby improving the signal-to-noise ratio and representativeness of the residual signals. The extracted features, such as spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate, more accurately reflect the acoustic characteristics of real environmental noise, making the subsequent matching process with the preset environmental noise feature template more precise and reliable.
[0068] Based on any of the above embodiments, please refer to the digital audio data acquisition method. Figure 2 The present invention also provides a digital audio data acquisition system, which includes a first acquisition module 210, a fusion module 220, an update module 230, a second acquisition module 240, a determination module 250, and an identification acquisition module 260.
[0069] The first acquisition module 210 is used to acquire acoustic change information of multiple auxiliary microphones, position information of multiple auxiliary microphones and timestamps during digital audio playback.
[0070] The fusion module 220 is used to fuse all acoustic change information, location information and timestamps to obtain fused acoustic information.
[0071] The update module 230 is used to update the preset acoustic profile using fused acoustic information to obtain an updated acoustic profile.
[0072] The second acquisition module 240 is used to identify abnormal acoustic component information in the acoustic update profile, and then acquire the acquired sound data obtained by the auxiliary microphone corresponding to the abnormal acoustic component information.
[0073] The determination module 250 is used to determine the predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information based on the collected sound data and the acoustic updated profile.
[0074] The identification and acquisition module 260 is used to identify the acquisition results, including the reflection component information of the audio played by the digital speaker and the external environmental noise information, based on the predicted sound data and the acquired sound data.
[0075] In this embodiment, the first acquisition module 210 continuously collects multi-source acoustic data, which is then efficiently integrated by the fusion module 220, and the update module 230 dynamically maintains the environmental acoustic profile. When an abnormal acoustic event is detected, the second acquisition module 240 can selectively acquire high-resolution data, followed by the determination module 250 predicting the sound based on the environmental model, and finally, the identification acquisition module 260 accurately separates the reflection component information and the external environmental noise information. This systematic approach overcomes the limitations of traditional static calibration, achieving non-invasive, continuous perception and accurate analysis of the dynamic acoustic environment, thereby significantly improving the adaptability and user experience of the digital audio system.
[0076] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention specification.
Claims
1. A method of digital sound data acquisition, characterized by, The method comprises the following steps: Obtaining acoustic change information of multiple auxiliary microphones during digital audio playback, position information of the multiple auxiliary microphones, and a timestamp; Fusing all the acoustic change information, position information, and timestamp to obtain fused acoustic information; Updating a preset acoustic image using the fused acoustic information to obtain an acoustic updated image; After identifying abnormal acoustic component information in the acoustic updated image, obtaining collected sound data collected by an auxiliary microphone corresponding to the abnormal acoustic component information; Based on the collected sound data and the acoustic updated image, determining predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information; Based on the predicted sound data and the collected sound data, identifying a collection result including reflection component information of the digital audio playback audio and external environmental noise information.
2. The method of claim 1, wherein, The step of obtaining acoustic change information of multiple auxiliary microphones during digital audio playback comprises: Obtaining initial sound data collected by each auxiliary microphone within a period of time during digital audio playback and original audio information of the digital audio playback; Performing data processing on the initial sound data to obtain frequency distribution information and root mean square energy information; Based on the frequency distribution information, root mean square energy information, original audio information, and a preset audio rule, determining acoustic change information containing acoustic event summary information or state feature information.
3. The method of claim 1, wherein, After identifying abnormal acoustic component information in the acoustic updated image, the step of obtaining collected sound data collected by an auxiliary microphone corresponding to the abnormal acoustic component information comprises: After identifying abnormal acoustic component information in the acoustic updated image, sending a data collection instruction with a probe mode parameter to the auxiliary microphone corresponding to the abnormal acoustic component information; Based on the data collection instruction, obtaining collected sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information, the collected sound data containing high-resolution data information and high-precision timestamp information.
4. The method of claim 2, wherein, The step of determining predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information based on the collected sound data and the acoustic updated image comprises: Taking fine-tuned response data collected by the auxiliary microphone corresponding to the probe mode parameter as the collected sound data; Determining predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information based on the collected sound data and the acoustic updated image.
5. The method of claim 4, wherein, The step of identifying a collection result including reflection component information of the digital audio playback audio and external environmental noise information based on the predicted sound data and the collected sound data comprises: After determining a residual signal between the predicted sound data and the collected sound data, confirming that there is a quasi-periodic component information with a time synchronization causal relationship when fine-tuning the residual signal and the original audio information; After stripping the quasi-periodic component information from the residual signal, performing time-frequency analysis on the stripped residual signal to confirm the energy distribution of the stripped residual signal in multiple preset time windows and frequency subbands; Confirming a cross-correlation coefficient between the energy distribution of the stripped residual signal and the energy distribution of the original audio information; The time delay component with the cross-correlation coefficient exceeding the first preset threshold is taken as the collection result of the reflection component information; The spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the residual signal after stripping corresponding to the cross-correlation coefficient lower than the second preset threshold are matched with a preset environmental noise feature template to identify the collection result of the external environmental noise.
6. The method of claim 5, wherein, In the step of taking the time delay component with the cross-correlation coefficient exceeding the first preset threshold as the collection result of the reflection component information, the step of constructing the first preset threshold comprises: Based on the acoustic update image, confirm the reverberation time, spatial geometric structure change, or reflection surface material information recorded in the environmental acoustic characteristic description; Adjust the preset basic threshold using the recorded reverberation time, spatial geometric structure change, or reflection surface material information to obtain the first preset threshold.
7. The method of claim 5, wherein the digital audio data is collected in a format of a.wav file. The step of taking the time delay component with the cross-correlation coefficient exceeding the first preset threshold as the collection result of the reflection component information comprises: Perform spatial consistency verification on the cross-correlation coefficient to obtain coefficient verification information; Confirm that the coefficient verification information is greater than the verified cross-correlation coefficient corresponding to the preset verification threshold; Take the time delay component with the verified cross-correlation coefficient exceeding the first preset threshold as the collection result of the reflection component information.
8. The method of claim 7, wherein, The step of performing spatial consistency verification on the cross-correlation coefficient to obtain coefficient verification information comprises: Perform data quality evaluation on the cross-correlation coefficient to obtain a quality evaluation result; When the quality evaluation result is abnormal, correct the cross-correlation coefficient to obtain a corrected cross-correlation coefficient; Perform spatial consistency verification on the corrected cross-correlation coefficient to obtain coefficient verification information.
9. The method of claim 3, wherein, The step of matching the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate of the residual signal after stripping corresponding to the cross-correlation coefficient lower than the second preset threshold with a preset environmental noise feature template to identify the collection result of the external environmental noise comprises: Perform spatial weighted averaging on the residual signal after stripping corresponding to the cross-correlation coefficient lower than the second preset threshold to obtain a weighted average residual signal; Extract the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate from the weighted average residual signal; Match the spectral centroid, spectral flatness, zero-crossing rate, and energy envelope change rate in the weighted average residual signal with a preset environmental noise feature template to identify the collection result of the external environmental noise.
10. A digital acoustic data acquisition system characterized by, The system comprises: A first acquisition module configured to acquire acoustic change information of a plurality of auxiliary microphones, position information of the plurality of auxiliary microphones, and a timestamp during digital sound playback; A fusion module configured to perform fusion processing on all the acoustic change information, position information, and timestamp to obtain fused acoustic information; An update module configured to update a preset acoustic image using the fused acoustic information to obtain an acoustic update image; A second acquisition module configured to, after identifying abnormal acoustic component information in the acoustic update image, acquire collected sound data obtained by sound collection of an auxiliary microphone corresponding to the abnormal acoustic component information; A determination module configured to determine predicted sound data collected by the auxiliary microphone corresponding to the abnormal acoustic component information based on the collected sound data and the acoustic update image; and A sound collection module configured to collect sound data based on the predicted sound data. The recognition acquisition module is configured to recognize, based on the predicted sound data and the acquired sound data, an acquisition result including reflection component information of the digital sound playback audio and external environmental noise information.