Intelligent configuration system of multi-channel sound equipment

By analyzing the signal and occlusion changes of the multi-channel audio system and combining the timing characteristics of temperature and echo delay drift, dynamic spatial reconstruction is achieved, solving the problem of the multi-channel audio system's inability to adapt to environmental changes and improving the sound field optimization capability and listening experience.

CN120602839AInactive Publication Date: 2025-09-05SHENZHEN HENGDA ZHITONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737636.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Multi-channel audio systems lack a real-time response mechanism to changes in acoustic signals and spatial occlusion, and are unable to compensate for timing characteristics through temperature drift and echo delay drift, resulting in unreasonable spatial sound field distribution, uneven sound coverage, and a reduced listening experience.

Method used

Through the spatial architecture modeling method based on the analysis of signal change characteristics and occlusion change characteristics, combined with the timing characteristic fusion algorithm of temperature drift and echo delay drift, dynamic space reorganization is performed to achieve adaptive sound field optimization.

Benefits of technology

It improves the real-time perception capability of spatial environment changes and the adaptive optimization level of sound field reconstruction, reduces the risk of acoustic distortion caused by misjudgment of spatial layout or drift of environmental parameters, and improves the systematic processing flow of acoustic data collection - environmental assessment - layout reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602839A_ABST
    Figure CN120602839A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent configuration system of a multi-channel sound box, relates to the technical field of audio signal processing and intelligent control, and is used for solving the problems of unreasonable space sound field distribution, non-uniform sound coverage and reduced auditory experience caused by incapability of intelligently disassembling and recombining according to actual space change. The method comprises the following steps: acquiring acoustic output signals of a multi-channel sound box in real time, synchronously monitoring the influence of space shielding on reflected signals, extracting signal change characteristics and shielding change characteristics of each channel, establishing a data analysis model, calculating a space architecture coefficient, comparing the space architecture coefficient with a preset threshold value, and determining a sound production field reconstruction condition result. Corresponding temperature sensor data and echo delay data are collected, a time sequence drift coefficient is obtained according to a time sequence algorithm, the reasonability of a preset reconstruction space layout is evaluated according to the time sequence drift coefficient and a space architecture coefficient, layout disassembly and recombination are carried out according to weighted assignment and space weight, and the real-time sensing capacity of space environment changes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio signal processing and intelligent control, and more particularly to an intelligent configuration system for multi-channel audio. Background Art

[0002] With the continuous advancement of intelligent audio technology, spatial acoustic modeling, and wireless communication technologies, multi-channel audio systems have been widely adopted in home entertainment, smart conferencing, virtual reality, public broadcasting, and other fields. Traditional multi-channel audio configuration systems typically rely on static layout assumptions and initial environmental modeling, rarely undergoing dynamic adjustments or optimizations after deployment. As a result, they are unable to adapt to changes in acoustic characteristics brought about by changes in the spatial environment. Furthermore, dynamic factors such as spatial obstruction, ambient temperature fluctuations, and echo reflection characteristics directly affect the overall sound field quality of the audio system.

[0003] The existing technology has the following deficiencies:

[0004] Currently, multi-channel audio systems lack a real-time response mechanism to changes in acoustic signals and spatial occlusions. They are unable to compensate for timing characteristics through temperature drift and echo delay drift. Spatial architecture judgment and reconstruction methods are limited, and they cannot intelligently disassemble and reassemble according to actual spatial changes. This leads to irrational spatial sound field distribution, uneven sound coverage, and a degraded auditory experience. Therefore, an intelligent configuration system for multi-channel audio is proposed.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0006] To overcome the above-mentioned shortcomings of the prior art, an embodiment of the present invention provides an intelligent configuration system for multi-channel audio, which solves the problems raised in the above-mentioned background technology through a spatial architecture modeling method based on the analysis of signal change characteristics and occlusion change characteristics, a timing characteristic fusion algorithm combining temperature drift and echo delay drift, and an adaptive sound field optimization mechanism that dynamically reorganizes the space based on spatial rationality evaluation and weight allocation.

[0007] To achieve the above objectives, the present invention provides the following technical solutions: an intelligent configuration system for multi-channel audio, comprising a data collection module, a spatial architecture model, a time sequence analysis module, and an architecture reorganization module, with signal connections between the modules;

[0008] The data collection module is used to collect the acoustic output signals of multi-channel speakers in real time and simultaneously monitor the impact of spatial occlusion on the reflected signals to obtain the signal change characteristics and occlusion change characteristics of each channel speaker;

[0009] The spatial architecture module is used to obtain the signal change characteristics and occlusion change characteristics of each channel sound, establish a data analysis model, obtain the spatial architecture coefficient, and compare it with the preset spatial threshold to analyze the risk of spatial distribution misjudgment and determine the sound field reconstruction condition results;

[0010] The timing analysis module is used to obtain the sound field reconstruction condition results, collect the corresponding temperature sensor data and echo delay data and fuse them to obtain the air temperature drift and echo delay drift of each speaker. Based on the timing algorithm, the timing drift coefficient is obtained;

[0011] The architecture reorganization module is used to evaluate the rationality of the preset reconstructed spatial layout based on the spatial architecture coefficient and the drift misjudgment coefficient, disassemble the preset reconstructed spatial layout, assign weights to each disassembled space, reorganize the space according to the weight values, and adjust the preset reconstructed spatial layout.

[0012] In a preferred embodiment, the signal change characteristics include the reflection signal characteristic change rate per unit time and the difference between the mean value of the spectrum and the mean value of the reference spectrum per unit time;

[0013] The echo signals of each audio channel are collected in real time, the signal characteristics are extracted, and the numerical differentiation method is applied to the continuously sampled characteristic sequence to calculate the change rate of the reflection signal characteristics per unit time;

[0014] The echo signal of each audio channel is collected in real time, and the spectrum information is extracted through fast Fourier transform. The spectrum mean in the current unit time is calculated, and the difference between the spectrum mean and the preset reference spectrum mean is calculated to obtain the mean difference between the spectrum mean and the reference spectrum in the unit time.

[0015] The occlusion change characteristics include the maximum duration of significant attenuation of channel sound energy and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction;

[0016] Collect echo sound energy data of each audio channel in real time, set the sound energy attenuation judgment threshold, and when the echo sound energy is less than the sound energy attenuation judgment threshold, it is counted as attenuation state, and the duration of continuous attenuation state is recorded. The longest duration is taken as the maximum duration of significant attenuation of channel sound energy;

[0017] Sound pressure level data were collected at fixed spatial position nodes under the preset layout and reconstructed layout states, and the consistency score results were calculated using the Pearson correlation coefficient to obtain the consistency score of the spatial sound pressure distribution before and after the preset reconstruction.

[0018] In a preferred embodiment, the data analysis model is a logistic regression model;

[0019] The characteristic change rate of the reflected signal per unit time, the difference between the mean of the spectrum and the mean of the reference spectrum per unit time, the maximum duration of significant attenuation of channel sound energy, and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction were standardized and substituted into the logistic regression formula to calculate the data correlation coefficient.

[0020] In a preferred embodiment, the spatial structure coefficient is compared with a preset structure threshold to obtain the sound field reconstruction condition result. If the spatial structure coefficient is greater than or equal to the structure threshold, the probability of spatial misjudgment is lower, the more there is a spatial structure change, and the more sufficient the sound field reconstruction condition is. If the spatial structure coefficient is less than the structure threshold, the probability of spatial misjudgment is higher, the less likely there is a spatial structure change, and the less sufficient the sound field reconstruction condition is.

[0021] In a preferred embodiment, after obtaining the sound field reconstruction condition result, the sound corresponding to the spatial architecture coefficient smaller than the architecture threshold is screened out, the remaining sound is marked as the remaining sound, and the temperature sensor data and echo delay data corresponding to the remaining sound are collected;

[0022] The air temperature drift is obtained by calculating the difference between the preset reconstructed space temperature and the preset reference temperature in real time;

[0023] Based on the transmission timestamp and the current reception timestamp in the preset reconstruction space, the current echo delay is calculated and subtracted from the reference delay value to obtain the echo delay drift.

[0024] In a preferred embodiment, the air temperature drift and the echo delay drift are normalized and substituted into the timing algorithm;

[0025] The time series algorithm is based on exponential sliding weighting. By performing exponential decay weighted fusion processing on the historical sequence sampling values, a dynamically updated time series feature estimation result is constructed to obtain the drift misjudgment coefficient.

[0026] In a preferred embodiment, the drift misjudgment coefficient is mapped to the inverse of the drift misjudgment coefficient and marked as the drift misjudgment coefficient mapping value;

[0027] Substitute the spatial structure coefficient and the drift misjudgment coefficient mapping values ​​into the geometric mean method to calculate and obtain the reasonable coefficient of the preset reconstruction space;

[0028] The preset reconstruction space rationality coefficient is compared with the preset space rationality threshold. If the preset reconstruction space rationality coefficient is greater than or equal to the space rationality threshold, the preset reconstruction space design is more reasonable, more stable and reliable, and the current round of system operation is terminated, and the current preset reconstruction space parameter configuration is recorded. If the preset reconstruction space rationality coefficient is less than the space rationality threshold, the preset reconstruction space design is more unreasonable and more likely to cause sound field drift, and the preset reconstruction space disassembly mechanism is activated.

[0029] In a preferred embodiment, the preset reconstruction space decomposition mechanism divides the area into multiple independent spaces according to the preset reconstruction layout, by collecting environmental noise data and echo delay difference;

[0030] Based on the non-excited sound signal data collected by the audio channel per unit time, by setting the noise detection cycle, under the condition of no active sound wave excitation, the background noise signal of the audio in each independent space in the current spatial environment is collected in real time, and substituted into the weighted filtering algorithm to obtain the environmental noise data;

[0031] For all adjacent pairs of speakers in the preset independent space, the echo delays between the adjacent pairs of speakers are calculated to obtain the echo delay differences;

[0032] The ambient noise data and echo delay difference are standardized and substituted into the weighted average method to obtain the spatial decomposition coefficient.

[0033] In a preferred embodiment, the space decomposition coefficient is compared with a preset decomposition threshold to obtain a preset reconstructed space layout decomposition result. If the space decomposition coefficient is greater than or equal to the decomposition threshold, the current corresponding independent space is decomposed. If the space decomposition coefficient is less than the decomposition threshold, the original space layout is maintained.

[0034] Each disassembled space is weighted according to its size and the spatial damping degree in the direction of the sound.

[0035] In a preferred embodiment, the size of the disassembled space is obtained based on the sound field structure diagram and the spatial boundaries of each speaker in the disassembled space.

[0036] The spatial damping degree in the direction of the sound is obtained by calculating the difference between the reference sound pressure level of the sound in an unobstructed environment and the actual sound pressure level at the current sound pressure receiving point, and then calculating the ratio with the straight-line distance in the direction of the sound wave propagation.

[0037] The weight assignment is based on the normalized fusion algorithm, which is obtained by substituting the size of the disassembled space and the normalized results of the spatial damping degree of the sound direction of the sound;

[0038] Compare the weight assignments of each adjacent disassembled space, spatially reorganize adjacent disassembled spaces with the same weight assignments, and adjust the preset reconstructed spatial layout.

[0039] Technical effects and advantages of the present invention:

[0040] 1. The present invention collects the acoustic output signals of multi-channel audio in real time and simultaneously monitors the impact of spatial occlusion on the reflected signal, extracts the signal change characteristics and occlusion change characteristics of each channel, establishes a data analysis model, calculates the spatial structure coefficient and compares it with the preset threshold, determines the sound field reconstruction condition results, collects the corresponding temperature sensor data and echo delay data, obtains the timing drift coefficient based on the timing algorithm, and evaluates the rationality of the preset reconstructed spatial layout with the spatial structure coefficient, and disassembles and reorganizes the layout based on weighted assignment and spatial weight, thereby improving the real-time perception capability of spatial environment changes and the adaptive optimization level of sound field reconstruction, reducing the risk of acoustic distortion caused by misjudgment of spatial layout or drift of environmental parameters, and improving the systematic processing flow of acoustic data collection-environmental assessment-layout reconstruction and the intelligent configuration mechanism of multi-channel audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a module diagram of an intelligent configuration system for multi-channel audio according to the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0043] Example 1

[0044] An intelligent configuration system for multi-channel audio, such as Figure 1 As shown, it includes a data collection module, a spatial architecture model, a timing analysis module, and an architecture reorganization module, and the signals between the modules are connected;

[0045] The data collection module is used to collect the acoustic output signals of multi-channel speakers in real time and simultaneously monitor the impact of spatial occlusion on the reflected signals to obtain the signal change characteristics and occlusion change characteristics of each channel speaker;

[0046] Deploy an array of high-sensitivity audio units in space, numbered M j , where j∈{1,2,…,M}, M is the number of speakers. Specifically, each speaker corresponds to its spatial coordinate (x j ,y i ,z j );

[0047] For each channel C i , where i∈{1,2,…,N}, N is the number of audio channels, and the original acoustic signal S emitted by it is collected in real time i(t), receiving the reflected and direct sound signals R through the sound array ij (t);

[0048] The sampling signal model is: R ij (t) = H ij (t)*S i (t)+n j (t);

[0049] Among them, from channel C i To Speaker M j The channel transfer function (including direct sound and reflected sound), n j (t) is the background noise of the sound, * is the convolution operation;

[0050] Use a unified clock source CLK to synchronize the sampling timing of each speaker, ensuring that all audio data has a unified timestamp T k , the sampling period is ΔT;

[0051] That is, every moment T k Collected:

[0052]

[0053] Microsecond-level synchronization error control is achieved through time synchronization protocol;

[0054] It should be noted that the time synchronization protocol achieves microsecond-level synchronization error control by configuring a unified hardware clock source for distributed clock synchronization, ensuring that each acquisition node has a unified and high-precision time reference during the data sampling process, thereby effectively reducing the multi-node sampling timing drift error and improving the data synchronization consistency of the entire system. This will not be elaborated here;

[0055] The signal change characteristics include the reflection signal characteristic change rate per unit time and the difference between the spectrum mean and the reference spectrum mean per unit time;

[0056] The reflection signal characteristic change rate per unit time is the rate at which the characteristic value of the echo signal of each audio channel changes per unit time, reflecting the dynamic fluctuations of the spatial acoustic environment on a short-term scale. Its acquisition logic is to collect the echo signals of each audio channel in real time, extract the signal characteristics, and apply the numerical differentiation method to the continuously sampled characteristic sequence to calculate the reflection signal characteristic change rate per unit time.

[0057] The unit time limit was set by our experimenters based on the typical response time constant of acoustic changes in the experimental scene and the system signal sampling rate, and will not be elaborated here.

[0058] Specifically, the calculation formula for the reflection signal characteristic change rate per unit time is:

[0059]

[0060] Where, ΔR i (t) is the characteristic change rate of the reflection signal per unit time, R i (t+Δt) is the reflection signal characteristic of the next unit time sampling interval, R i (t) is the reflection signal characteristic of the current unit time sampling interval, Δt is the sampling interval per unit time;

[0061] Furthermore, the signal features including reflected acoustic energy, echo delay, and spectrum center frequency are integrated into a single signal feature through feature vector splicing, and then integrated into a continuously sampled feature sequence through time series stacking;

[0062] It should be noted that the numerical differentiation method uses the first-order forward difference method, which is common knowledge among the experimenters and will not be described here;

[0063] The mean difference between the spectral mean and the reference spectrum per unit time is the difference between the spectral mean of the real-time acquired sound echo signal and the preset reference spectrum mean within the sampling interval per unit time. It is used to characterize the degree of change in the spectral characteristics of the current sound field environment. Its acquisition logic is to acquire the echo signal of each sound channel in real time, extract the spectrum information through fast Fourier transform, calculate the spectral mean within the current unit time, and calculate the difference between this spectral mean and the preset reference spectrum mean to obtain the mean difference between the spectral mean and the reference spectrum per unit time.

[0064] Among them, the spectrum information is extracted by fast Fourier transform, and its calculation formula is:

[0065]

[0066] Where X(f) is the frequency domain signal, x(n) is the sampling value of the time domain echo signal, N is the total number of sampling points, f is the frequency component, and j is the imaginary unit;

[0067] Specifically, the preset reference spectrum mean is set by the experimenters based on the statistical expectation value of the mean of multiple sampling spectrums in a static environment and the outlier elimination criteria, which will not be elaborated here;

[0068] Furthermore, the statistical expectation value of the mean of multiple sampling spectra in a static environment is obtained. Under the condition that the environmental noise and the structural layout remain static, multiple groups of echo signal sampling are performed, the spectrum data of each group of samples are extracted respectively, the mean of each frequency component is calculated, and the arithmetic mean of all sample spectrum means is taken as the reference spectrum mean;

[0069] The occlusion change characteristics include the maximum duration of significant attenuation of channel sound energy and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction;

[0070] The maximum duration of significant channel acoustic energy attenuation is the longest duration per unit time that the echo acoustic energy level of a particular audio channel is below a set threshold (e.g., a certain percentage of the baseline acoustic energy value). This is used to characterize the stability and abnormal characteristics of the channel under short-term obstruction or environmental interference. Its acquisition logic is to collect echo acoustic energy data from each audio channel in real time, set an acoustic energy attenuation determination threshold, and when the echo acoustic energy is less than the acoustic energy attenuation determination threshold, it is considered to be in an attenuation state. The duration of the continuous attenuation state is recorded, and the longest duration is taken as the maximum duration of significant channel acoustic energy attenuation.

[0071] Specifically, the threshold for determining acoustic energy attenuation is set by the experimenter based on historical echo acoustic energy data. Optionally, during long-term operation, in order to cope with the benchmark drift caused by environmental changes, the new reference acoustic energy mean and standard deviation under the barrier-free state can be periodically (for example, every 24 hours) re-collected and the threshold updated to maintain the stability of the system's discrimination ability. This is not detailed here.

[0072] The consistency score of the spatial sound pressure distribution before and after the preset reconstruction refers to the similarity score of the sound pressure level distribution of each key spatial node sampled before and after the spatial layout reconstruction. It is used to measure the consistency between the reconstructed sound field environment and the original sound field environment. Its acquisition logic is to collect sound pressure level data at fixed spatial position nodes under the preset layout and reconstructed layout states respectively, and use the mean square error between the two sets of data as the consistency score result to obtain the consistency score of the spatial sound pressure distribution before and after the preset reconstruction;

[0073] The specific calculation formula for the consistency score is the Pearson correlation coefficient:

[0074]

[0075] Where score is the consistency score of the spatial sound pressure distribution before and after the preset reconstruction, SPL pre,n and SPL post,n is the sound pressure level of the nth spatial node before and after reconstruction, and are the mean sound pressures of the spatial nodes before and after reconstruction, N is the number of spatial nodes, ranging from [-1, 1]. The closer to 1, the better the consistency.

[0076] Among them, for the collection of fixed spatial position nodes, the position nodes are determined based on the physical structure characteristics of the space and the sound field coverage requirements. Usually, the position nodes are fine-tuned using the spatial node layout strategy. Optionally, fine-tuning is performed through reflection path integrity analysis and signal coverage redundancy analysis, which will not be elaborated here.

[0077] The spatial architecture module is used to obtain the signal change characteristics and occlusion change characteristics of each channel sound, establish a data analysis model, obtain the spatial architecture coefficient, and compare it with the preset spatial threshold to analyze the risk of spatial distribution misjudgment and determine the sound field reconstruction condition results;

[0078] Specifically, the data analysis model is a logistic regression model;

[0079] In order to analyze the risk of misjudgment of spatial distribution and determine the conditions for reconstruction of the sound field, a logistic regression model was established to determine the spatial architecture coefficient;

[0080] The change rate of the reflection signal characteristics per unit time, the difference between the mean value of the spectrum per unit time and the mean value of the reference spectrum, the maximum duration of significant attenuation of channel sound energy, and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction are standardized. All input variables will be converted to the same range to ensure that each input contributes to the model in a balanced manner.

[0081] It should be noted that the standardization methods include but are not limited to standard linear transformation based on interval scaling, Z-Score standardization method based on statistics, or normalization method based on nonlinear mapping function. The application methods of standardization are not described in detail here.

[0082] The data correlation coefficient is calculated by substituting the characteristic change rate of the reflected signal per unit time, the difference between the mean value of the spectrum per unit time and the mean value of the reference spectrum, the maximum duration of significant attenuation of the channel sound energy, and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction into the logistic regression formula. The specific formula is expressed as follows:

[0083]

[0084] Where L is the result of logistic regression, i.e., the spatial framework coefficient, e is the natural base, and y is the linear combination term of the logistic regression model. Specifically, y is set as:

[0085]

[0086] Where β0 is the bias term, Cs j Sp is the characteristic change rate of the reflection signal of the j-th sound in unit time, j Jc is the difference between the mean value of the spectrum of the j-th sound per unit time and the mean value of the reference spectrum, jLs is the maximum duration of the significant attenuation of the channel sound energy of the j-th sound source. i is the consistency score of the spatial sound pressure distribution before and after the preset reconstruction of the j-th speaker, β1, β2, β3, and β4 are the regression coefficients of the characteristic change rate of the reflection signal per unit time, the difference between the mean value of the spectrum and the reference spectrum per unit time, the maximum duration of significant attenuation of channel sound energy, and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction, respectively.

[0087] Among them, if the change rate of the reflection signal characteristics per unit time, the difference between the mean value of the spectrum and the reference spectrum per unit time, and the maximum duration of significant attenuation of channel sound energy are larger, the increase in these three characteristic quantities usually indicates that the sound field environment has undergone a more obvious dynamic change or occlusion change, reflecting the actual change of the acoustic propagation path or obstacle layout in the space. In this case, the larger the spatial structure coefficient, the more likely there is a change in the spatial structure.

[0088] On the contrary, if the consistency score of the spatial sound pressure distribution before and after the reconstruction is greater, the degree to which the spatial layout remains consistent before and after the reconstruction is higher, the smaller the spatial structure coefficient is, the more similar the spatial and sound pressure distributions are after reconstruction, and the less likely there will be changes in the spatial structure;

[0089] Compare the spatial structure coefficient with the preset structure threshold to obtain the sound field reconstruction condition result. If the spatial structure coefficient is greater than or equal to the structure threshold, the probability of spatial misjudgment is lower, the spatial structure change is more likely to occur, and the sound field reconstruction condition is more sufficient. If the spatial structure coefficient is less than the structure threshold, the probability of spatial misjudgment is higher, the spatial structure change is less likely to occur, and the sound field reconstruction condition is less sufficient.

[0090] Sending the sound field reconstruction condition results to the timing analysis module;

[0091] It should be noted that the architectural threshold was determined by our researchers based on historical environmental sound field stability test results and analysis of abnormal sound pressure distribution detection rates in multiple scenarios, and will not be elaborated on here.

[0092] The timing analysis module is used to obtain the sound field reconstruction condition results, collect the corresponding temperature sensor data and echo delay data and fuse them to obtain the air temperature drift and echo delay drift of each speaker. Based on the timing algorithm, the timing drift coefficient is obtained;

[0093] After obtaining the sound field reconstruction condition results, the sound corresponding to the spatial architecture coefficient smaller than the architecture threshold is screened out, the remaining sound is marked as the remaining sound, and the temperature sensor data and echo delay data corresponding to the remaining sound are collected;

[0094] The temperature sensor data is obtained by real-time measurement through a high-precision digital temperature acquisition module embedded in the audio or audio unit, and is obtained by weighted smoothing filtering of multiple node temperature values ​​based on the sampling window corresponding to each unit time;

[0095] The echo delay data is obtained by combining the echo ranging algorithm under synchronous trigger control with the delay difference between the echo signal received by the speaker and the preset transmission unit time. The echo starting point is determined by triggering based on the set amplitude threshold and signal-to-noise ratio strategy.

[0096] The air temperature drift is the difference between the real-time temperature value of the remaining audio nodes in the current unit time and the temperature value at the preset reference time point. The acquisition logic is to calculate the difference between the preset reconstructed space temperature and the preset reference temperature in real time to obtain the air temperature drift;

[0097] The preset reference temperature is obtained by our experimenters based on the system initial calibration experimental data and the typical space environment temperature stability reference standard, which will not be described in detail here;

[0098] Echo delay drift refers to the delay change of the echo signal received by the same audio node before and after the reconstruction of the sound field. Its acquisition logic is based on the transmission timestamp and the current reception timestamp in the preset reconstruction space. The current echo delay is calculated and subtracted from the reference delay value to obtain the echo delay drift.

[0099] Specifically, in the process of calculating the difference between the current echo delay and the reference delay value, the determination of the starting receiving point is triggered by the set amplitude threshold and signal-to-noise ratio determination logic;

[0100] The air temperature drift and echo delay drift are normalized and substituted into the timing algorithm to obtain the timing drift coefficient.

[0101] The standardization process has been described in the above content and will not be repeated here;

[0102] Specifically, the time series algorithm is based on exponential sliding weighting;

[0103] Generally, exponential sliding weighting is a method used to smooth time series data;

[0104] Set the initial sliding weighted value to be equal to the initial sampling data, where the recursive formula of exponential sliding weighting is:

[0105] S(t)=α·x(t)+(1-α)·S(t-1);

[0106] Where S(t) is the exponential sliding weighted value at the current moment t, x(t) is the original sampled data at the current moment t, S(t-1) is the exponential sliding weighted value at the previous moment, and α is the smoothing factor used to control the impact of new data on the results.

[0107] Optionally, when α is large, the algorithm responds faster to new data changes, but the noise is amplified. When α is small, the smoothing effect is stronger, but the response to the latest changes is slower. α is usually set by system experience, and can also be dynamically adjusted according to the statistical characteristics of data noise;

[0108] Furthermore, exponential sliding weighting actually applies exponential decay weighting to historical data, which can be expanded to obtain:

[0109]

[0110] Where S(t) is the exponential sliding weighted value at the current moment t, which is used as the smoothed feature representation result, α is the smoothing factor, x(tk) is the original sample value at the kth unit time point before the current moment, k is the backtracking time index, which indicates the lag degree of historical data, (1-α) k is the decay weight of historical data. The larger k is (the farther back in time), the smaller the weight is.

[0111] The weight index of older data decreases, and the latest data has the highest weight, which meets the demand of "new data first" in a dynamic environment.

[0112] Specifically, the exponential sliding weighted formula constructs a class of dynamically updated time series feature estimation results by performing exponential decay weighted fusion processing on the historical sequence sampling values, where (1-α) k The decay factor reflects the decreasing trend of the contribution of historical data to the current results over time, and α controls the relative weight of the fusion of new and old data;

[0113] It should be noted that the relationship between the air temperature drift and echo delay drift of the remaining sound and the drift misjudgment coefficient is as follows: the larger the air temperature drift and echo delay drift, the larger the drift misjudgment coefficient, and the lower the rationality of the preset reconstruction space layout;

[0114] The architecture reorganization module is used to evaluate the rationality of the preset reconstructed spatial layout based on the spatial architecture coefficient and the drift misjudgment coefficient, and to disassemble the preset reconstructed spatial layout, assign weights to each disassembled space, reorganize the space based on the weight values, and adjust the preset reconstructed spatial layout;

[0115] The drift misjudgment coefficient is mapped to the inverse of the drift misjudgment coefficient and marked as the drift misjudgment coefficient mapping value, so that the numerical change trend of this indicator is consistent with the direction of spatial rationality assessment. That is, the larger the drift misjudgment coefficient, the smaller the mapping value, and thus the greater the irrationality of the preset reconstructed spatial layout in the rationality coefficient fusion calculation;

[0116] Substitute the spatial structure coefficient and the drift misjudgment coefficient mapping values ​​into the geometric mean method to calculate and obtain the reasonable coefficient of the preset reconstruction space;

[0117] Specifically, the geometric mean method is common knowledge among the experimenters and will not be described in detail here;

[0118] Compare the preset reconstruction space rationality coefficient with the preset space rationality threshold. If the preset reconstruction space rationality coefficient is greater than or equal to the space rationality threshold, the preset reconstruction space design is more reasonable, more stable and reliable, and the current round of system operation is terminated, and the current preset reconstruction space parameter configuration is recorded. If the preset reconstruction space rationality coefficient is less than the space rationality threshold, the preset reconstruction space design is more unreasonable and more likely to cause sound field drift, and the preset reconstruction space disassembly mechanism is activated;

[0119] The preset reconstruction space decomposition mechanism divides the area into multiple independent spaces according to the preset reconstruction layout. It collects ambient noise data and echo delay differences, normalizes them, and then uses the weighted average method to obtain the space decomposition coefficient. The space decomposition coefficient is compared with the preset decomposition threshold to obtain the decomposition result of the preset reconstructed space layout.

[0120] The logic for acquiring ambient noise data is based on the non-excited sound signal data collected by the audio channel per unit time. By setting a noise detection cycle and without active sound wave excitation, the background noise signal of the audio in each independent space in the current spatial environment is collected in real time. The noise signal is then substituted into the weighted filtering algorithm to obtain the ambient noise data.

[0121] Among them, the experimenters set the noise detection period based on the spatial static acoustic fluctuation standard and the noise stability time window modeling results; the no active sound wave excitation condition means that all audio units of the system are in a non-excitation state and no test signal or background music signal is sent during the detection period; the background noise signal is the background sound signal generated by non-system excitation sources such as the natural environment, human voice, and equipment operation; the weighted filtering algorithm uses a sliding weighted average or exponential weighted average method to suppress instantaneous pulse interference and improve the stability of noise assessment, which will not be described in detail here;

[0122] The echo delay difference reflects whether there are structural obstructions, echo overlap, or sound wave path disturbances in the preset spatial layout. Its acquisition logic is to calculate the difference in echo delay between all adjacent speaker pairs in the preset independent space to obtain the echo delay difference.

[0123] The ambient noise data and echo delay difference are normalized and substituted into the weighted average method to obtain the spatial decomposition coefficient;

[0124] The standardization process has been described above and will not be elaborated here.

[0125] Compare the space decomposition coefficient with the preset decomposition threshold to obtain the preset reconstructed space layout decomposition result. If the space decomposition coefficient is greater than or equal to the decomposition threshold, the current corresponding independent space is decomposed. If the space decomposition coefficient is less than the decomposition threshold, the original space layout is maintained.

[0126] It should be noted that the disassembly threshold was determined by our experimenters based on their experimental test experience and the results of multiple historical spatial layout evaluation experiments, and will not be elaborated here.

[0127] Optionally, the boundaries of the disassembled and segmented areas can be based on the obstacle layout information in the spatial geometric structure and the acoustic signal occlusion path map, or based on the sound pressure distribution equipotential line fitting results and the spatial sound intensity gradient variation law. Specifically, the number of disassembled and segmented areas is related to the spatial node density parameter and the perceptible sound field abnormality evaluation index. Based on the above content, the experimenters can further limit and calibrate the number of segmented areas and the division boundaries, which will not be elaborated here.

[0128] Assign weights to each disassembled space based on the size of the disassembled space and the spatial damping degree of the sound direction;

[0129] The logic for obtaining the size of the disassembled space is based on the sound field structure diagram, combined with the spatial boundaries of each speaker in the disassembled space to obtain the size of the disassembled space;

[0130] The logic for obtaining the spatial damping degree in the sound direction of the speaker is to calculate the difference between the reference sound pressure level of the speaker in an unobstructed environment and the actual sound pressure level at the current sound pressure receiving point, and then calculate the ratio with the straight-line distance in the direction of sound wave propagation to obtain the spatial damping degree in the sound direction of the speaker;

[0131] The weight assignment is based on the normalized fusion algorithm, which is obtained by substituting the size of the disassembled space and the normalized results of the spatial damping degree of the sound direction of the sound;

[0132] Specifically, the normalized fusion algorithm is expressed as follows:

[0133]

[0134] Where W k Assign a value to the b-th decomposition space weight, A b is the size of the b-th disassembly space, η b The damping degree of the sound direction of the b-th disassembled space sound, and is the weight distribution coefficient, satisfying

[0135] Compare the weight assignments of each adjacent disassembled space, spatially reorganize adjacent disassembled spaces with the same weight assignments, and adjust the preset reconstructed spatial layout;

[0136] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0137] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0138] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0139] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0140] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0142] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0143] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0144] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0145] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An intelligent configuration system for multi-channel audio, characterized by: It includes data collection module, spatial architecture model, timing analysis module and architecture reorganization module, and signal connections between each module; The data collection module is used to collect the acoustic output signals of multi-channel speakers in real time and simultaneously monitor the impact of spatial occlusion on the reflected signals to obtain the signal change characteristics and occlusion change characteristics of each channel speaker; The spatial architecture module is used to obtain the signal change characteristics and occlusion change characteristics of each channel sound, establish a data analysis model, obtain the spatial architecture coefficient, and compare it with the preset spatial threshold to analyze the risk of spatial distribution misjudgment and determine the sound field reconstruction condition results; The timing analysis module is used to obtain the sound field reconstruction condition results, collect the corresponding temperature sensor data and echo delay data and fuse them to obtain the air temperature drift and echo delay drift of each speaker. Based on the timing algorithm, the timing drift coefficient is obtained; The architecture reorganization module is used to evaluate the rationality of the preset reconstructed spatial layout based on the spatial architecture coefficient and the drift misjudgment coefficient, disassemble the preset reconstructed spatial layout, assign weights to each disassembled space, reorganize the space according to the weight values, and adjust the preset reconstructed spatial layout.

2. The intelligent configuration system for multi-channel audio according to claim 1, characterized in that: The signal change characteristics include the reflection signal characteristic change rate per unit time and the difference between the spectrum mean and the reference spectrum mean per unit time; The echo signals of each audio channel are collected in real time, the signal characteristics are extracted, and the numerical differentiation method is applied to the continuously sampled characteristic sequence to calculate the change rate of the reflection signal characteristics per unit time; The echo signal of each audio channel is collected in real time, and the spectrum information is extracted through fast Fourier transform. The spectrum mean in the current unit time is calculated, and the difference between the spectrum mean and the preset reference spectrum mean is calculated to obtain the mean difference between the spectrum mean and the reference spectrum in the unit time. The occlusion change characteristics include the maximum duration of significant attenuation of channel sound energy and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction; Collect echo sound energy data of each audio channel in real time, set the sound energy attenuation judgment threshold, and when the echo sound energy is less than the sound energy attenuation judgment threshold, it is counted as attenuation state, and the duration of continuous attenuation state is recorded. The longest duration is taken as the maximum duration of significant attenuation of channel sound energy; Sound pressure level data were collected at fixed spatial position nodes under the preset layout and reconstructed layout states, and the consistency score results were calculated using the Pearson correlation coefficient to obtain the consistency score of the spatial sound pressure distribution before and after the preset reconstruction.

3. The intelligent configuration system for multi-channel audio according to claim 2, characterized in that: The data analysis model was a logistic regression model; The characteristic change rate of the reflected signal per unit time, the difference between the mean of the spectrum and the mean of the reference spectrum per unit time, the maximum duration of significant attenuation of channel sound energy, and the consistency score of the spatial sound pressure distribution before and after the preset reconstruction were standardized and substituted into the logistic regression formula to calculate the data correlation coefficient.

4. The intelligent configuration system for multi-channel audio according to claim 3, characterized in that: The spatial structure coefficient is compared with the preset structure threshold to obtain the sound field reconstruction condition result. If the spatial structure coefficient is greater than or equal to the structure threshold, the probability of spatial misjudgment is lower, the spatial structure change is more likely to occur, and the sound field reconstruction condition is more sufficient. If the spatial structure coefficient is less than the structure threshold, the probability of spatial misjudgment is higher, the spatial structure change is less likely to occur, and the sound field reconstruction condition is less sufficient.

5. The intelligent configuration system for multi-channel audio according to claim 4, characterized in that: After obtaining the sound field reconstruction condition results, the speakers corresponding to the spatial architecture coefficients smaller than the architecture threshold are screened out, the remaining speakers are marked as remaining speakers, and the temperature sensor data and echo delay data corresponding to the remaining speakers are collected; The air temperature drift is obtained by calculating the difference between the preset reconstructed space temperature and the preset reference temperature in real time; Based on the transmission timestamp and the current reception timestamp in the preset reconstruction space, the current echo delay is calculated and subtracted from the reference delay value to obtain the echo delay drift.

6. The intelligent configuration system for multi-channel audio according to claim 5, characterized in that: Normalize the air temperature drift and echo delay drift and substitute them into the timing algorithm; The time series algorithm is based on exponential sliding weighting. By performing exponential decay weighted fusion processing on the historical sequence sampling values, a dynamically updated time series feature estimation result is constructed to obtain the drift misjudgment coefficient.

7. The intelligent configuration system for multi-channel audio according to claim 1, characterized in that: Map the drift misjudgment coefficient to the inverse of the drift misjudgment coefficient and mark it as the drift misjudgment coefficient mapping value; Substitute the spatial structure coefficient and the drift misjudgment coefficient mapping values ​​into the geometric mean method to calculate and obtain the reasonable coefficient of the preset reconstruction space; The preset reconstruction space rationality coefficient is compared with the preset space rationality threshold. If the preset reconstruction space rationality coefficient is greater than or equal to the space rationality threshold, the preset reconstruction space design is more reasonable, more stable and reliable, and the current round of system operation is terminated, and the current preset reconstruction space parameter configuration is recorded. If the preset reconstruction space rationality coefficient is less than the space rationality threshold, the preset reconstruction space design is more unreasonable and more likely to cause sound field drift, and the preset reconstruction space disassembly mechanism is activated.

8. The intelligent configuration system for multi-channel audio according to claim 7, characterized in that: The preset reconstruction space disassembly mechanism divides the area into multiple independent spaces according to the preset reconstruction layout, by collecting environmental noise data and echo delay difference; Based on the non-excited sound signal data collected by the audio channel per unit time, by setting the noise detection cycle, under the condition of no active sound wave excitation, the background noise signal of the audio in each independent space in the current spatial environment is collected in real time, and substituted into the weighted filtering algorithm to obtain the environmental noise data; For all adjacent pairs of speakers in the preset independent space, the echo delays between the adjacent pairs of speakers are calculated to obtain the echo delay differences; The ambient noise data and echo delay difference are standardized and substituted into the weighted average method to obtain the spatial decomposition coefficient.

9. The intelligent configuration system for multi-channel audio according to claim 8, characterized in that: Compare the space decomposition coefficient with the preset decomposition threshold to obtain the preset reconstructed space layout decomposition result. If the space decomposition coefficient is greater than or equal to the decomposition threshold, the current corresponding independent space is decomposed. If the space decomposition coefficient is less than the decomposition threshold, the original space layout is maintained. Each disassembled space is weighted according to its size and the spatial damping degree in the direction of the sound.

10. The intelligent configuration system for multi-channel audio according to claim 9, characterized in that: Based on the sound field structure diagram and the spatial boundaries of each speaker in the disassembled space, the size of the disassembled space is obtained; The spatial damping degree in the direction of the sound is obtained by calculating the difference between the reference sound pressure level of the sound in an unobstructed environment and the actual sound pressure level at the current sound pressure receiving point, and then calculating the ratio with the straight-line distance in the direction of the sound wave propagation. The weight assignment is based on the normalized fusion algorithm, which is obtained by substituting the size of the disassembled space and the normalized results of the spatial damping degree of the sound direction of the sound; Compare the weight assignments of each adjacent disassembled space, spatially reorganize adjacent disassembled spaces with the same weight assignments, and adjust the preset reconstructed spatial layout.