Three-dimensional space data dynamic sampling control system and method

By performing time-frequency analysis and optimizing HOA encoding on microphone array signals, constructing a channel reliability index, and dynamically adjusting radial back filtering, the problems of pseudo-spatial information and misjudgment of radial back filtering in multi-channel synchronous acquisition are solved, thereby improving the reliability and stability of three-dimensional spatial audio data.

CN121531270APending Publication Date: 2026-02-13XIAMEN YIJIE AUDIOVISUAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511946208.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing multi-channel synchronous acquisition and signal processing systems are prone to false spatial information mixing due to channel anomalies and noise pollution when facing three-dimensional spatial audio scenarios with multi-unit intelligent microphone arrays. Radial back filtering over-amplifies high-order noise, causing misjudgment of spatial complexity and deterioration of high-order high-frequency signal-to-noise ratio.

Method used

The microphone array channel signals are acquired by the signal acquisition module. After time and frequency analysis, the HOA encoding method is optimized, the channel reliability index is constructed, the radial back filter is adjusted, the configuration parameters of the three-dimensional spatial data sampling are dynamically adjusted, abnormal channels and strong noise channels are weakened, the real spatial complexity index is constructed, and the configuration order and bit rate are adjusted in conjunction.

Benefits of technology

It effectively reduces the probability of high-order spherical harmonic coefficients being contaminated by pseudo-structures, improves the authenticity and stability of spherical harmonic coefficients, suppresses the excessive amplification of high-order noise by radial back filtering, ensures the reliability and adaptability of spatial coding, and reduces the transmission of invalid high-order data and computational burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531270A_ABST
    Figure CN121531270A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional space data dynamic sampling control system and method, and belongs to the technical field of three-dimensional space data control, and the system comprises a signal collection module which is used for obtaining an original signal of each microphone channel in a microphone array, carrying out the time frequency analysis of the original signal of each microphone channel, and obtaining a time frequency analysis result; obtaining a sound pressure vector of each microphone channel; the spherical harmonic coding module is used for optimizing the HOA coding mode, coding the sound pressure vector of each microphone channel according to the optimized HOA coding mode, and obtaining a multi-order spherical harmonic coefficient vector to form a spherical harmonic coefficient sequence; and the dynamic control module is used for adjusting radial inverse filtering, acquiring an order energy sequence based on the spherical harmonic coefficient sequence, weighting high-order energy in the order energy sequence to construct a space complexity index, comparing the space complexity index with an input task demand complexity index, and dynamically adjusting configuration parameters of three-dimensional space data dynamic sampling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional space data control, and in particular to a three-dimensional space data dynamic sampling control system and method. BACKGROUND

[0002] With the increasing requirements of virtual reality and industrial hearing applications for spatial presence perception and sound source positioning accuracy, multi-channel synchronous data acquisition as the source of spatial sound field reconstruction usually relies on shared master clock, hardware trigger or high-precision time synchronization mechanism to realize cross-channel sampling alignment, and forms multi-channel sound pressure data blocks with consistent timestamps and frame lengths through multi-channel analog-digital conversion links with uniform sampling rate and bit depth, so that multi-channel sound pressure vectors can be reliably assembled within the same control frame or time frequency unit, providing calibratable and traceable input conditions for subsequent spatial feature extraction and high-order encoding. On this basis, three-dimensional spatial audio acquisition and high-order ambisonic encoding based on multi-unit intelligent microphone arrays gradually become an important technical route to realize high-fidelity spatial sound field representation, which can map multi-channel sound pressure into multi-order spherical harmonic coefficients under the constraint of array geometry, thereby providing a unified and extensible three-dimensional spatial data basis for subsequent decoding rendering and cross-platform transmission.

[0003] For example, the multi-channel multi-trigger mode synchronous data acquisition method and system disclosed in the Chinese patent CN118210266B includes: a control terminal sets the measurement point information and the trigger mode and sends them to an ARM system; a collection device uploads the sampling data of each sampling rate corresponding to each trigger channel to the ARM system; the ARM system synchronously searches for the sampling data in the data packets of the sampling data of each trigger channel according to the trigger position information, records the sampling data that meets the set conditions, and records the trigger position information in one-dimensional data; after the sampling data collection is completed, the trigger channel information, the sampling data corresponding to the sampling rate, and the trigger position information are recombined into three-dimensional data.

[0004] For example, the intelligent substation digital sampling synchronization conversion method and device disclosed in the Chinese patent CN109596949B includes: an optical fiber transceiver module is used for sampling digital quantity signals; an FPGA module is used for receiving digital quantity sampling signals; a main control CPU module is used for configuring adaptive conversion parameters according to the received digital quantity sampling signals, and the adaptive conversion parameter configuration includes determination of a rated delay time, determination of a rated phase voltage, and determination of a channel mapping relationship; an adjusted output sampling value and an integral period delay time are determined according to the adaptive conversion parameters; a D / A module is used for converting the output sampling value into an output analog quantity according to a preset output control ratio, and delaying the output according to the integral period delay time.

[0005] The above-mentioned technology has at least the following technical problems: the existing multi-channel synchronous acquisition and related signal processing system can realize the time sequence consistency of cross-channel data through the means of master clock sharing, trigger alignment or time synchronization, and support the recombination, mapping or parameter adaptive configuration of the sampling data, but there are still obvious deficiencies in the three-dimensional spatial audio scene of multi-unit intelligent microphone array: on the one hand, when there are channel abnormalities, noise pollution and improper numerical compensation in the acquisition link, a large amount of pseudo spatial information unrelated to the real sound field is often mixed in the high-order HOA coefficient, and the decision is made directly according to the rough indicators such as high-order energy distribution and coefficient variance, which is easy to misjudge the channel abnormality and noise pseudo structure as the change of spatial field complexity, so that the abnormal channel is easy to introduce error accumulation after entering the subsequent spatial encoding or feature calculation link; on the other hand, since the microphone of the microphone array is a rigid sampling device, the radial inverse filtering is originally used to compensate the influence of the spherical shell scattering on the signal, and in the high-order and high-frequency scene, the radial transfer function value corresponding to the radial inverse filtering is small, and the corresponding radial inverse filtering is large, which not only amplifies the real signal component of the current order, but also amplifies the high-order noise component introduced by the channel noise, channel mismatch and numerical error in the same proportion, so that the signal-to-noise ratio of the high-order high-frequency subspace is not only not improved, but also significantly deteriorated, further leading to the misjudgment of the spatial complexity. SUMMARY

[0006] In order to solve the technical problems of pseudo spatial information and radial inverse filtering over-amplification causing spatial complexity misjudgment in the prior art, the embodiments of the present application provide a three-dimensional spatial data dynamic sampling control system and method. The technical solution is as follows:

[0007] On the one hand, a three-dimensional spatial data dynamic sampling control system is provided, which comprises: a signal acquisition module, configured to acquire the original signals of each microphone channel in a microphone array, and perform time-frequency analysis on the original signals of each microphone channel to obtain a sound pressure vector of each microphone channel; a spherical harmonic coding module, configured to optimize the HOA coding mode, encode the sound pressure vector of each microphone channel according to the optimized HOA coding mode, and obtain a spherical harmonic coefficient sequence formed by a multi-order spherical harmonic coefficient vector; a dynamic control module, configured to adjust the radial inverse filtering, and then obtain an order energy sequence based on the spherical harmonic coefficient sequence, weight the high-order energy in the order energy sequence to construct a spatial complexity index, compare the spatial complexity index with an input task requirement complexity index, and dynamically adjust the configuration parameters of the three-dimensional spatial data dynamic sampling.

[0008] In another aspect, a three-dimensional spatial data dynamic sampling control method is provided, which comprises: S1. obtaining original signals of microphone channels in a microphone array, and performing time-frequency analysis on the original signals of the microphone channels to obtain sound pressure vectors of the microphone channels; S2. optimizing an HOA coding mode, encoding the sound pressure vectors of the microphone channels according to the optimized HOA coding mode to obtain a spherical harmonic coefficient sequence formed by a plurality of spherical harmonic coefficient vectors; S3. adjusting radial de-filtering, and then obtaining an order energy sequence based on the spherical harmonic coefficient sequence, weighting high-order energy in the order energy sequence to construct a spatial complexity index, comparing the spatial complexity index with an input task requirement complexity index, and dynamically adjusting configuration parameters of three-dimensional spatial data dynamic sampling.

[0009] Advantages

[0010] The technical scheme provided by the embodiments of the present application has at least the following advantages:

[0011] 1. The three-dimensional spatial data dynamic sampling control system and method provided by the present application introduce a channel reliability index before coding and construct a sound vector weighting coefficient based on the channel reliability index, so as to realize hierarchical soft weakening or hard removal of abnormal channels, strong noise channels and gain-phase mismatch channels, complete reliability pre-constraint of a multi-channel sound pressure vector before entering a spherical harmonic coding link, reduce the probability of pollution of high-order spherical harmonic coefficients by channel pseudo-structure from the source, and avoid the problem of misjudgment of channel abnormalities as spatial complexity changes only according to the coarse index of high-order energy or coefficient variance in the prior art, thereby improving the authenticity, stability and interpretability of the spherical harmonic coefficient sequence.

[0012] 2. The present application introduces a radial de-filtering adaptive adjustment mechanism driven by an order reliability index for scattering compensation of a rigid spherical array, which can weaken the original design radial gain in a high-order, low signal-to-noise ratio or frequency band close to the physical effective order limit, or improve the inversion stability in an adaptive regularization manner, so as to suppress the excessive amplification of radial de-filtering to high-order noise and spatial aliasing artifacts, while retaining the necessary compensation strength in the task critical frequency band and the high reliability order range, so that the output standard spherical harmonic coefficient sequence obtains a more robust engineering compromise between the noise upper bound and the spatial detail fidelity.

[0013] 3、The application constructs order energy sequence and real spatial complexity index based on the spherical harmonic coefficient sequence corrected by channel weighting and radial compensation, and forms a supply-demand deviation criterion with the input task requirement complexity index, and then adjusts the configuration order and configuration code rate, so that the switching of order and code rate is jointly constrained by multiple evidences such as reliable high-order energy margin, reconstruction residual of candidate order and marginal improvement, thereby significantly reducing invalid high-order data, reducing transmission and computing power burden, avoiding the risk of uncontrollable quality fluctuation caused by false order increase and false high code rate preservation, and enhancing the adaptability and reproducible engineering effect of the system under different array forms and complex sound field conditions. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0015] Figure 1 A structural schematic diagram of a three-dimensional spatial data dynamic sampling control system provided by the embodiment of the present application;

[0016] Figure 2 A flowchart of a three-dimensional spatial data dynamic sampling control method provided by the embodiment of the present application;

[0017] Figure 3 A channel reliability evaluation and weighting encoding flowchart provided by the embodiment of the present application;

[0018] Figure 4 A radial de-filtering adjustment and dynamic control flowchart provided by the embodiment of the present application. DETAILED DESCRIPTION

[0019] In the following, some terms in the present application are explained. It should be noted that these explanations are for the convenience of understanding by those skilled in the art, and do not limit the scope of protection required by the present application.

[0020] At least one of the embodiments of the present application includes one or more; wherein the plurality means greater than or equal to two. In addition, it needs to be understood that in the description of the specification, the words "first", "second", "third", etc. are used only for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance. For example, the first device and the second device do not represent the importance of the two or the order of the two, but only for the purpose of distinguishing the description. In the embodiments of the present application, "and / or" is only to describe the relationship between the two, which means that there are three kinds of relationships, for example, A and / or B, which can represent the existence of A alone, the existence of A and B, and the existence of B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a kind of "or" relationship.

[0021] The orientation terms mentioned in the embodiments of the present application, such as "up", "down", "left", "right", "in", "out" and the like, are only the direction of the drawings, therefore, the orientation terms used are for better and clearer description and understanding of the embodiments of the present application, and are not indicative or implied that the devices or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, therefore, cannot be understood as a limitation on the embodiments of the present application.

[0022] The reference "one embodiment", "in some examples" or "some embodiments" and the like described in the embodiments of the present application means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the specification. Therefore, the statements "in some examples", "in one embodiment", "in some embodiments", "in other some embodiments", "in other some embodiments" and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0023] In order to make the technical problems, technical solutions and advantages of the present application more clear, the following will be described in detail in conjunction with the drawings and specific embodiments.

[0024] As Figure 1As shown, a structural schematic diagram of a three-dimensional space data dynamic sampling control system provided by an embodiment of the present application, the system comprises the following modules: a signal acquisition module, configured to acquire original signals of microphone channels in a microphone array, and perform time-frequency analysis on the original signals of the microphone channels to obtain sound pressure vectors of the microphone channels; a spherical harmonic coding module, configured to optimize an HOA coding mode, encode the sound pressure vectors of the microphone channels according to the optimized HOA coding mode, and obtain a spherical harmonic coefficient sequence formed by a multi-order spherical harmonic coefficient vector; and a dynamic control module, configured to adjust a radial de-filtering, and then acquire an order energy sequence based on the spherical harmonic coefficient sequence, weight high-order energy in the order energy sequence to construct a spatial complexity index, compare the spatial complexity index with an input task requirement complexity index, and dynamically adjust configuration parameters of three-dimensional space data dynamic sampling.

[0025] The original signal refers to a pre-coding basis sampling signal obtained by the microphone array channels at the same acquisition time base, including a synchronous digital pickup signal, channel acquisition configuration and channel state information, a mute frame and a low energy segment signal. The signal acquisition module performs parallel sampling on the analog sound and electricity output of each microphone in the array through a multi-channel synchronous acquisition link, and uses a shared master clock, hardware trigger alignment or an equivalent time synchronization mechanism to make each channel have consistent sampling start and end boundaries, time stamps and frame structures in the same control time window, thereby forming multi-channel basis data that can be directly used for time-frequency analysis and subsequent spherical harmonic coding. The channel acquisition configuration information refers to a series of parameters that are preset and fixed before the acquisition starts and determine the basic attributes of the signal, mainly including: a sampling rate, a channel gain for converting the original count value of an analog-to-digital converter (ADC) into a physical sound pressure unit, array geometry calibration parameters, and a cutoff frequency of an anti-aliasing filter, etc., which collectively define the basic framework of the signal digitization process. The channel gain is a kind of conversion from the original count value of an analog-to-digital converter to a physical sound pressure unit. The array geometry calibration parameters are, for example, the accurate three-dimensional coordinates of each microphone unit in a preset coordinate system. The channel state information refers to dynamic metadata generated and associated in real time during the synchronous acquisition process, which is used to monitor the signal quality and system health status. The core parameters usually include: a synchronization state flag, a real-time signal level statistics, an overload indication flag, and a signal-to-noise ratio estimate value. The synchronization state flag indicates whether the channel sampling clock is locked with the master clock. These two types of information and the synchronous digital pickup signal, mute frame and low energy segment signal itself together constitute the original signal, providing an indispensable physical scale basis and signal quality judgment basis for subsequent time-frequency analysis, spherical harmonic coding and even dynamic sampling control, ensuring the integrity and reliability of the processing from the original sound pressure to the high-order spatial representation.

[0026] It should be noted that after obtaining the spherical harmonic coefficient sequence, the system aggregates the multi-order spherical harmonic coefficients in the order to form the corresponding order sub-vector, specifically, for the multi-order spherical harmonic coefficient vector corresponding to any time frame or time-frequency unit, the system identifies the order mark according to the preset spherical harmonic index sequence, extracts all spherical harmonic components under the same order, and rearranges and combines the coefficients in the order according to the predetermined order to form the HOA sub-vector of the order, so that the sub-vector only contains the independent contribution of the order to the spatial sound field; further, in the preset control time window, the system aligns and merges the HOA sub-vectors of the same order according to the time stamp or time frame index, thereby forming a set of order sub-vectors or an order sub-vector sequence that is continuous in time, ensuring that subsequent statistics of the energy of the order band signal, the noise energy of the silent frame, and the signal-to-noise ratio of the order band can be based on the same order, the same frequency band, and the interpretable data subspace. Through the above order aggregation mechanism, the spatial detail contributions of different orders are explicitly separated from the original full-order mixed coefficients, which not only provides a clear data structure basis for the calculation of the order-specific reliability index, but also enables the radial inverse filtering adjustment factor to specifically suppress the risk in the high-order low-reliability area at the order frequency band granularity.

[0027] In order to facilitate the understanding of the order selection, reliability evaluation and radial compensation adjustment logic of the present application, in the present specification, the maximum order is used to describe the highest spherical harmonic expansion order limit allowed by the system geometry and design; the configuration order is used to describe the encoding order actually used by the system in the current control time window and satisfies the configuration order less than or equal to the maximum order; the candidate order is used to describe the discrete alternative order set used for comparison in the multi-order reconstruction residual evaluation; the channel reliability index is used to describe the reliability of a single channel in sampling the real sound field in the current time window; the order-specific reliability index is used to describe the interpretable contribution reliability of a certain order to the real spatial details under the constraints of a specific frequency band and channel reliability; the unified convention of the above terms is used to ensure that the sound vector weighting, radial inverse filtering adjustment, real spatial complexity construction, and configuration order and configuration code rate linkage control are closed-loop operated under the same interpretable semantic system.

[0028] Embodiment one of the present application: provide a three-dimensional space data dynamic sampling control system and method, applied to the multi-unit intelligent microphone array arranged around the sound field collection area. The system first acquires the synchronous digital pickup signal of each microphone channel in the array, and according to the installation position information of the microphone in the three-dimensional space, converts the position of each microphone into the azimuth and elevation of the spherical coordinate form, calculates the corresponding spherical harmonic function value based on the preset maximum order to construct the spherical harmonic sampling matrix; then the time and frequency analysis of each channel signal is performed to obtain the multi-channel sound pressure vector under each time frame and each frequency, for example, a multi-channel synchronous collection structure of 48kHz, 24bit can be used, and the sampling alignment is realized through the shared master clock, hardware trigger or IEEE1588PTP method, so that in any 10ms control frame, each channel forms a digital pickup segment with consistent length and consistent timestamp, which can be directly assembled into a multi-channel sound pressure vector under the same time and frequency unit. The 10ms control frame is used to describe the minimum alignment granularity of collection and time frequency analysis, and ensures that the multi-channel sound pressure vector has comparability and assemblability on the same time and frequency unit. The control time window is used to describe the minimum control scale of the strategy decision, and the control time window can be composed of several continuous control frames, for example, a plurality of 10ms control frames are spliced to form a statistical window of 50ms, 100ms or 200ms, so as to obtain sufficient stable observation samples when evaluating the channel credibility, estimating the signal-to-noise ratio of the frequency band of the order, constructing and configuring the order credibility index, and making the order rate linkage decision, so as to avoid the false triggering caused by single-frame transient noise, short-time phase jitter or accidental channel frame drop to the order and radial inverse filtering adjustment. In order to suppress the misdirection of high-order HOA (Higher Order Ambisonics) coding caused by channel abnormalities, noise pollution, gain and phase offset.

[0029] As Figure 3As shown, in this embodiment, the reliability of the original signals of each channel is evaluated before encoding to obtain the reliability index of each channel. Within each control time window, factors directly related to the reliability of spatial coding, such as noise level, gain deviation, and phase shift, are evaluated for each microphone channel simultaneously. For example, the equivalent noise energy of a channel is estimated using silent frames or low-energy segments. Specifically, an adaptive threshold can be set based on the statistical distribution of short-time energy of a channel or the joint energy of multiple channels to mark time frames that are significantly lower than the background level as noise estimation samples. To reduce the risk of mistaking weak sound source segments as silent and thus overestimating noise, secondary screening can also be performed by combining cross-channel consistency, bandwidth occupancy sparsity, or historical noise baseline. This results in a more reliable equivalent noise energy estimate, which serves as a unified and consistent baseline in the calculation of channel reliability index and order band signal-to-noise ratio. The online measured amplitude response is compared with the reference response obtained from factory or field calibration to obtain the gain deviation. Then, the phase offset is calculated based on multi-channel synchronization and the reference phase. The reference response describes the frequency response benchmark of the channel under normal installation conditions and without significant faults. The reference phase describes the relative phase consistency benchmark of multiple channels under the same sound field excitation or synchronous sampling conditions. During online operation, the system can estimate the channel amplitude response or characteristic amplitude statistics within a control time window and compare it with the reference response. Line matching is used to obtain the gain deviation, and the relative phase shift is estimated based on the cross-correlation, coherence, or reference source calibration segment of multiple channels. This ensures that the calculation of gain deviation and phase shift has a clear reference source and a repeatable calibration path. The above multi-dimensional evaluation quantities are normalized and fused into a confidence index in the range of 0 to 1 according to preset weights. This index is used to characterize the reliability of the channel in sampling the real sound field within the current time window, thereby providing a quantitative basis for subsequent soft preservation, soft attenuation, or hard rejection. This avoids abnormal channels being amplified into high-order pseudo-spatial structures after entering the spherical harmonic coding stage. In this embodiment, noise level, gain deviation, and phase shift are constructed as normalized evaluation quantities, and a weighted fusion method is used to obtain... The channel reliability index is derived by characterizing the noise term as a monotonically mapped quantity of the ratio of the equivalent noise energy estimated based on the silent frame to the energy of the current frame, the gain deviation term as a monotonically penalized quantity of the deviation of the online amplitude response relative to the reference response, and the phase offset term as a monotonically penalized quantity of the offset of the channel relative to the reference phase model or the multi-channel consistency model. After normalization, the above evaluation quantities are fused according to preset weights to obtain the channel reliability index in the range of 0 to 1. This index reflects the combined impact of noise-dominated risk, gain mismatch risk, and phase instability risk on the reliability of spatial coding, thereby ensuring that the basis for acoustic vector weighting has an interpretable physical meaning and a reproducible engineering calculation path.After obtaining the confidence index of each channel, the system maps the confidence index to a sound vector weighting coefficient based on preset channel preservation thresholds and channel rejection thresholds. This weighting coefficient, ranging from 0 to 1, represents the retention strength of the channel in spatial coding within the current control time window. When the confidence index is higher than the preservation threshold, the corresponding weighting coefficient takes a larger value to maintain the complete sound pressure contribution of the channel. When the confidence index is lower than the rejection threshold, the corresponding weighting coefficient is zero to achieve hard rejection. When the confidence index is between the two thresholds, the corresponding weighting coefficient changes smoothly in a monotonically continuous manner to achieve soft attenuation and avoid abrupt weight changes. Subsequently, the system uses the multi-channel sound pressure vector obtained from time-frequency analysis as the weighting object, matching and multiplying the sound pressure components of each channel in the sound pressure vector with their corresponding sound vector weighting coefficients one by one. This ensures that channels with higher confidence still maintain their main energy contribution after weighting, while channels with lower confidence are significantly suppressed after weighting, thus obtaining a weighted multi-channel sound pressure vector for spatial coding. Then, it is multiplied with the pre-calibrated HOA encoding matrix to obtain a weighted multi-order spherical harmonic coefficient vector, which is then concatenated in time order to form a weighted spherical harmonic coefficient sequence as HOA three-dimensional spatial audio data.

[0030] like Figure 4As shown, for the implementation using a rigid spherical array structure, this embodiment further aggregates the weighted spherical harmonic coefficients into HOA sub-vectors of each order according to the spherical harmonic order. Within a preset control time window, the signal energy and noise energy of each order in each frequency band are statistically analyzed to estimate the frequency band signal-to-noise ratio (SNR) index of the order. Combined with channel reliability, an order reliability index is obtained, and an order-related radial back filter adjustment factor is generated from the order reliability index. The radial back filter gain of the original design is adaptively weakened or retained to avoid excessive amplification of noise and spatial aliasing artifacts by the radial back filter in high-order, low SNR, or physically unusable frequency bands. This results in the output of a corrected standard high-order spherical harmonic coefficient sequence. The order frequency band SNR index can be defined based on the intra-order aggregated energy of the weighted spherical harmonic coefficients as the order signal energy obtained by summing the energy of all components of the same order or robust statistics within the control time window. The noise energy of the same order and frequency band estimated by the silent frame or low-energy segment is used as the signal energy. A baseline is established, thus forming a signal-to-noise ratio (SNR) estimate corresponding to both the order and frequency band indices. This index reflects the observability of the true sound field details at the target frequency band for that order, and also provides a frequency band-level risk assessment basis for the regularization intensity of the radial back filter adjustment factor. This allows high-order, low SNR regions to be clearly marked as high-risk noise amplification regions and prioritized for suppression during the compensation phase. The channel credibility aggregation result can be expressed as a weighted average or robust average of the credibility indices of the effective channels involved in the coding of that order. The order-frequency band SNR index is used to describe the observable upper bound of the true sound field at the target frequency band for that order. By normalizing and fusing the above two types of evidence, an order credibility index that varies with the order and frequency band can be obtained. This index can simultaneously characterize the dual conditions of whether the source channel is reliable and whether the order frequency band has real interpretable energy, thus providing a unified credibility scale for radial back filter adjustment, construction of the true spatial complexity index, and configuration of order and code rate linkage.Based on the HOA data after weighted and radial gain adaptive correction, this embodiment obtains energy sequences from order 0 to the highest order within the control time window. The higher-order energies are weighted using order reliability to construct a true space complexity index, which is then compared with the input task requirement complexity index. The system dynamically adjusts the configuration order and configuration code rate of the HOA accordingly. Specifically, the adjustment process is as follows: when there is a positive margin between the reliable higher-order energy and the task requirement (i.e., within the current control time window), the reliable higher-order energy obtained after channel weighting and radial inverse filtering gain adaptive correction is greater than the task requirement, and the difference between the reliable higher-order energy and the task requirement is positive and exceeds a preset stable threshold. When the margin threshold is reached, the positive margin refers to the degree of availability of reliable high-order spatial information obtained after adaptive correction of channel weighting and radial back filtering gain relative to the task requirements. The configuration order is increased or maintained at a higher order to meet the spatial resolution requirements. When the high-order energy is mainly contributed by low-reliability channels or low signal-to-noise ratio frequency bands, or when the deviation indicates that the spatial complexity is lower than the minimum requirements of the task, the high-order radial gain is reduced first, the configuration order is lowered, and the configuration bit rate is reduced simultaneously. This is done to reduce invalid high-order data and computing power consumption while ensuring the availability of spatial information for the task. This is achieved through the closed-loop mechanism of the adaptive order and bit rate linkage of the channel-level weighted order frequency band radial back filtering.

[0031] It should be noted that within the control time window, the system uses the spherical harmonic order as a grouping index to perform intra-order separation and aggregation of multi-order spherical harmonic coefficients in the spherical harmonic coefficient sequence. Specifically, based on a predetermined spherical harmonic index order, the system analyzes the multi-order spherical harmonic coefficient vectors of each time frame or time frequency unit within the time window, identifies all coefficient components belonging to the same order, and extracts and rearranges the coefficients within that order in a predetermined order to form the spherical harmonic coefficient sub-vectors of that order. Subsequently, the system aligns and merges the spherical harmonic coefficient sub-vectors of the same order in all time frames within the control time window according to timestamps or time frame indices, forming a set of order sub-vectors or a sequence of order sub-vectors within the control time window. This allows the spatial information contribution of different orders to be explicitly split from the original full-order mixed coefficients into independent and interpretable order subspace representations, thereby performing energy statistics and constructing order energy sequences within the target frequency band.

[0032] It should also be noted that when the system converts the microphone's position into azimuth and elevation angles in spherical coordinates based on its installation location information in three-dimensional space, the array geometric center or a preset reference point can be used as the origin to read the three-dimensional coordinates (x, y, z) of each microphone in the installation calibration file or mechanical design drawing. i ,y i ,z i ), and calculate its radial distance from the origin. Based on this, the projection vector located in the horizontal plane is used as the reference for azimuth angle calculation to obtain the azimuth angle. The degree of microphone elevation relative to the horizontal plane can be characterized as the elevation angle, which can be taken as... Alternatively, an equivalent definition could be used, such that each microphone obtains a set of (ϕ) values ​​related to the array center. i ,θ i The angle parameter provides a unified geometric input for subsequent evaluation of spherical harmonic basis functions and construction of sampling matrices.

[0033] After obtaining the azimuth and elevation angles corresponding to each microphone, when constructing the spherical harmonic sampling matrix by calculating the corresponding spherical harmonic basis function values ​​based on the preset maximum order Lmax, the azimuth and elevation angles can be calculated for each microphone position (ϕ). i ,θ i Calculate the spherical harmonics from order 0 to Lmax sequentially. The sampling vectors are then arranged in a predetermined order for all spherical harmonic functions (l,m) of the same microphone. The sampling vectors of the M microphones in the array are then stacked row-wise to form a sample vector of size [size missing]. The spherical harmonic sampling matrix describes the discrete sampling relationship of the ideal spherical harmonic sound field basis at a given order in the array geometry. Specifically, it can be constructed by offline calibrating the sound field, scanning a known reference sound source, or normalizing the calibration signal to establish the mapping relationship between the actual multi-channel response and the ideal spherical harmonic basis functions. Then, a coding matrix consistent with the array geometry is obtained by using least squares, constrained pseudo-inverse, or regularized stable solution strategies. This allows the subsequent channel confidence-driven weighted coding to accurately correspond geometrically to the interpretable link from microphone space sampling to spherical harmonic domain coefficient estimation. It can also correspond structurally with the HOA coding matrix obtained by subsequent calibration, giving the coding process clear geometric interpretability and calibrability.

[0034] In this embodiment one, the spherical harmonic function It can adopt any convention of real spherical harmonics or complex spherical harmonics, and can pre-fix the ordering rules and normalization specifications within the order to ensure the consistency of the encoding matrix, decoding matrix and subsequent order energy statistics; for example, a unified ordering convention compatible with the engineering HOA implementation can be adopted to arrange the m values ​​of each order in a predetermined order and maintain a one-to-one correspondence with the column index of the encoding matrix. At the same time, the same amplitude scale as the calibration stage is maintained in the normalization method, thereby avoiding the problems of order energy deviation, order reliability assessment distortion or unstable decision basis for order promotion or demotion caused by inconsistent spherical harmonic conventions.

[0035] The true spatial complexity index in this embodiment can be composed of the energy sequence from order 0 to the highest order within the control time window. Order reliability weighting is introduced for higher-order energies to suppress false complexity increases caused by channel pseudo-structures and low signal-to-noise ratio frequency bands. For example, each order energy and its corresponding order reliability can be multiplied or weighted and fused to form the overall complexity metric, making the index closer to the scale of spatial details in a real sound field that can be reliably sampled and safely compensated. Correspondingly, the task requirement complexity index can be obtained by mapping the target application's minimum requirements for spatial resolution and frequency band sensitivity. For example, a configurable requirement curve can be generated from the minimum spatial order requirements for speech intelligibility, sound source localization error upper limit, or immersive rendering. This ensures that the deviation between true complexity and task requirements not only has interpretable engineering significance but also directly drives the controllable adjustment of the configuration order and configuration bitrate.

[0036] Example A: Within each control time window, the system first performs a comprehensive evaluation of the noise level, gain deviation, and phase shift of the raw signal from a certain microphone channel to obtain the channel's reliability index. When the channel's reliability index is greater than or equal to the channel's full preservation threshold, the system determines that the channel's sampling reliability of the real sound field within the current time window is sufficient. Therefore, the weighting coefficient of the corresponding acoustic vector of the channel is directly set to 1, so that the sound pressure component corresponding to the frequency unit of the channel participates in the construction of the multi-channel sound pressure vector without being weakened. Subsequently, the system multiplies the weighted multi-channel sound pressure vector with the pre-calibrated HOA encoding matrix to obtain the weighted multi-order spherical harmonic coefficient vector of the time and frequency unit, and outputs and concatenates them in chronological order to form a weighted spherical harmonic coefficient sequence as HOA three-dimensional spatial audio data. This ensures that the contribution of the high-reliability channel to spatial information is fully preserved and stably supports subsequent energy statistics and spatial complexity assessment.

[0037] Example B: The system first performs a credibility assessment on the target channel and obtains its credibility index. When the credibility index of the channel is less than or equal to the channel removal threshold, the system determines that the channel is in a state of obvious abnormality or strong noise dominance. If it continues to participate in the encoding, it will significantly increase the risk of high-order pseudo-spatial structure injection. Therefore, the system sets the acoustic vector weighting coefficient of the channel to 0, so that it is effectively removed from the multi-channel sound pressure vector. On this basis, the system only uses the weighted sound pressure components of the remaining effective channels to perform a product operation with the HOA encoding matrix to obtain a more reliable weighted multi-order spherical harmonic coefficient vector, and forms a weighted spherical harmonic coefficient sequence in chronological order. This achieves source blocking of the influence of low-credibility channels and avoids the abnormal channel energy from being amplified into high-order energy anomalies or spatial aliasing artifacts through encoding and radial compensation.

[0038] Example C: When the credibility index C of a certain channel falls between the channel preservation threshold C1 and the channel rejection threshold C2, and the channel preservation threshold is greater than the channel rejection threshold, the system adopts a continuously adjustable soft weighting mechanism to balance information retention and risk suppression, so that the acoustic vector weighting coefficient of the channel decreases smoothly as credibility decreases. Optionally, the acoustic vector weighting coefficient can be calculated in an exponential mapping form, with the specific expression as follows:

[0039] ;

[0040] Where w is the acoustic vector weighting coefficient and n is the exponential coefficient, a preset empirical parameter determined based on array type, number of channels, tolerance of the target task to high-order pseudostructures, and ambient noise level. It is used to adjust the sensitivity of weight attenuation; a larger n indicates more aggressive suppression of low-to-medium confidence channels, while a smaller n indicates a greater tendency to retain the remaining effective information of suspicious channels. In this way, although channels are not hard-canceled, their contribution to the multi-channel sound pressure vector is moderately weakened according to confidence level. Subsequently, the system multiplies the weighted multi-channel sound pressure vector with the pre-calibrated HOA encoding matrix to obtain the weighted multi-order spherical harmonic coefficient vector for that time-frequency unit. Using the sampling timestamp or time frame index within the control time window as the sorting criterion, the weighted multi-order spherical harmonic coefficient vectors corresponding to each time frame are concatenated in ascending order from early to late. When outputting across multiple control time windows, the weighted multi-order spherical harmonic coefficient vectors are concatenated according to the time of each control window. The window-level sequence is spliced ​​by incrementing the starting timestamp of the time window to form a weighted spherical harmonic coefficient sequence with clear time index consistency as HOA three-dimensional spatial audio data. This allows channels that are questionable but still valuable to be included in the encoding process in an interpretable way, reducing their misleading influence on higher-order judgments while retaining their potential effective sampling contribution to the real sound field. The weighted spherical harmonic coefficient vectors corresponding to each time frame are concatenated according to the incrementing sampling timestamp. While maintaining one-to-one alignment with the original multi-channel pickup frames, discrete instantaneous estimates can be organized into a continuously analyzable temporal object. This facilitates the performance of energy aggregation, robust statistics, order band signal-to-noise ratio estimation, and order credibility calculation on coefficients of the same order within a preset control time window. It also provides a stable, reproducible, and streaming-suitable temporal basis for subsequent radial inverse filtering adjustment, construction of real spatial complexity indicators, and configuration of closed-loop adaptive order / bitrate.

[0041] As shown in Examples A, B, and C, the coding adjustment method based on sound vector weighting coefficients prioritizes the channel reliability assessment results. Before coding, it applies hierarchical weighting to the multi-channel sound pressure vectors, allowing high-reliability channels to participate in sound pressure vector construction with full weight to fully preserve effective spatial sampling contributions. Low-reliability or abnormal channels are effectively removed from the sound pressure vectors with zero weight to block the risk of pseudo-spatial structure injection. Suspicious channels between two thresholds are smoothly weakened through continuously adjustable soft weights to balance information preservation and risk suppression. Based on this, the system inputs the weighted multi-channel sound pressure vectors into a pre-calibrated high-order panoramic sound coding matrix to obtain the weighted multi-order spherical harmonic coefficients of the corresponding time and frequency units, which are then concatenated in chronological order to form a weighted spherical harmonic coefficient sequence. This achieves a coding adaptive adjustment mechanism based on channel reliability, using weighted pre-suppression as a means, and aiming to improve the authenticity and stability of high-order spatial information. This allows subsequent energy statistics, spatial complexity assessment, and order and bit rate linkage control to be established on a more reliable spatial data foundation.

[0042] Example D: In the adaptive attenuation scenario of radial back filtering, after the system aggregates the weighted spherical harmonic coefficients by order, it is statistically found within the control time window that the signal energy of a certain high-order spherical harmonic coefficient in the target frequency band is only slightly higher than the noise energy of the silent frame, resulting in a low frequency band signal-to-noise ratio for that order. At the same time, the mean channel confidence level corresponding to that order is at a general level or there is short-term instability in some local channels. At this time, the order confidence index is evaluated as a medium-low value. Based on this, the system generates a small radial back filtering adjustment factor to significantly reduce the radial gain of the original design. Even if the original energy of that order shows a local increase, it will be judged as a high-risk area of ​​potential high-order energy anomalies, thereby avoiding the radial back filtering from further amplifying noise and spatial aliasing artifacts, and ensuring that the output standard high-order spherical harmonic coefficient sequence has a more robust noise upper bound and a lower probability of pseudo spatial structure injection in the high-order region.

[0043] Example E: In the adaptive preservation scenario of radial back filtering, the system observes that the signal energy of a certain mid-to-high order in the critical task frequency band is significantly higher than the noise energy of the silent frame within the same control time window. The signal-to-noise ratio of the order in the band is at a sufficient level, and the overall reliability of the channels involved in the encoding of this order is high, with gain and phase shift within a controllable range. At the same time, this frequency band is located within the array's physically available bandwidth and the theoretical effective order limit and is marked as a task-sensitive frequency band. In this case, the order reliability index is evaluated as a high value, and the system generates a radial back filtering adjustment factor close to 1, so that the radial back filtering gain of this order in this frequency band basically retains the original design compensation strength, so as to fully correct the scattering effect of the rigid sphere array and maintain the resolution capability of this order for real spatial details, thereby maintaining stable support for the spatial resolution required by the target task without introducing additional high-order noise amplification.

[0044] Examples D and E above illustrate different adjustment directions for radial back filtering. As can be seen from Examples D and E, this invention adjusts radial back filtering by estimating the frequency band signal-to-noise ratio (SNR) index of the order, combining it with channel reliability to obtain an order reliability index, and then generating an order-related radial back filtering adjustment factor based on the order reliability index. When higher-order frequency bands exhibit low reliability indices, channel reliability is average, or short-term instability exists, the system significantly weakens the original radial gain design using a smaller radial back filtering adjustment factor to suppress the amplification of noise and spatial aliasing artifacts and avoid the injection of pseudo-spatial structures. When higher-order frequencies have high reliability indices at mission-critical frequencies, the channels are generally reliable, and the system is within the effective order limit of the array and the physically available bandwidth, the system generates an adjustment factor close to 1 to retain the original design compensation strength, thereby ensuring scattering correction effect and spatial detail resolution capability. This achieves a balance between high-order risk suppression and effective spatial information fidelity, making subsequent spatial complexity assessment based on high-order energy and order and code rate linkage adjustments more reliable, interpretable, and stable.

[0045] This first embodiment can eliminate the interference of channel anomalies and noise pseudo-structures on spatial complexity judgment by moving them forward, so that the dynamic adjustment of order and bit rate is more in line with the real sound field structure and task requirements, thereby reducing the risk of false order upgrades, false high bit rate preservation and amplification of high-order noise, and improving the stability and controllability of three-dimensional spatial data acquisition and transmission.

[0046] Embodiment 2 of the present invention: In addition to the adjustment method of constructing sound vector weighting coefficients based on channel reliability to realize pre-weighted HOA coding in Embodiment 1, this embodiment further provides a coding adjustment method based on multi-order reconstruction residual driving. Specifically, within the control time window, the system performs HOA coding according to different candidate orders based on the same multi-channel sound pressure vector to obtain the spherical harmonic coefficient set of the corresponding order, and uses the spherical harmonic sampling relationship matching the candidate order to perform reverse reconstruction of the multi-channel sound pressure, and then calculates the normalized residual between the reconstructed sound pressure and the original sound pressure and the marginal improvement brought about by the increase of the order; when the system detects that it is difficult to significantly reduce the reconstruction residual by continuing to increase the order, or the residual improvement is insufficient to cover the data cost and noise amplification risk introduced by the new order, it determines the minimum sufficient order to meet the current sound field and task requirements, and adaptively adjusts the HOA configuration order and configuration bit rate accordingly. In this way, the order-up behavior is limited to scenarios that can substantially improve the interpretable reconstruction capability of the real sound field, while the order-down behavior corresponds to scenarios where the added higher-order order is insufficient to interpret the real sound field. This effectively distinguishes higher-order energy anomalies, channel pseudo-structures, or noise amplification from changes in real spatial complexity, expanding the basis for coding adjustment from simple energy and SNR thresholds to reconstruction evidence with physical meaning.

[0047] It should be noted that the reconstruction residual is a normalized measure of the energy difference or norm difference between the original sound pressure level of each channel and the reconstructed sound pressure level of the candidate order within the control time window. Furthermore, the marginal improvement in the residual between adjacent candidate orders can be calculated as direct evidence of the interpretive gain of the new order on the real sound field. When the system detects that the marginal improvement is lower than a preset threshold or shows a rapid decay trend among multiple consecutive candidate orders, it can be determined that continuing to increase the order is not valuable enough for the interpretable reconstruction of the real sound field. Then, the minimum sufficient order that meets the current sound field and task constraints is determined and matched with the corresponding bit rate, thereby improving the order selection from a simple energy heuristic to an interpretable decision based on reconstruction evidence.

[0048] Example F: When the array is in a complex sound field with multiple sound sources or strong reflections, and the system performs encoding and reconstruction comparisons on the same multi-channel sound pressure vector at different candidate orders within the same control time window, it will be observed that as the order increases, the degree of fit between the reconstructed sound pressure and the original sound pressure shows a continuous and significant improvement. This improvement is not only reflected in the overall decrease of the residual, but also in the fact that the marginal contribution of the newly added order to the decrease of the residual is still at a high level, thus indicating that higher-order components have real spatial interpretation value. At this time, the system determines that there are spatial details in the current sound field that can be effectively characterized by higher orders, and takes the order that satisfies the significant improvement in reconstruction as the minimum sufficient configuration order available, and matches the corresponding encoding bitrate, so that the order-up behavior has direct evidence to support the ability to interpret and reconstruct the real sound field, rather than blindly increasing the configuration simply because the apparent increase in higher-order energy.

[0049] Example G: When there is a local channel noise increase or short-term phase instability within a certain time window, resulting in an abnormal increase in higher-order energy, after the system performs encoding and reverse reconstruction with different candidate orders, it is often found that the improvement of the reconstruction residual of the newly added higher order on the original sound pressure is very limited, and even the marginal improvement decays rapidly. This indicates that these higher-order energies are more likely to come from noise amplification or channel pseudo-structure rather than the real sound field structure. Therefore, the system constrains the order selection to a range where the residual improvement is insufficient to prove the real contribution of the newly added order, and prioritizes maintaining or regressing to a lower configuration order while simultaneously suppressing the corresponding bitrate. This makes the order reduction behavior explainable by the insufficient interpretation gain of the newly added higher order on the real sound field, thereby achieving targeted correction of the misjudgment of abnormal higher-order energy as a change in real spatial complexity.

[0050] As can be seen from Examples F and G, this second embodiment is consistent with the first embodiment in terms of overall objectives and control framework. Both use the control time window as the basic decision metric and aim to suppress the misleading effect of channel anomalies and noise pseudostructures on high-order panoramic sound coding, and serve the dynamic linkage adjustment of configuration order and configuration bitrate, thereby ensuring the effectiveness and stability of 3D spatial data under the constraints of task requirements. However, the key difference between the two lies in the different types and priorities of evidence used for adjustment. The second embodiment focuses on using channel credibility as a prerequisite constraint and combining order energy and the signal-to-noise ratio of the order's frequency band to form a credibility-weighted energy supply and demand. The criteria make order raising and lowering more of an engineering screening of channel reliability and energy risk. Implementation Example 2 introduces multi-order reconstruction residuals and their marginal improvements as evidence of physically interpretable sound field interpretation gain, in addition to such engineering criteria. This limits order raising behavior to scenarios that can significantly improve the interpretable reconstruction capability of the original multi-channel sound pressure, while order lowering behavior corresponds to scenarios where the newly added higher order does not contribute enough to the interpretation of the real sound field. Thus, when faced with complex situations where the apparent energy of higher orders increases but the real spatial information gain is limited, it can more reliably distinguish between changes in real spatial complexity and energy anomalies caused by noise and pseudostructure.

[0051] Embodiment 3 of the present invention: In addition to the adjustment method in Embodiment 1, which generates a radial back-filter adjustment factor from the order confidence index and scales the original radial back-filter gain, this embodiment further provides a radial back-filter adjustment method based on adaptive regularization stabilization. Specifically, for scattering compensation of rigid sphere arrays, the system no longer simply weakens or retains the existing back-filter gain in an amplitude-based manner, but instead expresses the radial inversion process as a regularized stable back-filter form. Within the control time window, it considers the frequency band signal-to-noise ratio of the order, the channel confidence aggregation result, the order confidence, and the physical available bandwidth and theoretical effective order limit, etc. Information is used to generate regularization intensity parameters related to the order and frequency band. The generation of these parameters can be performed within each control time window, using the frequency band of the order as the basic index unit. First, the system aggregates and statistically analyzes the energy of the same order within the target frequency band based on weighted spherical harmonic coefficients. Then, it combines this with noise energy of the same order and frequency band estimated from silent frames or low-energy segments to form the signal-to-noise ratio (SNR) characterization for that order. Simultaneously, it aggregates the weighted average or robust average of the reliability indices of the effective channels involved in coding for that order to obtain a channel-level reliability characterization. Finally, it combines the aforementioned order-level frequency band SNR characterization, channel-level reliability characterization, and order-level regularization intensity parameters. The credibility index is uniformly normalized to form a frequency band risk assessment quantity of order under a unified dimension. Based on this, the system further introduces constraint information of physically available bandwidth and theoretically effective order limits to mark frequency bands or order regions that exceed the array's physically interpretable range. Additional penalties are applied to the risk assessment quantity of such regions. Subsequently, the system generates corresponding regularization strength parameters based on the frequency band risk assessment quantity of that order using a monotonic mapping method. This results in higher-risk, higher-order, low-SNR, low-credibility, or physically unavailable regions corresponding to larger regularization strengths, thereby more strongly suppressing the amplification of noise by the inverse problem in the stable inverse filtering solution. Frequency units of orders with lower risk and located in the mission-critical frequency band correspond to smaller regularization intensities, making the inverse filtering closer to ideal scattering compensation while preserving the ability to resolve real spatial details. When the system determines that a certain order is in a low signal-to-noise ratio, high aliasing risk, or low confidence state in the corresponding frequency band, the regularization intensity is automatically increased to make the inverse filtering more conservative, thereby compensating for scattering effects while suppressing the amplification of noise by the inverse problem. When the system determines that the order has sufficient signal-to-noise ratio and high order confidence within the mission-critical frequency band, the regularization intensity is decreased to make the inverse filtering closer to ideal compensation, thereby maintaining the ability to resolve real spatial details. Through this adaptive regularization intensity mechanism, the risk control of radial inverse filtering is upgraded from simple gain scaling to interpretable stabilization inverse problem adjustment, further enhancing the ability to suppress the excessive amplification of high-order noise and improving engineering robustness.

[0052] Example H: When the system evaluates that the signal energy of a certain mid-to-high order signal in the target frequency band is only slightly higher than the noise energy of the silent frame within the control time window, and the channel reliability aggregation result of this order is at a general level or shows local short-term fluctuations, the system determines this frequency band as an inversion region with a high risk of noise amplification, and automatically increases the corresponding regularization intensity in the radial inverse filtering calculation, so that the radial inversion changes from an aggressive state that approaches the ideal inverse filtering to a conservative and stable flexible state. This significantly weakens the amplification capability of the inverse problem on noise while still having the significance of scattering compensation, and avoids the high-order coefficients of this frequency band being pushed into aliasing artifacts or pseudo-spatial structures by uncontrolled radial gain.

[0053] Example I: When the system observes a certain mid-to-high-order frequency band with sufficient signal-to-noise ratio in the mission-critical frequency band within the same control time window, and the overall credibility of the channels involved in encoding this order, with gain and phase shift within a controllable range, and this frequency band is within the array's physically available bandwidth and theoretically effective order limit, the system will automatically reduce the regularization intensity of this order frequency band, making the radial inversion closer to the ideal compensation state, so as to fully correct the scattering effect of the rigid sphere array and maintain the ability to resolve real spatial details; it can be seen that this embodiment achieves the stabilization adjustment of radial inverse filtering through adaptive regularization intensity, so that the real spatial information that needs to be retained and the high-order noise risk that needs to be suppressed can obtain a consistent and interpretable engineering processing path within the same inversion framework.

[0054] As can be seen from Examples H and I, this Embodiment 3 and Embodiment 1 share commonalities in the control object and criterion source of radial back-filter adjustment. Both are aimed at the scattering compensation requirements of rigid spherical arrays, and both use interpretable indicators such as the frequency band signal-to-noise ratio of the order, the channel reliability aggregation result, and the order reliability within the control time window to identify high-order high-risk regions and mission-critical reliable regions, thereby achieving risk constraints on the radial inversion amplification effect and faithful support for real spatial details. However, the core difference between the two lies in the different mechanism levels of the adjustment methods. Embodiment 1 uses the adjustment factor generated by the order reliability index to adjust the original radial back filter design. The gain is adaptively reduced or retained in an amplitude-type manner, which is a controllable scaling of the existing inverse filter gain. However, Example 3 further describes the radial inversion process as a stable inverse filter form with regularization, and changes the stability and conservatism of the inversion by adaptively adjusting the regularization intensity. This makes noise suppression no longer rely solely on gain multiplicative scaling, but achieves continuous interpretable control within the same inversion framework by stabilizing the inverse problem, making the high-risk frequency band more conservative and the reliable frequency band closer to the ideal compensation. This results in a more robust noise upper bound and a lower probability of aliasing artifact amplification in scenarios with high-order low signal-to-noise ratio or close to the physical effective order limit.

[0055] In this third embodiment, the stabilized radial inverse filtering can be expressed as an inverse problem with regularization. That is, when solving the inversion gain of rigid sphere array scattering compensation, a regularization intensity parameter related to the order and frequency band is introduced to make the inversion proactively conservative in low signal-to-noise ratio or high aliasing risk regions. The above-mentioned regularization intensity parameter can be generated by mapping the order frequency band signal-to-noise ratio, channel credibility aggregation result and order credibility index fusion result. When the system determines that a certain order is in a low credibility or close to the theoretical effective order limit in the corresponding frequency band, the regularization intensity is increased to suppress the upper bound of noise being amplified by inverse compensation. When the system determines that the order has sufficient signal-to-noise ratio in the mission key frequency band and the channel is stable as a whole, the regularization intensity is decreased to make the inverse filtering closer to the ideal compensation, thereby achieving continuous interpretable adjustment of risk suppression and detail fidelity within the same mathematical framework.

[0056] Embodiment 4 of the present invention: Combining Embodiments 2 and 3, this embodiment provides a three-dimensional spatial data dynamic sampling control system and method. While maintaining the channel reliability assessment, acoustic vector weighting, and order energy-task requirement linkage framework of Embodiment 1, it introduces two enhancement mechanisms: multi-order reconstruction residual evidence and adaptive regularized radial back filtering. This forms a multi-evidence closed-loop control across channel, order, and inversion stabilization dimensions. Within the control time window, the system determines the minimum sufficient configuration order to satisfy the current sound field and task constraints by using the reconstruction residuals and marginal improvement magnitude of candidate orders, and simultaneously matches the configuration bitrate. On the other hand, it adaptively adjusts the regularization intensity of the radial back filtering through order reliability, the frequency band signal-to-noise ratio of the order, and physical effectiveness constraints. This ensures that the formation of higher-order coefficients is subject to both the pre-constraint of source channel reliability and the reconstruction evidence constraint of order selection, while simultaneously achieving stable noise suppression protection during the radial compensation stage. Therefore, this embodiment, while ensuring the required spatial resolution and the fidelity of key frequency band information, further reduces the risks of incorrect order upscaling, incorrect high bit rate preservation, and high-order pseudo-energy injection caused by channel anomalies, noise pseudo-structures, or instability in rigid sphere inversion.

[0057] As can be seen from Example 4, Example 4 and Example 1 are consistent in their technical purpose and control objectives. Both aim to reduce the risk of channel anomalies and noise pseudostructures misleading the judgment of high-order panoramic sound coding and spatial complexity in the three-dimensional spatial data acquisition process of multi-unit intelligent microphone arrays, and to achieve dynamic optimization of configuration order and configuration bit rate while meeting the spatial resolution required for the task, while suppressing the excessive amplification of high-order noise and spatial aliasing artifacts by the radial compensation of rigid sphere arrays. However, the two methods of achieving the above-mentioned common objectives differ. Example 1 mainly uses channel reliability-driven hierarchical weighting of acoustic vectors, order reliability-driven adaptive weakening or retention of radial inverse filtering gain, and base The system uses the linkage control of credibility-weighted order energy and task demand supply-demand deviation to form an interpretable, configurable, and engineering-friendly closed-loop adjustment path. Implementation 4, while using the above basic framework, further introduces multi-order reconstruction residuals and their marginal improvements as interpretable supplementary evidence for order increases and decreases. It also adopts an adaptive regularized and stabilized radial inverse filtering adjustment method to replace or enhance the simple gain scaling path. This achieves a strengthened solution to the same technical problem by using reconstruction to explain gain constraints on order and stabilization to constrain radial amplification. This expands the adjustment of order and bit rate, as well as the suppression of high-order noise risks, from a path dominated by a single engineering criterion to a composite control path with more sufficient evidence and stronger robustness.

[0058] This fourth embodiment allows setting evidence priorities or consistency verification rules. When there is an inconsistency between the order energy supply and demand criterion weighted by channel credibility and the multi-order reconstruction residual evidence, control conflicts can be avoided. For example, when the reconstruction residual shows that the interpretable gain of the newly added order is insufficient while the order energy appears high, the reconstruction evidence should be adopted first to suppress the risk of false order upgrades caused by noise or pseudo-structures. When both the reconstruction evidence and the order credibility show that a certain order has a real contribution in the key frequency band of the task, the order energy supply and demand criterion can then be allowed to trigger order upgrades and bit rate increases. This ensures that the multi-evidence closed-loop control can maintain an agile response to real complex sound fields and provide stronger interpretable constraints on false high-order values ​​caused by channel anomalies and radial inversion instability.

[0059] All embodiments of the present invention employ a rigid spherical array structure. For embodiments employing a rigid spherical array structure, the aforementioned multi-unit intelligent microphone array refers to the uniform or quasi-uniform installation of multiple pickup units along the spherical surface on the surface of a rigid spherical shell or its adjacent region, with the spherical shell serving as a stable acoustic boundary to form interpretable scattering and radial response characteristics. Under this structure, the spherical harmonic sampling relationship and scattering compensation of the array have clear physical interpretability, enabling the radial back-filter to establish a calculable correspondence with the spherical harmonic order, frequency band, and physical effective order limit under known boundary conditions in the form of inverse compensation. Therefore, it is more suitable to introduce radial back-filter adaptive attenuation or retention driven by the order band signal-to-noise ratio and order credibility within the framework of the present invention, so as to systematically suppress the excessive amplification of noise and spatial aliasing artifacts in high-order, low signal-to-noise ratio, or physically unusable frequency bands. It should be noted that the present invention does not use rigid spherical arrays as the only implementation method. In addition to rigid spherical arrays, open-aperture acoustic transparent spherical arrays, near-free-field spherical arrays, double-layer spherical arrays, or other array structures that meet the requirements of three-dimensional spherical sampling can also be used. When the above-mentioned non-rigid boundary or array form that is closer to the free field is used, the construction of the spherical harmonic sampling matrix and the channel credibility weighting mechanism before encoding can still be consistent. However, the source of radial compensation, the physical effective order limit, and the inversion stability constraint will change. At this time, the radial adjustment logic generated based on order credibility in the present invention can be migrated to the gain constraint or regularization intensity constraint on the radial kernel of the free field or mixed boundary, so that the system can still achieve interpretable suppression of the risk of high-order noise amplification. In comparison, the main reason for choosing a rigid sphere array as the preferred implementation method is that its structural boundaries are stable, its scattering effects are predictable, and it is easy to calibrate. It can provide a more consistent geometric and acoustic benchmark for high-order spherical harmonic coding, and it can make the closed-loop mechanism of channel-level weighting, order reliability, radial compensation risk control, and order and code rate linkage of the present invention easier to achieve consistency verification and effect reproduction in engineering. If an array structure with unclear boundaries or weaker scattering characteristics is used instead, and the radial back filtering of rigid spheres is still used, it may lead to insufficient compensation or mismatch, resulting in increased reconstruction error and order energy judgment deviation. Therefore, it is necessary to reselect or adaptively fuse the corresponding radial compensation according to the array physical boundary to ensure that the risk suppression and spatial detail fidelity goals of the present invention can be equivalently achieved under different array configurations.

[0060] like Figure 2The diagram shown is a flowchart of a three-dimensional spatial data dynamic sampling control method provided in an embodiment of this application, including: S1. Acquiring the original signals of each microphone channel in the microphone array, and performing time-frequency analysis on the original signals of each microphone channel to obtain the sound pressure vector of each microphone channel; S2. Optimizing the HOA encoding method, encoding the sound pressure vector of each microphone channel according to the optimized HOA encoding method to obtain a multi-order spherical harmonic coefficient vector to form a spherical harmonic coefficient sequence; S3. Adjusting the radial back filter, and then obtaining the order energy sequence based on the spherical harmonic coefficient sequence, weighting the high-order energy in the order energy sequence to construct a spatial complexity index, comparing it with the input task requirement complexity index, and dynamically adjusting the configuration parameters of the three-dimensional spatial data dynamic sampling.

[0061] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0063] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0064] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0065] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A three-dimensional spatial data dynamic sampling control system, characterized in that, include: The signal acquisition module is used to acquire the raw signals of each microphone channel in the microphone array, and to perform time-frequency analysis on the raw signals of each microphone channel to obtain the sound pressure vector of each microphone channel. The spherical harmonic coding module is used to optimize the HOA coding method. It encodes the sound pressure vector of each microphone channel according to the optimized HOA coding method to obtain multi-order spherical harmonic coefficient vectors to form a spherical harmonic coefficient sequence. The dynamic control module is used to adjust the radial back filter, then obtain the order energy sequence based on the spherical harmonic coefficient sequence, and construct a space complexity index by weighting the higher-order energies in the order energy sequence. This index is then compared with the input task requirement complexity index to dynamically adjust the configuration parameters for dynamic sampling of three-dimensional spatial data. The original signal refers to the pre-encoded basic sampling signal obtained by each channel of the microphone array under the same acquisition time base, including synchronous digital pickup signal, channel acquisition configuration and channel status information, silence frame and low energy segment signal.

2. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The time-frequency analysis of the raw signals from each microphone channel is performed as follows: S11. Denoise the original signal to remove static noise and low-frequency interference. Use the short-time Fourier transform method to convert the original signal between the time domain and the frequency domain, obtain the frequency components in each time frame, and calculate the amplitude and phase information of each frequency band. S12. Based on the preset frequency range and bandwidth, analyze the energy distribution of each frequency component to obtain the sound pressure response of each channel in all combinations of time period and frequency. S13. The sound pressure response of each channel at all combinations of time period and frequency is fused to obtain a multi-channel sound pressure vector.

3. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The optimization of the HOA encoding method includes: A sound vector weighted optimization method based on channel credibility and a multi-order reconstruction residual optimization method based on candidate order; The channel credibility-based acoustic vector weighting optimization method is used to construct acoustic vector weighting coefficients based on the channel credibility index before encoding; The multi-order reconstruction residual optimization method based on candidate order is used to compare the reconstruction residuals and marginal improvement of residuals of different candidate orders within a control time window.

4. The three-dimensional spatial data dynamic sampling control system as described in claim 3, characterized in that: The specific optimization process of the acoustic vector weighted optimization method based on channel reliability and the multi-order reconstruction residual optimization method based on candidate order is as follows: The channel reliability-based acoustic vector weighting optimization method obtains the channel reliability index based on the evaluation results of silence frames, gain deviation and phase shift in the original signal within the control time window. The channel reliability index is compared with the channel full preservation threshold and the channel rejection threshold to determine the acoustic vector weighting coefficient of graded or continuous mapping, and the multi-channel sound pressure vector is weighted to optimize the HOA coding method. The multi-order reconstruction residual optimization method based on candidate orders optimizes the HOA coding method by performing spherical harmonic coding and reverse reconstruction on the multi-channel sound pressure vectors within the same control time window according to multiple candidate orders, calculating the normalized reconstruction residual of each candidate order and the marginal improvement of the residual between adjacent candidate orders, and determining the minimum sufficient order to meet the current sound field and task requirements based on the criterion that the residual decrease slows down or is lower than a preset threshold.

5. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The sound pressure vector of each microphone channel is encoded according to the optimized HOA encoding method. The specific encoding process is as follows: S21. Within the control time window, call or generate a HOA coding matrix that matches the microphone array geometry and the preset maximum order; S22. Determine the weighting coefficients of the acoustic vectors corresponding to each channel based on the channel credibility index, and weight the multi-channel sound pressure vectors under the same time frame and the same frequency to obtain the weighted multi-channel sound pressure vectors; S23. Perform matrix operations on the weighted multi-channel sound pressure vector and the HOA coding matrix to obtain a multi-order spherical harmonic coefficient vector covering the zeroth order to the preset maximum order; S24. Output the multi-order spherical harmonic coefficient vectors corresponding to each time frame and frequency in chronological order and cascade them to form a spherical harmonic coefficient sequence.

6. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The adjustment of the radial inverse filter includes: Radial gain scaling adjustment based on order confidence index and stable radial inversion adjustment based on adaptive regularization strength; The radial gain scaling adjustment method based on the order confidence index is used to adaptively weaken or retain the original radial back filter gain. The stabilized radial inversion adjustment method based on adaptive regularization intensity is used to express radial inverse filtering as a stable inversion with regularization.

7. The three-dimensional spatial data dynamic sampling control system as described in claim 6, characterized in that: The specific adjustment process of the radial gain scaling adjustment method based on the order confidence index and the stable radial inversion adjustment method based on the adaptive regularization strength is as follows: The radial gain scaling adjustment method based on the order confidence index obtains the order sub-vector by aggregating the spherical harmonic coefficient sequence by order. Within the control time window, the signal-to-noise ratio of the order in the frequency band is statistically analyzed and combined with the aggregation result of the channel confidence index to obtain the order confidence index. The order confidence index generates a radial back filter adjustment factor that varies with the order and frequency band, which weakens or retains the original radial back filter gain. The stable radial inversion adjustment method based on adaptive regularization intensity generates regularization intensity parameters related to the order and frequency band by using the frequency band signal-to-noise ratio of the order, the channel credibility aggregation result, the order credibility index, and the information of physical available bandwidth and theoretical effective order limit. In low credibility or high risk areas, the regularization intensity is increased to make the inversion more conservative, while in mission-critical and high credibility areas, the regularization intensity is decreased to make the inversion approach ideal compensation.

8. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The process of constructing a space complexity index by weighting higher-order energies in the energy sequence is as follows: Within the control time window, multi-order spherical harmonic coefficients are grouped and aggregated according to their order based on the spherical harmonic coefficient sequence to obtain the spherical harmonic coefficient sub-vectors corresponding to each order. Energy statistics are performed on the spherical harmonic coefficient subvectors corresponding to each order within the control time window to form an order energy sequence from zero to the maximum order; The order confidence index is obtained based on the aggregation results of the channel confidence index and the signal-to-noise characteristics of each order in the target frequency band; The order energy within the preset high-order range is multiplicatively or weighted and fused with the corresponding order credibility index to obtain the credibility-weighted high-order energy. The higher-order energy and lower-order energy weighted by the credibility are fused according to a preset rule to construct a spatial complexity index for characterizing the scale of real interpretable spatial details.

9. The three-dimensional spatial data dynamic sampling control system as described in claim 1, characterized in that: The configuration parameters for dynamically adjusting the three-dimensional spatial data sampling are specifically adjusted as follows: Within the controlled time window, the space complexity metric is compared with the input task requirement complexity metric; When the space complexity index has a positive margin relative to the task requirement complexity index and the order reliability index corresponding to the higher order energy meets the preset reliability conditions, the configuration order is increased or maintained, and the configuration code rate is increased or maintained synchronously according to the correspondence between the configuration order and the configuration code rate. When the space complexity index is less than the task requirement complexity index or the order reliability index corresponding to the higher energy is lower than the preset reliability condition, the risk of higher order amplification is suppressed by radial back filtering, and the configuration order is reduced and the configuration code rate is reduced simultaneously. When the complexity deviation index obtained by subtracting the task requirement complexity index from the space complexity index is within the preset stable range, the current configuration order and configuration code rate remain unchanged.

10. A method applied to a three-dimensional spatial data dynamic sampling and control system according to any one of claims 1-9, characterized in that, include: S1. Obtain the raw signals of each microphone channel in the microphone array, and perform time-frequency analysis on the raw signals of each microphone channel to obtain the sound pressure vector of each microphone channel; S2. Optimize the HOA encoding method, and encode the sound pressure vector of each microphone channel according to the optimized HOA encoding method to obtain the multi-order spherical harmonic coefficient vector to form a spherical harmonic coefficient sequence; S3. Adjust the radial back filter, then obtain the order energy sequence based on the spherical harmonic coefficient sequence, weight the higher-order energies in the order energy sequence to construct a space complexity index, compare it with the input task requirement complexity index, and dynamically adjust the configuration parameters of the dynamic sampling of the three-dimensional spatial data.

Citation Information

Patent Citations

  • A method and device for digital sampling synchronization conversion in intelligent substations

    CN109596949B

  • Multi-channel multi-trigger mode synchronous data acquisition method and system

    CN118210266B